Model or dataset
strelov1/freehire avatar
strelov1/freehire

freehire: a self-hosted job search engine that crawls company ATS pages

freehire — the open-source search engine for job seekers

795 stars121 forksGoMIT

At a glance

What is it?
freehire is an MIT-licensed Go and SvelteKit job aggregator that pulls postings from ATS platforms like Workday, Greenhouse and Lever and exposes them through Meilisearch. It is built for people who want the raw catalogue, not another hosted job board.
Who is it for?
Adopt freehire if you want a job catalogue you control and can query over HTTP, and you are willing to run Postgres, Meilisearch, Redis and object storage. Do not adopt it if you need a hosted board with no ops work, or if your listings are not on ATS platforms the harvest layer already covers.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem freehire addresses: job data that has been copied too many times

Most job boards are aggregators of aggregators. A role is posted on a company's Workday tenant, scraped by one site, reposted by another, and eventually appears in a search result that links to a page which links to the original. By that point the listing may be months stale, the apply button may go nowhere, and a recruiter's name sits between the candidate and the employer. freehire takes the opposite route: it crawls company career pages directly, through the ATS platforms that host them. The README names Workday, Greenhouse, Lever, Ashby and iCIMS among 92 ATS platforms, with 225 live sources in total and a long tail of aggregators and direct feeds.

The audience is narrower than a consumer job board. This is for engineers and data people who want the catalogue itself: a queryable set of postings in their own Postgres, with a clean HTTP API on top. The hosted instance at freehire.me exists, but the repository is the product. The README is explicit that the pipeline and the data are both in the open, and that adding a company is one line of YAML. That is the pitch: a job index you can fork, extend and re-crawl on your own schedule rather than a feed you rent.

How the ingest pipeline and search stack fit together

The repository layout shows a Go module at the root with several binaries under cmd/: server, ingest, enrich, reindex, tg-ingest, tg-extract, backfill-derive, liveness, notify, import-collections, recount-companies and migrate. The Dockerfile builds them into one image, with the HTTP server as the default entrypoint and the run-once workers invoked separately, which the Dockerfile comment describes as a scheduled `docker compose run --rm app /app/<worker>` in production. So ingestion is a batch job, not a long-running crawler daemon.

Storage is split by purpose. PostgreSQL with pgvector holds the normalized postings, filtering data and semantic embeddings; the migrations folder is mounted into the Postgres init directory so schema setup happens on first volume creation. Meilisearch backs full-text and faceted search for the search endpoint and the reindex command, and the compose file notes that its huggingFace embedder downloads and runs the model inside the container, caching it in a named volume. Redis is described as a required dependency for the shared rate limiter, not optional like Meilisearch. Go dependencies include chromedp for browser-driven crawling, a TLS client, sqlc-generated database access, langchaingo for LLM calls over any OpenAI-compatible endpoint, and Bifrost as an LLM gateway with per-user virtual keys. The frontend under web/ is SvelteKit 2 with Svelte 5 runes and Tailwind 4, server-rendered.

The data model claim worth noting is deduplication. The README states that the same role posted to three boards collapses into one entry, normalized into a single shape and deduplicated on a stable key. Facets for region, work mode, seniority, skills and salary come from curated dictionaries rather than inference. That last point is the design stance of the project: derived attributes are looked up, not guessed.

Installing freehire with Docker Compose and running a first query

The README's quick start assumes Docker and a Makefile. One command builds and starts the whole stack: api, web, postgres, meilisearch, redis and minio.

bash
make up        # build + start the whole stack in Docker:
               # api, web, postgres, meilisearch, redis, minio
curl localhost:8080/health
curl localhost:8080/api/v1/jobs

After the stack is up, the health endpoint should answer and the jobs endpoint should return postings. If port 8080 is already in use, the README gives an override rather than a config file edit:

bash
HIRE_HOST_PORT=8090 make up

Migrations run automatically the first time the Postgres volume is initialized, because the migrations directory is mounted into the init directory. The README is direct about the consequence: changing a migration does not re-apply to an existing volume. You either recreate the volume or apply pending files by hand.

bash
docker compose down -v && make up
make migrate

For backend work without the full stack, the README's local development path starts only the database and then runs the server from the host:

bash
docker compose up -d db   # database only
make run

What you should see is a server on the host talking to the containerized Postgres. The compose file also fixes the database image as pgvector/pgvector:pg18 rather than stock Postgres, because migration 0092 runs `CREATE EXTENSION vector`, which the stock image does not ship. That detail matters if you plan to swap in your own managed Postgres: it needs the pgvector extension available.

Where freehire is the wrong tool

The crawler model is the limitation. freehire reads from ATS platforms and direct feeds. A company that posts roles on its own hand-rolled careers page, or that hires only through a recruiter, or that lists jobs behind a login, is not in the catalogue unless someone writes a source for it. The README frames the source list as 92 ATS platforms plus a long tail of aggregators, which is broad but not universal. If your target market is small local employers who post to a regional board, this pipeline has nothing to crawl.

The operational surface is the second constraint. Docker Compose brings up six services: the API, the web frontend, Postgres with pgvector, Meilisearch, Redis and MinIO. Redis is a required dependency for rate limiting, so you cannot drop it. Meilisearch is described as optional relative to Redis, but without it you lose the faceted search endpoint and the reindex command. That is a real amount of infrastructure for what a reader might expect to be a single binary, and the README does not document rollback for a failed migration beyond recreating the volume, which discards data.

There is also a scope question. The repository has grown well past a job index: the README lists a CV builder, an application tracker with a mail inbox, an in-process agent with five presets, and tracer links. Each of those needs configuration, and the README notes that some features require an LLM endpoint and draw on AI credits. If you only want postings in Postgres, you are carrying code and configuration for features you will not enable. The README does not document a supported way to build the ingest path alone.

freehire compared with running a scraper against one ATS API

The obvious alternative is to skip the aggregator and call a single ATS API directly. Greenhouse, Lever and Ashby all expose per-company job endpoints, and a script that loops over a list of company tokens and writes to a table is a weekend of work. That approach wins on size: no Meilisearch, no Redis, no MinIO, no frontend. It also wins on clarity, because you control the schema and there is nothing to normalize.

What it does not give you is the part freehire spends its complexity on. Cross-ATS normalization is the expensive bit: the same role on Workday and Greenhouse has different field names, different location formats and different employment-type vocabularies. freehire's stated deduplication on a stable key, plus facets derived from curated dictionaries, is the work you would otherwise redo per platform. If you crawl three companies, write the script. If you want 294,000 companies and 92 platforms behind one query shape, the normalization layer is the product, and rebuilding it is not a weekend.

The middle option is a hosted job API. Those exist, and they are cheaper in engineering time than either path here, but they reintroduce the problem freehire was built to avoid: you get someone else's copy of the posting, with their staleness and their attribution requirements. freehire's README makes the trade explicit by linking every listing to the original posting. That guarantee is only available if you own the crawl.

Licence, maintenance and the cost of staying current

freehire is MIT-licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. The repository carries a LICENSE file at the root, and the README badge repeats the MIT identifier. Nothing in the repository suggests a dual licence or a contributor agreement that would complicate a fork. This is not legal advice; if you plan to redistribute a modified version, read the LICENSE file itself.

The maintenance question is concrete. The last push to the default branch was on 2026-09-10, and the most recent tagged release is extension-v0.1.6 from 2026-08-17. The release history shows three extension releases within two days in mid-August, which suggests the browser extension is the part moving fastest right now. The Go module targets go 1.26.8 and pins a large dependency set, including chromedp, the AWS SDK, langchaingo and Meilisearch's client. Upgrade cost will track those upstreams, and the Dockerfile pins the Typst version used for CV PDF rendering to 0.15.0 with a comment noting it was matched against an ATS extraction test.

The hidden cost is the crawl itself, not the code. A self-hosted instance is only as fresh as its ingest schedule. The README describes the workers as scheduled invocations, so freshness is a function of how often you run ingest and how tolerant the target sites are of your request rate. The compose file ships a development Meilisearch key and development environment setting, which you would replace before exposing anything. The README does not document a production deployment topology beyond the scheduled worker invocation, so sizing the ingest cadence is on you.

Editorial conclusion

Adopt freehire if you want a job catalogue you control and can query over HTTP, and you are willing to run Postgres, Meilisearch, Redis and object storage. Do not adopt it if you need a hosted board with no ops work, or if your listings are not on ATS platforms the harvest layer already covers. Verify first that the crawler reaches your target companies: add one line of YAML to the sources file, run the ingest worker, then confirm the posting appears in GET /api/v1/jobs/search with the facets you expect.

Frequently asked questions

What is freehire?

freehire is an MIT-licensed, self-hostable job search engine written in Go, with a SvelteKit frontend. It crawls postings directly from company ATS platforms such as Workday, Greenhouse, Lever, Ashby and iCIMS, normalizes them into one schema, deduplicates them on a stable key, and serves faceted search over Meilisearch. The project also runs a hosted instance at freehire.me.

Is freehire actually free to use and self-host?

The repository is MIT-licensed, which the README states permits self-hosting and building on top of it, and the pipeline and data are both in the open. Running your own instance still costs you the infrastructure: Docker Compose starts the API, web frontend, Postgres with pgvector, Meilisearch, Redis and MinIO. Some features listed in the README require an LLM endpoint to be configured.

What do I need installed to run freehire locally?

The README's quick start uses Docker and the repository Makefile: `make up` builds and starts the whole stack, and `curl localhost:8080/health` checks it. If port 8080 is taken, the README shows `HIRE_HOST_PORT=8090 make up`. For backend work you can start only the database with `docker compose up -d db` and then run `make run`.

How do I add a company to freehire?

The README states that adding a company is one line of YAML in the sources configuration. After adding it you run the ingest worker, which the Dockerfile describes as a run-once binary invoked on a schedule, and then reindex so the new postings appear in search. The README does not document a dry-run mode for verifying a new source before it writes to the catalogue.

Does freehire include reposted or recruiter listings?

No. The README describes every listing as crawled directly from a company's own ATS and linking to the original posting, with no recruiter reposts and no aggregator middlemen. The catalogue is built from 225 live sources across 92 ATS platforms plus a long tail of aggregators and direct feeds.

What happens if I change a database migration in freehire?

Migrations are applied automatically only on first Postgres volume initialization, because the migrations directory is mounted into the init directory. The README states that changing a migration does not re-apply to an existing volume; you either recreate it with `docker compose down -v && make up` or apply pending files with `make migrate`.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. strelov1/freehire on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/strelov1-freehire.svg)](https://hysenlabs.com/projects/strelov1-freehire)