DeepGit: a six-stage GitHub search agent that spends LLM calls only when free signals disagree
Deep research agent to help you find the best GitHub repositories 🕵️!
At a glance
- What is it?
- A MIT licensed Python research engine that reads intent, sweeps topics, prefilters on free signals, and only opens source code when a confidence gate says the answer is contested. Sound shape, but the packaging, benchmarks and env defaults are all looser than the README implies.
- Who is it for?
- DeepGit is a reasonable pick if your GitHub search is a recurring task and you want the ranking judged on requirements rather than stars, and if you can absorb a Python 3.11 install with a token and one LLM provider key.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 36 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two LLM calls, and code reading only for contested picks
The pipeline has six stages and only two of them can spend a token. First, one LLM call turns your request into a plan: search angles, GitHub topics, hard must-haves, soft nice-to-haves, and anti-patterns. Second, gathering runs at zero cost through batched GitHub GraphQL search across keyword queries plus star sorted topic sweeps, which is how repos like apache/kafka surface even when the description omits the obvious words. Third, semantic recall upserts everything seen into a local vector index and pulls in near neighbours from the accumulated corpus. Fourth, a prefilter narrows the field using activity, health, license, tests and CI, and relevance. Fifth, one batched, evidence grounded ranking runs over compact cards made of README text plus those signals, and it enforces the language and hardware constraints you named. Sixth, a zero-cost confidence gate decides whether the top picks are contested; only then does DeepGit read the real source, tests and manifests, re-judge from that code, and hand the final order to a reflection critic.
So a C library never wins a C++ query, and a GPU-only tool never tops a plain fast X ask. Two calls is the design target, not a measured average; the measured number lives in the benchmark table.
The quickstart needs a token, and the semantic extra is the heavy part
Getting started takes one clone, one virtualenv and one editable install:
git clone https://github.com/zamalali/DeepGit.git
cd DeepGit
python -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activate
pip install -e ".[semantic]" # the semantic extra is optional but recommended
export GITHUB_API_KEY=ghp_xxx # a PAT with public-repo read accessRequirements are Python 3.11 or newer and a GitHub token. The token is what lets the gathering stage run GraphQL search instead of scraping, so a working install is tied to a credential with public repo read access.
The bracketed extra decides what you get. It pulls in lancedb, fastembed and truststore for persistent CPU-only embeddings. Without it the pipeline degrades gracefully to keyword and topic search, so the semantic recall described in the feature table is an opt-in cost rather than a default. That is a sensible split for a first run, but it also means the headline behavior of the project only exists for people who read the parenthetical comment.
The search itself runs from the terminal as `deepgit search "a fast, embeddable key-value store in Go"`, or from Python as a library.
Two dependency lists keep the same thirteen pins
pyproject.toml and requirements.txt list the same thirteen runtime dependencies at the same floors: dspy>=2.5, deepagents>=0.0.2, langgraph>=0.2.60, langchain-core>=0.3, langchain-groq>=0.2, langchain-google-vertexai>=2.0, httpx>=0.27, anyio>=4.0, pydantic>=2.0, pydantic-settings>=2.0, python-dotenv>=1.0, numpy>=1.25, and mcp>=1.2,<2. requirements.txt opens by telling you to prefer the editable install, then repeats the list and comments the three optional semantic packages back out as hints. Two lists with no generation step between them means a version bump has to be made twice, and nothing in the repo checks that they agree.
The pins themselves carry information. mcp is held below version 2 while everything else uses an open upper bound, so a breaking MCP release is anticipated and deliberately not absorbed. deepagents is still at 0.0.2, an early tag for a package the agent layer depends on. And the two provider integration packages that ship by default are langchain-groq and langchain-google-vertexai, while .env.example advertises four provider values: groq, openai, anthropic, vertex_ai. Choosing OpenAI or Anthropic means relying on what dspy itself ships rather than on a pinned integration package here, so the non-default providers carry a different upgrade risk than the two that have code in the dependency list.
The Docker image starts an MCP server on 0.0.0.0:8080
The image is python:3.12-slim with the semantic extra installed unconditionally, so the container is heavier than the minimal local install. It sets MCP_TRANSPORT to streamable-http, HOST to 0.0.0.0, PORT to 8080, exposes 8080, declares a volume at /data, points DEEPGIT_INDEX_DIR at /data/lancedb, and enters through deepgit-mcp rather than the CLI:
docker run --rm -p 8080:8080 -e GITHUB_API_KEY=ghp_xxx deepgitThat published port is the operational detail to sit with. The server binds every interface and the GitHub token travels in as an environment variable, and the documented command has no authentication flag, no token gate and no localhost-only variant. Anything that can reach port 8080 can spend your token's rate limit through the agent. The image does offer the safer shape in a comment: for a local stdio client, run `docker run --rm -i -e GITHUB_API_KEY=... deepgit deepgit-mcp`. Publishing that port to the internet is a decision to make deliberately rather than inherit from the README example.
Two more Docker notes. Persistence is opt-in through the volume mount, so without `-v deepgit-data:/data` the accumulated vector index dies with the container and semantic recall restarts from nothing. And the README section that explains the Model Context Protocol setup stops one clause into the sentence linking to the specification, so the client configuration promised by the feature table is not in the file.
pytest is pointed at bench/, so the benchmark is also the test suite
The pytest configuration reads `testpaths = ["bench"]` with `asyncio_mode = "auto"`, and no tests directory appears at the top level of the repo. Whatever tests exist are collected from the benchmark directory, which means the suite that verifies behaviour and the suite that produces the headline numbers are the same code path. The dev extra carries pytest, pytest-asyncio, ruff and mypy, so the tooling is there; the placement is the choice to question.
Three benchmark entry points are documented, each writing a full per-case report to _bench_report.md:
python -m bench.run_bench --heldout # held-out set (unseen queries)
python -m bench.run_bench # core set
python -m bench.run_bench --edge # robustness / edge casesThe numbers the project publishes come from the held-out run: 12 unseen queries with the adaptive controller on, Hit@1 at 75%, Hit@3 at 92%, MRR at 0.833, wrong pick at position one at 0%, and an average of 2.0 LLM calls per query. Twelve queries is a small sample, and the 0% wrong-pick figure is a claim about a dozen judgments rather than a measured error rate. Lint is configured separately: ruff targets py311 at line length 100 with E, F, I, N, UP and B selected, and mypy runs with strict off but warn_return_any on.
Version 4.0.0 with no tagged release and a Beta classifier
pyproject.toml says version 4.0.0, requires-python >=3.11, and classifies the project as Development Status 4 - Beta with Python 3.11 and 3.12 listed. The GitHub repository has no releases at all, so there is no tag to install from and no changelog entry to diff against. A package can reach a 4.x number without anyone having cut a 4.x artifact; if you pin, you are pinning a commit.
The version number is also the one place the rebuild is not visible. The README announces a ground-up rebuild with semantic recall and an MCP server as the new parts, but it never attaches that claim to a version, so there is no way to tell from the README which release predates the rewrite. The last commit on the default branch is dated 2026-08-30, roughly a month before this writing, which is recent enough that the absence of releases looks like a process choice rather than abandonment.
Platform coverage is uneven in the same file. Local installs accept 3.11 and up, and classifiers name 3.11 and 3.12 only, while the container is built on python:3.12-slim. There is no 3.13 classifier, so a 3.13 environment is unclassified even though the floor would allow it.
Retrieval defaults live in commented-out lines
The knobs that shape search results are all commented out in .env.example:
# Optional: Override retrieval settings
# MAX_RESULTS=100
# RETRIEVAL_ALPHA=0.7
# RERANK_TOP_N=50The names are useful and the shapes are guessable, but the numbers shown are placeholders inside comments. Anyone using DeepGit as a library rather than through the CLI has three tunable parameters and no stated starting values anywhere in the documentation. Since the ranking is described as fit-based rather than star-based, these three settings are the difference between a fast recall-heavy search and a precise one, and they are exactly the settings you would want written down.
The rest of the example file has a similar shape. GITHUB_API_KEY and GROQ_API_KEY are both marked required, the Groq key with a note that it serves DSPy LLM calls, which means a Vertex AI setup still faces a required-looking Groq line it will never use. LANGSMITH_API_KEY sits empty next to LANGCHAIN_TRACING_V2=false, so tracing is off unless you ask for it. DEEPGIT_INDEX_DIR is the one variable the Dockerfile sets and .env.example never mentions, which is how the Docker volume mount keeps its meaning.
The wheel ships only src/deepgit, so bench and agent.yaml stay in the checkout
The build backend is hatchling and the wheel target packages exactly one directory, src/deepgit, with two console scripts: deepgit pointing at deepgit.__main__:main and deepgit-mcp at deepgit.mcp_server:main. That is the right shape for shipping a CLI and a server. It also means everything outside that directory is absent from an installed copy.
The repo root holds bench/, assets/, SOUL.md, agent.yaml, the Dockerfile, requirements.txt and the README. SOUL.md and agent.yaml are referenced by no packaging metadata and appear nowhere in pyproject.toml, so whatever they describe ships only for people working from a source clone. bench/ is in the same position, with a practical consequence: `python -m bench.run_bench --heldout` is a module invocation against the source tree, so reproducing the published numbers needs a checkout rather than the installed package. The CLI and the MCP server behave the opposite way and work from a plain install.
The project also publishes a homepage at deepgit.vercel.app. Nothing in the README describes what that deployment runs, whether it is a demo of the same pipeline or something separate, so treat it as an unverified entry point rather than documentation.
Editorial conclusion
DeepGit is a reasonable pick if your GitHub search is a recurring task and you want the ranking judged on requirements rather than stars, and if you can absorb a Python 3.11 install with a token and one LLM provider key. It is not a fit for offline use, for air-gapped environments, or for anyone who needs a stable release stream: version 4.0.0 sits in pyproject.toml with no tagged GitHub release, the classifier still says Beta, and the last commit on the default branch is dated 2026-08-30. Before you build on it, check four things. Read the benchmark numbers as the project's own claim on 12 held-out queries, not as a general accuracy figure. Confirm the retrieval defaults, because MAX_RESULTS, RETRIEVAL_ALPHA and RERANK_TOP_N appear only as commented lines in .env.example. Decide which LLM provider you will set, since Groq is what .env.example ships a key for. And if you deploy the container, remember it binds port 8080 on all interfaces with your GitHub token inside.
Frequently asked questions
What does DeepGit need before it can search?
Python 3.11 or newer, a GitHub personal access token with public repo read access exported as GITHUB_API_KEY, and a key for one LLM provider. The keyword and topic stages run without the semantic extra; lancedb, fastembed and truststore arrive with the optional [semantic] install.
How many LLM calls does one DeepGit search make?
The project reports an average of 2.0 calls per query on its held-out benchmark of 12 unseen queries. Gathering, semantic recall and prefiltering run on free signals, and a zero-cost confidence gate decides whether the real source, tests and manifests get read and re-judged.
Can DeepGit run without the semantic extra?
Yes. The pipeline degrades gracefully to keyword and topic search when lancedb and fastembed are missing, and pyproject.toml says so in a comment next to the extra. The Docker image is the exception, since it installs the extra by default.
What do DeepGit's own benchmark numbers say?
On the held-out set of 12 unseen queries: Hit@1 75%, Hit@3 92%, MRR 0.833, wrong pick at position one 0%, and 2.0 average LLM calls per query. Three runs exist, core, held-out and edge, and each writes a per-case report to _bench_report.md.
Which LLM providers can DeepGit use?
Set LLM_PROVIDER to groq, openai, anthropic or vertex_ai. Groq is the default and the one .env.example ships a key for; Vertex AI authenticates through Application Default Credentials with VERTEX_PROJECT and GOOGLE_APPLICATION_CREDENTIALS.
How do you run DeepGit as an MCP server?
Run deepgit mcp locally, or start the Docker image, whose entrypoint is deepgit-mcp with MCP_TRANSPORT set to streamable-http on port 8080. For a local stdio client the image comment suggests docker run --rm -i -e GITHUB_API_KEY=... deepgit deepgit-mcp instead of publishing the port.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zamalali-deepgit)