ai-doc-gen runs five analyzers and three writers, and its container image contains neither its plugin nor a test runner
AI-powered multi-agent system that automatically analyzes codebases and generates comprehensive documentation. Features GitLab integration, concurrent processing, and multiple LLM support for better code understanding and developer onboarding.
At a glance
- What is it?
- A multi-agent documentation generator that analyses a repository, writes reusable analysis documents to .ai/docs, then turns them into a README and AI assistant rule files, with a GitLab cronjob mode that opens merge requests. The packaging tells a tighter story than the feature list: nine of eleven dependencies are pinned exactly, the Python range admits one minor version, and the Docker image copies the source and nothing else.
- Who is it for?
- Use ai-doc-gen when you want documentation regenerated on a schedule rather than written once, and the cronjob mode that opens merge requests is the part that changes how a repository stays documented. Two things to know before you wire it in.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 76 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Analysis and generation are separate commands with separate exclusion axes
The pipeline is two stages that the CLI keeps apart. Five specialised analysis agents run concurrently, one each for code structure, data flow, dependencies, request flow and APIs, and their output is reusable documents written to `.ai/docs/`. Generation agents then read those documents and produce a `README.md` plus the assistant configuration files `CLAUDE.md`, `AGENTS.md` and `.cursor/rules/`.
Keeping them apart is the design decision that matters, because it means you analyse once and regenerate as often as you like:
uv run src/main.py analyze --repo-path .
uv run src/main.py generate readme --repo-path .
uv run src/main.py generate ai-rules --repo-path .The three stages have different exclusion axes, which is easy to misread. The analysis stage excludes analyses, with flags for individual ones:
uv run src/main.py analyze --repo-path . --exclude-code-structure --exclude-data-flowThe README stage instead excludes README sections, and the names do not overlap with the analysis names: `--exclude-architecture` and `--exclude-c4-model`.
uv run src/main.py generate readme --repo-path . --exclude-architecture --exclude-c4-modelSo the C4 model is a README section rather than an analysis, and skipping an analysis does not automatically remove the section that depends on it. The analysis documents can also be skipped as input, with `--use-existing-readme` supplying the current README as context instead of a fresh look at the code.
A worker pool of 0 means auto-detect, so core count becomes API spend
Concurrency is a single setting with a default that expands on its own. The environment variable is `ANALYZER_MAX_WORKERS`, the configuration key is `analyzer.max_workers`, and the value `0` means auto-detect the CPU count. The CLI exposes the same thing as `--max-workers`.
uv run src/main.py analyze --repo-path . --max-workers 2Set that aside and the shape of the cost becomes obvious. There are five analyzer agents and they run concurrently, so an unpinned install on a large machine sends five analyzer calls at once, each one iterating over files and each one capable of retrying on a rate limit. The default is tuned for throughput rather than for restraint.
Which model each agent uses is configurable per agent, along with the endpoint. That is the lever that matters more than the worker count, because it lets a cheap model handle the mechanical passes and an expensive one handle the judgement calls. The provider is not fixed: any OpenAI-compatible API works, with OpenAI, Anthropic-compatible gateways, OpenRouter and local models all named as options.
Retries are the second cost knob, and the environment sample says so directly, with a comment headed for rate-limited environments:
ANALYZER_AGENT_RETRIES=5
TOOL_FILE_READER_MAX_RETRIES=5
HTTP_RETRY_MAX_ATTEMPTS=10Three separate retry settings for three different layers, which means a run against a throttled provider will retry the HTTP layer, the tool layer and the agent layer independently.
The Claude Code plugin path needs no Python and no key, and it is a different product
The repository presents two front ends for the same ideas, and their prerequisites do not overlap.
The lighter one treats the repository as a Claude Code plugin. It ships a `.claude-plugin/` directory and a `skills/` directory, and installs from inside Claude Code with two commands:
/plugin marketplace add divar-ir/ai-doc-gen
/plugin install ai-doc-gen@divarThat path is described as needing no API keys and no Python setup at all. It adds three skills that the assistant invokes by name: `analyze-codebase` for the multi-agent analysis producing `.ai/docs/` documents, `generate-readme` for a README from analysis or from direct exploration, and `generate-ai-rules` for `CLAUDE.md`, `AGENTS.md` and Cursor rules.
The heavier one is the CLI, and its prerequisites are Python 3.13, Git, and API access to an OpenAI-compatible provider.
git clone https://github.com/divar-ir/ai-doc-gen.git
cd ai-doc-genInstallation is either uv, described as the recommended route, which installs the tool by piping a shell script and then syncing:
curl -LsSf https://astral.sh/uv/install.sh | sh
uv syncor a plain editable install with `pip install -e .`. An `ai-doc-gen` console script is also installed, exposing the same CLI as `src.main:cli_main`, so the two invocations are equivalent.
The practical consequence is that the plugin route gives you the analysis and generation steps with the host providing the model, while the CLI route makes you supply the model yourself. The generated artefacts are the same, which is why the plugin path is the one to try first if you only want to see what comes out.
Config is three layers deep, and the line budgets for the two rule files differ by 4x
Configuration is found without being named on the command line. The tool looks for `.ai/config.yaml` or `.ai/config.yml` in the repository, and the precedence order is explicit: Pydantic defaults, then the YAML file, then CLI flags. Setting up a repository means copying the shipped example into place.
# Copy and edit environment variables (LLM API keys, base URLs, etc.)
cp .env.sample .env
# Copy and edit configuration
mkdir -p .ai
cp config_example.yaml .ai/config.yamlThe YAML layer covers which analyses to skip, the worker pool cap, which README sections appear, and the AI rules behaviour including whether to skip existing files and which detail level to use, with `minimal`, `standard` and `comprehensive` as the choices.
The line budgets are where the asymmetry sits. `--max-claude-lines` defaults into the six hundreds and `--max-agents-lines` into the low hundreds, and the documented example sets them at 600 and 150 respectively. So the same generation step is allowed to write a `CLAUDE.md` four times the size of an `AGENTS.md`, which is a deliberate difference between the two consumers rather than a rounding artefact.
uv run src/main.py generate ai-rules --repo-path . \
--detail-level comprehensive \
--max-claude-lines 600 \
--max-agents-lines 150The three skip flags are also separate rather than one flag, so you can protect a hand-written `AGENTS.md` while letting the generator rewrite the other two:
uv run src/main.py generate ai-rules --repo-path . \
--skip-existing-claude-md \
--skip-existing-agents-md \
--skip-existing-cursor-rulesThe image copies src and nothing else, and its default command has no subcommand
The Dockerfile is short, and every line of it narrows what the container can do.
FROM python:3.13
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
PYTHONPATH="/app/src"
# Install uv
RUN pip install --no-cache-dir uv
WORKDIR /app
# Copy dependency files first for better caching
COPY pyproject.toml uv.lock ./
# Install dependencies using uv sync
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev
# Copy shared library and source code
COPY src ./src
WORKDIR /app
CMD ["uv", "run", "ai-doc-gen"]Dependency files are copied before the source so a code change does not invalidate the dependency layer, and the sync is frozen with dev dependencies excluded, which pins the build to `uv.lock`. A cache mount on the uv directory keeps rebuilds from re-downloading.
What follows from the copy list is the interesting part. Only `pyproject.toml`, `uv.lock` and `src/` are copied. `skills/`, `.claude-plugin/`, `.cursor/`, `config_example.yaml`, `.env.sample` and `k8s/` are all absent, so inside the container the Claude Code plugin path does not exist and the example configuration is not there to copy. Anything that reads the checked-in example has to bring its own.
The default command is also worth noting. `CMD` runs `uv run ai-doc-gen` with no subcommand, so the container starts the CLI and stops, waiting for arguments. Every real invocation has to supply one of `analyze`, `generate readme`, `generate ai-rules` or `cronjob analyze`.
For scheduled use, the repository ships a Helm chart under `k8s/helm/` aimed at CronJob deployments, which pairs with the GitLab mode below.
Nine dependencies are pinned exactly, the Python range allows one minor, and dev has no test runner
The project metadata is unusually strict, and it has consequences.
[project]
name = "ai-doc-gen"
version = "1.2.0"
requires-python = ">=3.13,<3.14"The Python range admits a single minor version, so a 3.14 interpreter is excluded by declaration. Every dependency except two is pinned with an exact equals: `pydantic==2.12.4`, `jinja2==3.1.6`, `ujson==5.11.0`, `pydantic-ai==1.22.0`, `nest-asyncio==1.6.0`, `python-gitlab==7.0.0`, `gitpython==3.1.45`, `logfire==4.15.1` and `opentelemetry-instrumentation-httpx==0.58b0`. Only `pyyaml` and `python-dotenv` are floors.
The `==` pins are what make `uv.lock` meaningful in the container, and they are also what makes an unattended rebuild fragile: when a pinned library publishes a fix, nothing moves until someone edits the manifest. Combined with the narrow Python range, this is a project that expects to be maintained rather than resolved.
The dev group is where the gap is. It contains two entries, `ipython` and `ruff`, and no test runner at all.
[dependency-groups]
dev = [
"ipython>=9.6.0",
"ruff>=0.14.0",
]No pytest, no coverage tool, nothing that would execute the code. The Ruff configuration is thorough by contrast, with a 120 character line length, a py313 target, `fix = true`, `src` as the source root, string import detection enabled, dependency direction analysis, and docstring code formatting turned on.
One packaging detail explains the import paths. The wheel declares `packages = ["src"]` with a build source mapping `src` to `src`, so the importable top-level package is literally named `src`, which is why the console script is written `src.main:cli_main` and the environment sample sets `PYTHONPATH=src`.
The cronjob mode opens merge requests, and its activity window is one flag
The fourth command is the one that changes how a repository stays documented.
uv run src/main.py cronjob analyzeRather than analysing the repository you are standing in, this mode discovers recently active projects on a GitLab instance, runs the analysis against them, and opens merge requests carrying the fresh output. The window is a single option:
uv run src/main.py cronjob analyze --max-days-since-last-commit 14Three properties follow from that design. Activity is the filter, so a repository nobody has touched is not re-analysed, which is what keeps a scheduled job over a whole GitLab group from spending tokens on abandoned code. The write-back is a merge request rather than a push, so generated documentation arrives as a reviewable change with a diff, and the same exclusion and worker controls apply to every project in the batch. And because the tool holds a GitLab client rather than shelling out to a forge, the dependency list carries `python-gitlab` and `gitpython`.
Running it on a schedule is the deployment story the repository commits to, with a Helm chart under `k8s/helm/` described as being for containerized and scheduled Kubernetes CronJob deployments. The container image, the chart and the cronjob mode are three parts of one path: build, schedule, and hand the result to a human.
One thing the mode does not describe is what happens when analysis output conflicts with an existing generated file that someone has edited. The skip flags exist for the AI rules files and are per-target, and the README stage has its own context flag, but no policy for the merge is stated, so the conflict resolution is whatever GitLab does with the branch you push.
Telemetry ships switched off, and the repository commits the files the tool writes
Observability is configured in the environment sample rather than in the YAML, and it arrives in two layers.
ENABLE_LANGFUSE=false
OTEL_SDK_DISABLED=falseTracing is OpenTelemetry, routed through `logfire` as the always-present dependency, with Langfuse as an optional integration for LLM observability and analytics. When you turn Langfuse on it wants three values, a public key, a secret key and a host, and there is a separate exporter endpoint for the OpenTelemetry data:
LANGFUSE_PUBLIC_KEY=YOUR_PUBLIC_KEY_HERE
LANGFUSE_SECRET_KEY=YOUR_SECRET_KEY_HERE
LANGFUSE_HOST=https://YOUR_LANGFUSE_HOST_HERE
OTEL_EXPORTER_OTLP_ENDPOINT=YOUR_OTEL_ENDPOINT_HERENote the shape of that: the Langfuse layer is off by default while the OpenTelemetry SDK switch is on, so a default run emits trace data through logfire and sends nothing to Langfuse unless you opt in and supply an endpoint. The environment sample also sets `ENVIRONMENT=development` with the three options spelled out as development, staging and production, which is a reminder that the sample is the place most runtime decisions actually get made.
The repository itself is the clearest demonstration of the output. At the top level sit `CLAUDE.md`, `AGENTS.md`, `.cursor/` and `.ai/`, which are exactly the files the generation stage is built to produce, alongside `.claude-plugin/`, `skills/`, `config_example.yaml`, `.env.sample`, the `Dockerfile`, `uv.lock` and the `k8s/` chart. The tool's own assistant rule files are committed, which is the practical test of whether the generated output is worth keeping in version control.
Two last facts about the project rather than the code. It carries version 1.2.0 in its metadata and has no GitHub releases at all, so that number has never been published as a version you can install by tag. And the reasoning behind it is written up outside the repository, in an English post on Medium and a Persian post on Virgool, both linked from the project page. The last push to the default branch is dated 2026-07-21, and the licence is MIT.
Editorial conclusion
Use ai-doc-gen when you want documentation regenerated on a schedule rather than written once, and the cronjob mode that opens merge requests is the part that changes how a repository stays documented. Two things to know before you wire it in. The concurrency default is auto-detected from your CPU count, and each of the five analyzers is an LLM call, so a machine with many cores multiplies your spend unless you set a worker cap. And the dependency list is pinned with exact versions under a Python range that allows only one minor release, so plan for maintenance rather than assuming a rebuild will resolve. For the Claude Code plugin path you need neither Python nor a key, which makes it the cheapest way to try the analysis half.
Frequently asked questions
What does ai-doc-gen actually generate for a repository?
Five concurrent analysis agents cover code structure, data flow, dependencies, request flow and APIs, and write reusable documents to `.ai/docs/`. Generation agents then turn those into a `README.md` plus `CLAUDE.md`, `AGENTS.md` and `.cursor/rules/` files, with generated output placed in the repository root.
How do I run ai-doc-gen without setting up Python?
Install it as a Claude Code plugin, which the project says needs no API keys and no Python setup. The two commands are `/plugin marketplace add divar-ir/ai-doc-gen` and `/plugin install ai-doc-gen@divar`, and it adds the `analyze-codebase`, `generate-readme` and `generate-ai-rules` skills.
How many analyzer agents run at once in ai-doc-gen?
Five, one each for code structure, data flow, dependencies, request flow and APIs. Concurrency is capped by `ANALYZER_MAX_WORKERS` or `analyzer.max_workers`, where `0` auto-detects the CPU count, and can be set per run with `--max-workers`.
Which LLM providers can ai-doc-gen use?
Any OpenAI-compatible API, with per-agent model and endpoint settings. OpenAI, Anthropic-compatible gateways, OpenRouter and local models are all named as options. Retries sit at three layers, with `ANALYZER_AGENT_RETRIES`, `TOOL_FILE_READER_MAX_RETRIES` and `HTTP_RETRY_MAX_ATTEMPTS` set separately for rate-limited environments.
What does the GitLab cronjob mode in ai-doc-gen do?
It discovers recently active GitLab projects, runs the analysis on them, and opens merge requests with the fresh output. `--max-days-since-last-commit` restricts it to projects with commits inside a given window. A Helm chart under `k8s/helm/` is provided for running it as a scheduled job.
How does ai-doc-gen avoid overwriting hand-written AI rule files?
Three separate skip flags exist, `--skip-existing-claude-md`, `--skip-existing-agents-md` and `--skip-existing-cursor-rules`, so you can protect one target and let the generator rewrite the others. Generation detail is set with `--detail-level minimal`, `standard` or `comprehensive`, and the size budgets are separate per file, `--max-claude-lines` and `--max-agents-lines`.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/divar-ir-ai-doc-gen)