OneFileLLM: One Command to Turn Repos, Papers and Videos into LLM Context
Specify a github or local repo, github pull request, arXiv or Sci-Hub paper, Youtube transcript or documentation URL on the web and scrape into a text file and clipboard for easier LLM ingestion
At a glance
- What is it?
- OneFileLLM is a Python CLI that scrapes GitHub repos, pull requests, arXiv and Sci-Hub papers, YouTube transcripts and web docs into a single structured XML file copied to your clipboard. It is convenient for prompt building and thin on guarantees: the README documents no rollback, no cache and no output size control.
- Who is it for?
- Adopt OneFileLLM if you routinely paste several sources into a chat window and want one deterministic command instead of manual copying, and if you are comfortable reading onefilellm.py when the output surprises you. Do not adopt it if you need incremental sync, a documented cache, or a size budget before you send a prompt to a paid API.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 111 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OneFileLLM solves, and for whom
The friction it targets is mundane and real. You want to ask a model about a library, a paper and a conference talk at the same time. Doing that by hand means opening the repository, copying files, downloading the PDF, finding a transcript, and pasting four things into one prompt with no consistent structure. OneFileLLM collapses that into a single command and produces one XML document, which the README says is automatically copied to your clipboard. The intended user is an engineer or researcher who already has a model interface open and wants the context assembled before they start typing. It is not a retrieval system, not an index, and not a server. It is a scraper with a formatting step. The README's own description calls it a content aggregator, and the examples confirm the scope: local paths, GitHub repositories, pull requests, issues, YouTube URLs, arXiv identifiers, DOI, PMID, and documentation sites.
How the aggregation pipeline actually works
Every input is classified first, then routed to a fetcher for that source type. A GitHub URL without a tree segment is treated as a repository; a URL containing tree/ is parsed so the request carries a ref parameter for that branch or tag, which the README states explicitly. A pull request URL goes down a different path than an issue list, and the issues path reads the state query parameter: state=all by default, state=open, or state=closed. Academic identifiers use prefix syntax, so arxiv:1706.03762, doi:10.1038/s41586-021-03819-2 and PMID:35177773 each dispatch to a different resolver. Local files are handled by extension, with a format override available through -f. The dependency list in pyproject.toml tells you what each route uses: PyPDF2 for PDFs, nbformat and nbconvert for notebooks, youtube-transcript-api and yt-dlp for video, readability-lxml and lxml for web pages, pandas with openpyxl and xlrd for spreadsheets, and tiktoken, which suggests token counting somewhere in the output path. Crawling has its own option family: depth, page cap, concurrency, delay, include and exclude patterns, robots handling, and toggles for stripping JavaScript, CSS and comments. All of it converges on one XML file. The design is a fan-in pipeline with per-source adapters, and the XML is the only interface between the tool and your prompt.
Installing OneFileLLM and running a first aggregation
There are two install routes. The pip package gives you both the CLI and a Python API without cloning anything, which is the shortest path if you only want to use it. Cloning is the route to take if you intend to read or change the source.
pip install onefilellmAfter that, the onefilellm entry point is on your PATH. The README's synopsis is onefilellm [OPTIONS] [INPUT_SOURCES...], and a first real run can mix a local directory with a GitHub URL in one invocation:
onefilellm ./docs/ https://github.com/user/project/issues/123What you should see is a single XML document written to disk and copied to the clipboard. If you are working from a clone instead, the equivalent first step is the requirements install, and the script form of the same command:
pip install -r requirements.txt
python onefilellm.py research_paper.pdf config.yaml src/GitHub requests are rate limited without a token, so the README recommends exporting one before you scrape anything large:
export GITHUB_TOKEN="your_personal_access_token"If you would rather drive it from code, the pip package exposes a run function that takes the same list of inputs:
from onefilellm import run
run(["./docs/"])For repeated source sets, the alias system stores a named command string. A simple alias wraps one URL, and a complex one can hold several separated by spaces:
python onefilellm.py --alias-add mcp "https://github.com/anthropics/mcp"
python onefilellm.py --alias-add modern-web \
"https://github.com/facebook/react https://reactjs.org/docs/ https://github.com/vercel/next.js"Aliases also accept {} placeholders, so gh-search can be created against https://github.com/search?q={} and filled in at call time. Piped input works too: the README shows cat large_dataset.json | python onefilellm.py - --format json, and the same dash convention with curl for a JSON API response.
Where OneFileLLM breaks down
The README does not document a cache, an incremental mode, or a way to resume an interrupted run. If a scrape of a large repository fails halfway, the documented behaviour is silence on the subject, and you re-run the whole thing. That matters because the tool's cost is not CPU, it is the network round trips and the API rate limits you consume each time. There is also no documented output size control. A repository with a few hundred source files can produce an XML document far larger than a model's context window, and tiktoken being a dependency does not by itself mean the tool enforces a budget. Nothing in the README describes truncation, chunking, or a token cap, so you should assume the file is as big as the sources are. The format override is another sharp edge: -f accepts text, markdown, json, html, yaml, doculing and markitdown, and choosing wrong for a given file changes what the model sees without an error. Finally, Sci-Hub is one of the advertised sources. That is a deliberate design choice by the author, and it carries legal exposure that the MIT licence on the code does nothing to address. If your organisation has a policy against that source, the tool is the wrong choice regardless of how well the rest of it works.
OneFileLLM versus repomix and similar packers
The closest comparison is a repository packer such as repomix, which also walks a codebase and emits one file for a model. The difference is scope. A packer stays inside the repository: it respects ignore rules, counts tokens, and produces a code-only artifact. OneFileLLM is broader and shallower. It will pull a GitHub repository, but it will also pull the issue list for that repository, a pull request, an arXiv paper, a DOI, a YouTube transcript and a documentation site into the same XML envelope. That breadth is the whole point, and it is also why the output is less predictable: a packer's input is a directory tree you control, while OneFileLLM's input can be a live web page whose HTML changes between runs. If your task is 'explain this codebase', a dedicated packer gives you tighter output. If your task is 'compare this library's implementation with the paper it came from and the author's talk', OneFileLLM is doing work you would otherwise do by hand. The alias system is the feature that makes that second workflow repeatable, and it has no direct equivalent in a single-purpose packer.
Maintenance, licence and the upgrade question
The last push to the repository was on 2026-06-12, roughly three months before this writing, so the project is not abandoned, but the absence of any retrieved releases means there is no versioned changelog to read before upgrading. The pyproject.toml pins version 0.1.0, which is a signal about maturity rather than quality: expect the interface to move. Several dependencies are pinned exactly (requests==2.31.0, beautifulsoup4==4.11.1, PyPDF2==2.10.0, nltk==3.7, pyperclip==1.8.2), while others float (tiktoken, pandas, PyYAML, lxml). That mix means an upgrade can break on the pinned side when a transitive dependency conflicts, and drift on the unpinned side without a changelog entry to warn you. Python 3.8 or later is required. The licence is MIT, which permits commercial use and modification with attribution and no warranty; it says nothing about the copyright status of the content you scrape, and the Sci-Hub source makes that distinction worth stating plainly. Nothing here is legal advice, but the licence on the code and the legality of a given source are two separate questions, and the README addresses only the first.
Editorial conclusion
Adopt OneFileLLM if you routinely paste several sources into a chat window and want one deterministic command instead of manual copying, and if you are comfortable reading onefilellm.py when the output surprises you. Do not adopt it if you need incremental sync, a documented cache, or a size budget before you send a prompt to a paid API. Before relying on it, verify the three things the README leaves open: what the XML looks like for your source mix, how large the file gets on a full repository, and what the tool does when a source fails mid-run. Run it once on a small directory and a single GitHub URL, inspect the XML, and only then point it at a large repository.
Frequently asked questions
How do I install OneFileLLM without cloning the repository?
The README states the project is available as a pip package, so pip install onefilellm gives you both the CLI and the Python API. Cloning is only needed if you want to run the script directly or edit the source.
Can OneFileLLM scrape a specific branch or tag of a GitHub repository?
Yes. When the GitHub URL includes a tree segment, such as https://github.com/openai/whisper/tree/main/whisper, the tool parses that portion and sends the request with a ref parameter for the specified branch or tag.
What output format does OneFileLLM produce?
It combines all inputs into a single structured XML file and copies it to your clipboard. The -f flag overrides format detection for text input, with text, markdown, json, html, yaml, doculing and markitdown as the accepted values.
Does OneFileLLM need a GitHub token?
The README recommends exporting GITHUB_TOKEN with a personal access token for GitHub API access. It describes the token as recommended rather than required, which matters because unauthenticated GitHub requests are rate limited.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jimmc414-onefilellm)