repo2txt: Packaging a GitHub Repository for LLM Context
Web-based tool converts GitHub repository contents into a single formatted text file
At a glance
- What is it?
- repo2txt is a browser-based TypeScript tool that converts a GitHub repository, a local directory, or a zip file into a single plain-text output for LLM prompts. All processing runs in the browser with no server-side uploads, and the tool includes real-time GPT token counting, but GitLab and Azure DevOps support are still labeled Beta.
- Who is it for?
- Engineers preparing repository context for LLM prompts who want a zero-install option will find the hosted deployment at abinthomas.in/repo2txt ready to use immediately. Those working with private repositories should supply a GitHub personal access token to access private repositories and raise the rate limit from 60 to 5,000 API requests per hour.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 57 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Context-Preparation Problem repo2txt Addresses
Feeding a repository's source code into an LLM prompt requires concatenating dozens of files into a single text block while stripping irrelevant assets, respecting .gitignore patterns, and staying within the model's context window. Doing this by hand with shell scripts is fragile; most LLM interfaces do not provide a native way to ingest an entire codebase.
repo2txt was built specifically for this workflow. The README describes it as a fast, browser-based tool for AI-assisted development: paste a GitHub URL, select which files to include, click Generate Output, and receive a single formatted text file ready to paste into an LLM prompt or download as .txt.
The intended users are engineers and researchers who regularly ask LLMs to explain, review, or modify code, and who need to prepare that code quickly without installing a command-line tool or uploading code to a third-party server. The README notes that version 2.0.0-beta.1 supports not just GitHub public and private repositories but also local directories through the browser's directory picker, zip file uploads via drag-and-drop, and beta-stage support for GitLab and Azure DevOps.
The Browser-Only Architecture and Its Consequences
The README is explicit that repo2txt is 100 percent browser-based: no server uploads occur, all processing is local, and GitHub tokens are stored in sessionStorage only, not sent to any backend. This design means the code being processed never leaves the user's device, which matters when working with proprietary or pre-publication code.
The technical stack is React 19 with TypeScript, built with Vite 5, styled with Tailwind CSS 3, and using Zustand for state management. File content loading uses progressive streaming, and tokenization runs in background Web Workers using the gpt-tokenizer library so the main thread does not block during token counting. Virtual scrolling via TanStack Virtual handles repositories with 10,000 or more files without browser performance degradation.
The privacy-first architecture has one trade-off: everything depends on the GitHub API for remote repositories, which means the application is subject to GitHub's rate limits. Without a personal access token, the limit is 60 requests per hour. With one, it rises to 5,000 per hour. The README provides step-by-step usage instructions that mention adding a token as an optional step when working with private repositories or needing higher rate limits.
The README notes the bundle is approximately 330 KB gzipped and the first load takes less than two seconds on a 3G connection. Code splitting and lazy-loaded provider modules keep the initial payload small.
Running a Local Development Copy
The hosted version at abinthomas.in/repo2txt requires no installation. For teams that need to self-host or contribute to the project, the README documents the full local setup.
First, clone the repository and install dependencies:
git clone https://github.com/abinthomasonline/repo2txt.git
cd repo2txt
npm installStart the development server:
npm run devAfter starting, the app is available at http://localhost:5173/repo2txt/. To build a production bundle:
npm run buildThe output goes to the ./dist folder. The package.json scripts also include npm run test:unit for the Vitest unit tests and npm run test:e2e for the Playwright end-to-end tests. The CI pipeline runs tsc, eslint, and the unit tests in sequence via npm run ci.
The project targets Chrome 90 or newer, Firefox 88 or newer, Safari 14 or newer, and Edge 90 or newer. No server process is needed to run the production build: the output is a static site deployable to any CDN or static host.
Sources, Filtering, and the File Tree Interface
The README documents five input sources: GitHub public and private repositories, local directories via the browser's directory picker, zip file uploads, GitLab repositories (Beta), and Azure DevOps repositories (Beta). Each source uses the same filtering and export pipeline once files are loaded.
Filtering options include selecting or deselecting by file extension, enabling gitignore pattern respect automatically, adding custom ignore patterns, cherry-picking specific directories, and using a file tree preview with virtual scrolling for navigation. These controls let users exclude test fixtures, generated files, or large binary assets before generating output.
For GitHub inputs, the README documents accepted URL formats: the repository root, a specific branch using the tree/branch-name path, a specific subfolder within a branch, and branch names that contain slashes such as feature/test/branch-name. This last case is specifically noted as working correctly.
The export step copies the concatenated text to the clipboard or downloads it as a .txt file. The README does not document the exact output format in detail, but the project description specifies it produces a single formatted text file.
Token Counting and Handling Large Repositories
Real-time token counting is one of the features the README emphasizes. The gpt-tokenizer library runs in a Web Worker and counts tokens as files are loaded, reporting per-file token counts alongside line counts. This lets engineers see whether a repository fits within a given model's context limit before generating output.
The performance design targets repositories with 10,000 or more files. Virtual scrolling prevents the file tree UI from rendering thousands of DOM nodes at once, and smart caching reduces redundant memory use across file loads. Progressive loading streams file contents as they arrive from the GitHub API rather than waiting for the full tree to load.
The README reports Lighthouse performance scores of 95 or above across all metrics and a 100 percent end-to-end test pass rate. These are self-reported figures from the README.
One practical limit is GitHub API rate capping. A repository with a deeply nested file tree may require many API requests to enumerate all paths. At 60 requests per hour without a token, a large repository could exhaust the rate limit mid-load. The README does not document what happens in this case or whether partial results are shown.
Limitations Worth Testing Before Relying on It
The most significant limitation is that GitLab and Azure DevOps support are both labeled Beta. The README lists them as supported sources but does not describe the current state of those integrations or which operations are incomplete. Teams that need GitLab or Azure DevOps support should verify their specific workflow before committing to the tool.
The browser-based model also means there is no scripting or automation path. There is no CLI, no API, and no way to call repo2txt from a build pipeline or script. The classic command-line version has been moved to a separate repository (repo2txt-classic on GitHub), suggesting that the current version is focused exclusively on the browser use case.
For very large repositories where even virtual scrolling is not enough, or where offline access to repositories is needed, a local tool would be more appropriate. The README does not document maximum file size limits for zip uploads or local directory inputs.
repomix is a commonly cited alternative that offers a similar repository-to-text conversion but runs as a Node.js command-line tool rather than in the browser. The CLI approach makes repomix scriptable and suitable for CI pipelines, whereas repo2txt's browser-only design trades that flexibility for zero local installation and stronger privacy guarantees.
Project Status and License
The package.json shows version 2.0.0-beta.1 and a last push on 2026-08-05. The project is MIT licensed according to both the README and the package.json, with no restrictions on commercial use or modification.
The repository includes an AGENT.md file described as detailed design documentation for LLM agents, alongside a CONTRIBUTING.md with architecture guidance. These files suggest the project is set up to receive contributions from both humans and AI-assisted workflows.
There are no GitHub releases listed in the repository. The beta version label on 2.0.0 means the API and output format could change before a stable release. Teams building automation around the output format should check the IMPLEMENTATION_STATUS.md file in the repository for the current state of each planned feature.
Editorial conclusion
Engineers preparing repository context for LLM prompts who want a zero-install option will find the hosted deployment at abinthomas.in/repo2txt ready to use immediately. Those working with private repositories should supply a GitHub personal access token to access private repositories and raise the rate limit from 60 to 5,000 API requests per hour. Anyone depending on GitLab or Azure DevOps should confirm those integrations suit their needs, as both are labeled Beta in the README. The MIT license permits commercial use and modification without restriction.
Frequently asked questions
Does repo2txt require a GitHub personal access token?
A token is optional. Without one, the GitHub API rate limit is 60 requests per hour. With a personal access token, the limit rises to 5,000 requests per hour, and private repositories become accessible. The README notes that tokens are stored in sessionStorage only and never sent to a server.
Can repo2txt process local files without a GitHub URL?
Yes. The README documents a Local provider that accepts a directory via the browser's directory picker or a zip file via drag-and-drop. The same filtering and export options apply to local inputs as to GitHub repositories.
How does repo2txt count tokens?
Token counting uses the gpt-tokenizer library running in a background Web Worker so it does not block the main thread. The README reports real-time per-file token and line counts displayed in the file statistics panel during loading.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/abinthomasonline-repo2txt)