TaxHacker: self-hosted AI accounting for freelancers who want their receipts parsed locally
Self-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories
At a glance
- What is it?
- TaxHacker is a self-hosted TypeScript app that sends receipt and invoice images to an LLM and stores the extracted fields in a Postgres-backed table. It is aimed at freelancers and small businesses, and its main trade-off is that you own the prompts, the data and the OCR quality.
- Who is it for?
- TaxHacker fits freelancers, indie hackers and small businesses that already run Docker and want their receipt data in a Postgres database they control. It is a bad fit if you need double-entry bookkeeping, bank reconciliation or a supported product with an SLA.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem TaxHacker targets: receipt data trapped in images
The README frames the audience narrowly: freelancers, indie-hackers and small businesses. That is a group that files taxes but does not have an accountant on staff doing data entry. The task TaxHacker automates is the boring middle step: taking a photo of a receipt or a PDF invoice and turning it into a row with a date, an amount, a vendor, a currency and line items. The README states the app recognizes product names, amounts, items, dates, merchants and taxes, and stores them in what it calls a structured Excel-like database. It also handles documents in any language and any currency, which matters if you travel or invoice clients abroad.
The second problem is currency. TaxHacker detects the currency in a document and converts it to your base currency using historical exchange rates from the transaction date. The README claims support for 170+ world currencies and 14 cryptocurrencies including BTC, ETH, LTC and DOT. That is a real accounting concern: using today's rate for a purchase made three months ago produces a number your tax authority may not accept. Whether the rate source is reliable is not documented in the README, so treat that as something to verify against your own records.
The third problem is privacy. Sending every receipt to a cloud LLM means a third party sees your spending. TaxHacker's answer is self-hosting plus the option to point it at an OpenAI-compatible local endpoint such as Ollama, LM Studio, vLLM or LocalAI. The README is honest about the cost of that choice: "make sure that your local model is good in OCR tasks, results are not guaranteed."
How the extraction pipeline works, from upload to Postgres row
The repository layout shows a Next.js application (app/, components/, lib/, hooks/) with Prisma as the data layer (prisma/, prisma.config.ts) and LangChain as the LLM abstraction layer. The dependency list includes @langchain/openai, @langchain/google-genai and @langchain/mistralai, which matches the three cloud providers named in the README, plus langchain itself for the local OpenAI-compatible path.
The data flow implied by that stack: a document lands in an upload directory (UPLOAD_PATH, set to /app/data/uploads in the Docker Compose file), a record is created in Postgres, and when you trigger analysis the app renders the document into a prompt and sends it through LangChain to the configured provider. The model returns structured fields, which are written back to the transaction row. Custom fields work the same way: you define a field and a prompt, and that prompt is appended to the extraction request for matching documents.
The README describes documents as starting in an "unsorted" state until you process them manually or with AI assistance. That staging step is worth noting because it means ingestion and extraction are separate actions. You can bulk-upload a folder of receipts, then decide when to spend tokens on them. The Settings area exposes the system prompt template itself, which the README says you can modify, along with field-level and project-level prompts. That is unusual for this category: most receipt scanners hide the prompt entirely.
Installing TaxHacker with Docker Compose on port 7331
The repository ships three Compose files: docker-compose.yml (uses the prebuilt image ghcr.io/vas3k/taxhacker:latest), docker-compose.production.yml and docker-compose.build.yml. The plain docker-compose.yml is the shortest path. It defines three services: app, cron and postgres. The app publishes port 7331, sets SELF_HOSTED_MODE=true, mounts ./data to /app/data, and connects to postgresql://postgres:postgres@postgres:5432/taxhacker.
The .env.example file documents the variables you may want to set before starting. Note that BETTER_AUTH_SECRET is described there as recommended everywhere, and the Compose file comments say that in self-hosted Docker mode TaxHacker can auto-generate and persist one under ./data if it is omitted.
cp .env.example .env
# edit .env: set POSTGRES_PASSWORD and, if you want, BETTER_AUTH_SECRETAfter that, bring the stack up. The image is pulled from GitHub Container Registry, so no local build is required.
docker compose up -dOnce the containers report healthy, open http://localhost:7331 in a browser. The README's landing page and the Compose port mapping agree on 7331, so if the page does not load, check that nothing else on the host holds that port. Signup is controlled by DISABLE_SIGNUP, which .env.example sets to true; the README does not spell out the first-user flow, so if you cannot create an account, that variable is the first thing to inspect.
For a local development checkout instead of Docker, package.json requires Node >= 26 and provides a dev script that runs Next.js on port 7331 with Turbopack:
npm ci
npm run devYour first real use: upload one receipt photo, let it sit in unsorted, then run AI analysis on it and check the extracted date, amount and vendor against the paper. If the model mangles your language or handwriting, that is the signal to change models or edit the system prompt in settings.
Where TaxHacker breaks down
The README carries its own warning: "This project is still in early development. Use at your own risk!" That is not boilerplate. Version 0.8.5 shipped on 2026-07-20, and the last push to main was on 2026-09-07, so the codebase is moving, which also means schema and prompt behaviour can change between releases. The start script runs prisma migrate deploy before next start, so an image update can apply database migrations automatically. Back up ./pgdata and ./data before pulling a new tag.
The bigger limitation is what TaxHacker is not. Nothing in the README mentions double-entry bookkeeping, bank feeds, reconciliation, VAT return generation or e-invoicing formats. It extracts and stores transactions; it does not keep books in the accounting sense. If your accountant expects a trial balance, TaxHacker gives you a CSV export and attached documents, not a ledger.
OCR quality is entirely the model's problem. The README says results are not guaranteed with local models and warns that you alone are responsible for the quality and privacy of your data. A small local model that is fine with printed English receipts may fail on a crumpled handwritten receipt in another script. There is no documented fallback or confidence score described in the README, so a wrong extraction looks the same as a right one until you check it.
Finally, the README's own banner notes the author is looking for a job. That is a maintenance-risk signal worth weighing, though it is not the same as an archived repository.
TaxHacker versus Akaunting and other accounting tools
Akaunting is the natural comparison and appears in the related searches around this project. The difference is architectural, not cosmetic. Akaunting is a general-purpose double-entry accounting application: chart of accounts, journals, invoices you issue, bank reconciliation. Its input is structured data you type or import. TaxHacker starts at the other end. Its input is unstructured paper and PDFs, and its output is a flat transaction table with custom columns. TaxHacker does not replace Akaunting's ledger; it can feed it, because TaxHacker exports filtered transactions to CSV with attached documents.
Against cloud receipt scanners, the split is where the model runs and who holds the data. TaxHacker lets you point at Ollama, LM Studio, vLLM or LocalAI through an OpenAI-compatible endpoint, which means the receipt image never leaves your network. The cost is that you now own model selection, prompt tuning and extraction accuracy. A hosted scanner absorbs that work and charges for it.
Against writing your own LangChain script, TaxHacker's contribution is the surrounding application: Postgres schema, upload handling, unsorted staging, bulk operations, full-text search over recognized document content, multi-project grouping, categories, custom fields, CSV export and a cron container for scheduled jobs. That is a lot of plumbing you would otherwise rebuild.
Licence, upgrade cost and what self-hosting actually asks of you
TaxHacker is MIT licensed. In practical terms that means you can run it, modify it and use it commercially, and the repository includes the LICENSE file at the root. MIT gives no warranty, and the README's early-development warning reinforces that you are on your own for correctness of extracted numbers. This is not legal advice; if you plan to rely on the output for filings, read the licence and the README yourself.
Upgrade cost has three parts. First, the image: docker compose pull followed by docker compose up -d, with prisma migrate deploy running on start, so database migrations apply on boot. Second, the data volume: ./data holds uploads and, per the Compose comments, the persisted auth secret when BETTER_AUTH_SECRET is unset, while ./pgdata holds Postgres. Copy both before upgrading. Third, the LLM: model names are configurable through OPENAI_MODEL_NAME, GOOGLE_MODEL_NAME and MISTRAL_MODEL_NAME, so a provider deprecating a model is a config edit rather than a code change, but the extraction prompt may need retuning against the new model.
There is also an operational surface beyond the app container. The cron container mounts ./etc/crontab read-only and runs docker-cron-entrypoint.sh, and the Dockerfile installs cron and ghostscript in the runtime image, which suggests PDF processing happens in-process. The app also includes IMAP dependencies (imap-simple, mailparser) and an email:sync script, so an email-ingestion path exists in the codebase even though the README excerpt does not describe it in detail.
Editorial conclusion
TaxHacker fits freelancers, indie hackers and small businesses that already run Docker and want their receipt data in a Postgres database they control. It is a bad fit if you need double-entry bookkeeping, bank reconciliation or a supported product with an SLA. Before committing, verify that your chosen LLM handles your document language and handwriting, and check the docs for the backup and restore path for the ./data volume.
Frequently asked questions
Is TaxHacker an AI that will do my taxes?
No. TaxHacker extracts data from receipts, invoices and PDFs and stores it in a structured database with categories, projects and custom fields. The README describes it as simplifying reporting and making tax filing a bit easier, not as filing returns.
Is TaxHacker a ChatGPT for taxes?
It uses LLMs, but it is an application rather than a chat interface. You upload documents, the configured model extracts fields such as dates, amounts, vendors and line items, and the results are saved to a Postgres-backed table.
Can I use AI for my tax return with TaxHacker?
TaxHacker can prepare the underlying transaction data and export filtered transactions to CSV with attached documents, which the README says can be used for reports for an accountant or tax advisor. The README does not describe generating or submitting a tax return.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vas3k-taxhacker)