# AI Bank Statement Document Automation: YOLO, OCR and LLM Agents for PDF Statements

> A Jupyter-first repository that turns bank statement PDFs into structured, redacted financial data using YOLO layout detection, OCR, LiteLLM and three parallel agent harnesses. The design choices are opinionated, and the documentation leaves several operational gaps.

**johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction** — AI Bank Statement Document Automation By LLM model and Personal Finanical Analysis

- Repository: https://github.com/johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction
- Stars: 627 · Forks: 123
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/johnsonhk88-ai-bank-statement-document-automation-by-llm-and-personal-finanical-

## What the repository actually automates

Bank statement PDFs are hostile input. Layouts vary by issuer, tables span pages, and the numbers you want sit inside a visual grid that text extraction flattens into nonsense. This repository attacks that problem with a three-stage pipeline: YOLO for layout detection, OCR for text recovery, and an LLM for table extraction and structuring. The output feeds a financial analysis layer that categorizes income and expenses, looks for trends, and answers natural language questions over the extracted data.

The intended user is a developer or analyst who wants to run the whole chain locally. The README lists LM Studio and Ollama as first-class providers through LiteLLM, and the topics include ollama, gemma and rag. There is also a full-stack path: FastAPI, PostgreSQL 16, Celery with Redis, and a React 18 SPA served on port 80. That is a lot of surface area for one repository, and the README presents the Jupyter notebook as the primary entry point rather than the API. Treat the notebook as the reference implementation and the API as a work in progress.

## How the extraction and agent layers fit together

The document path is explicit in the README's feature list: YOLO layout detection, then OCR, then LLM-based table extraction. PyMuPDF and pymupdf4llm handle PDF parsing, pytesseract and opencv-python handle OCR, and ultralytics supplies YOLO. A sample file lives at data/bank-statement-document/Dummy-Bank-Statement.pdf.

Above that sit three agent harnesses that run the same workflow in parallel. CrewAI is the baseline and runs from the Jupyter notebook at backend/app/core/ai_agent_skills_dev.ipynb. Deep Agents uses the deepagents SDK with a run_e2e.py entry point under agents/deep-agents/. Hermes runs in a Docker sandbox through the Hermes CLI. Each harness keeps its own copies of the skills, and the README states there is no auto-sync between them. That means a fix to the PII redaction skill in one harness does not reach the other two. It is a deliberate isolation choice, and it is also a maintenance tax.

The RAG side is where the security posture is clearest. PII redaction happens before embedding into the vector store, which the README lists under Qdrant and Chroma. Redacting first is the right order. If you embed raw statement text, the vectors themselves become a PII store that is hard to audit and harder to delete.

## Installing the pinned stack and opening the notebook

The README does not give a full install sequence. It gives a pinned requirements.txt and a quickstart that begins with a cd into agents/deep and then stops. What can be confirmed is the dependency set and the notebook entry point.

The root requirements.txt pins the CrewAI stack. Install it into a virtual environment before opening the notebook:

```bash
pip install -r requirements.txt
```

The README notes that the file is pinned for reproducibility and suggests pipdeptree when conflicts appear:

```bash
pip install pipdeptree && pipdeptree
```

The API has its own pinned set under api/requirements.txt. The README describes the local model settings you must respect before running anything. Prefer 9B parameters or more, and 16K tokens of context or more, with 32K recommended. Below that, the README says agent instructions get dropped.

The README also points at a deterministic Deep Agents path that needs no LLM at all, extracting the PDF, redacting PII, storing vectors and answering from RAG. The command it gives is a cd into the deep-agents directory, and the text cuts off there, so the run command itself is not documented.

```bash
cd agents/deep
```

What you should see after installing is a working LiteLLM import path, since litellm==1.79.2 is pinned in requirements.txt. The README notes that reasoning models may return content in the reasoning_content field, so check that field before concluding a model returned nothing.

## Where the design breaks down

The most concrete limitation is stated by the project itself: local models often fail strict JSON output. The README recommends Markdown reports instead, with a second LLM call or the instructor package to post-process. That is a real constraint, not a footnote. If your downstream consumer expects a schema, you are adding a validation and repair step that the repository does not ship as a finished component.

The second limitation is the skill duplication across harnesses. Three copies of the same domain skills, with no sync mechanism, means the harnesses will drift. The README labels the multi-harness section experimental, which is honest, but it also means you should pick one harness and ignore the other two rather than treating them as interchangeable.

The third is scope. A repository that contains a Jupyter notebook, a FastAPI service, a Celery worker, Kubernetes manifests, a React SPA and three agent frameworks is not a library. There are no releases, no version tags and no published package. You are vendoring a snapshot. The README also does not document rollback, migration paths or how to upgrade between snapshots, which matters because Alembic migrations are in the tree and the API owns a PostgreSQL schema.

Finally, the README's GPU claim, that NVIDIA acceleration gives under 2x build time versus CPU, is a build-time comparison, not an inference benchmark. It says nothing about throughput on real statements.

## Choosing between this and a document AI service

The obvious alternative is a managed document AI service, or a general extraction library such as Docling, which this repository already depends on through langchain-docling. The difference in approach is where the intelligence sits. A managed service owns the layout model and the extraction model, and you send it documents. This repository owns nothing but the orchestration: it runs YOLO locally, calls OCR locally, and routes LLM calls through LiteLLM to whatever endpoint you configure, including LM Studio or Ollama on your own machine.

That matters for statements specifically. Bank statements contain account numbers, names, addresses and full transaction histories. With a local model, the document never leaves your network. With a managed service, it does, and you are relying on the vendor's retention policy. The trade-off is quality and effort. A vendor has tuned layout models for hundreds of statement formats. This repository gives you a yolo-base-layout-analysis directory and a dummy PDF, which means format coverage is your problem to solve.

If your statements come from one or two issuers, the local path is tractable. If they come from dozens, you will spend more time on layout detection than on the financial analysis the project is named for.

## Licence and the cost of keeping this running

The repository is Apache-2.0, which permits commercial use, modification and redistribution, and includes a patent grant. It also requires that you preserve the licence and notice files and state significant changes. This is not legal advice; check how the licence interacts with the model weights you pair it with, since Gemma and Llama carry their own terms that are separate from the code licence.

The maintenance picture is mixed. The last push was on 2026-08-02, which is recent, and the repository is not archived. But there are no releases, so there is no upgrade path other than pulling main. The dependency pins are aggressive: crewai==1.14.7, litellm==1.79.2, langchain==0.3.25, pymupdf==1.26.6, ultralytics==8.4.67. Pinning helps reproducibility and hurts currency. LangChain in particular moves fast, and a pinned 0.3.x line will accumulate incompatibilities with newer integration packages.

The requirements file itself notes the recommended future path: pyproject.toml with uv or pip-tools. That migration has not happened. Budget for dependency work before you budget for features.

## Conclusion

Adopt this repository if you are comfortable reading a Jupyter notebook and wiring your own LLM endpoint through LiteLLM, and you want a reference implementation of the YOLO plus OCR plus agent pipeline rather than a finished product. Do not adopt it if you need a supported, versioned release, a documented install path, or a hosted service; there are no releases and the README stops mid-sentence at the quickstart. Before committing, verify that your statement PDFs survive the YOLO layout step, that your chosen model meets the README's 9B parameter and 16K context floor, and that the PII redaction step runs before any embedding call in the path you intend to use.

## FAQ

### Can I use AI to analyze my bank statements with this project?

Yes, that is the stated purpose. The repository extracts transactions from statement PDFs with YOLO, OCR and an LLM, then runs categorization, trend analysis and natural language querying over the result. The README recommends a local model of 9B parameters or more with at least 16K tokens of context.

### Which AI tool is best for banking and finance in this repository?

The repository does not rank tools. It routes calls through LiteLLM, so the provider is a configuration choice, and the README lists LM Studio, Ollama, OpenAI and DeepSeek among the options. It does warn that local models often fail strict JSON output and suggests Markdown reports with post-processing instead.

### Does the AI Bank Statement Document Automation project handle PII in statements?

The README states that PII redaction happens before embedding into the vector database, which it lists as Qdrant or Chroma. There is a dedicated pii-handling skill in the shared skills directory. The README does not document what the redaction covers or how it is validated.

### What model size does the AI Bank Statement Document Automation project require?

The README recommends 9B parameters or more, naming Qwen2.5-14B, Qwen3-27B, Gemma-2-9B and Llama-3.1-8B+ as examples. It also recommends 16K tokens of context, with 32K preferred, and warns that low context drops agent instructions.

### Is the AI Bank Statement Document Automation project a released package?

No. There are no releases or version tags in the repository, and the README's quickstart is truncated. Adoption means pulling the main branch and installing from the pinned requirements.txt files at the root and under api/.

## Sources

- [Issues](https://github.com/johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction/issues)
- [johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction on GitHub](https://github.com/johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction)
- [License: Apache-2.0](https://github.com/johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction/blob/main/LICENSE)
- [README](https://github.com/johnsonhk88/AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/johnsonhk88-ai-bank-statement-document-automation-by-llm-and-personal-finanical-
