# LLMs-in-Finance: Notebooks for Agentic, RAG and Multimodal Finance Workflows

> A Jupyter Notebook collection from hananedupouy that demonstrates OpenAI Agents SDK, AutoGen, LlamaIndex, CrewAI, LangGraph and Anthropic patterns on finance tasks. It is a learning and prototyping resource, not a trading system.

**hananedupouy/LLMs-in-Finance** — LLMs in Finance - Generative AI - AI Agents 

- Repository: https://github.com/hananedupouy/LLMs-in-Finance
- Stars: 886 · Forks: 231
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-13 · Updated: 2026-09-13 · Language: en
- Canonical page: https://hysenlabs.com/projects/hananedupouy-llms-in-finance

## What LLMs-in-Finance actually solves, and for whom

Most engineers who want to apply generative AI to finance hit the same wall: the frameworks are documented with generic examples (weather lookups, travel planning), while the finance-specific problems are parsing a 200-page earnings report, reading a chart image, or coordinating several agents around a strategy. LLMs-in-Finance is a set of Jupyter notebooks that fills that gap with worked examples across four areas: AI Agents in Finance, RAG, Multimodals, and Papers in Finance. The repository description calls it "Exploring how to apply GenAI to real-world financial workflows with AI Agents, RAG, and Multimodal LLMs".

The audience is narrow and worth stating plainly. It is for analysts, quants and engineers who are comfortable in notebooks and want to see how a framework behaves before committing to it. It is not a library you import, not a service you deploy, and not a dataset. There is no package on any registry. The top-level entries are directories: Agents/, Datasets/, IMA_livre_blanc/, Multimodal_llms/, Papers/, RAG/, plus LICENSE and README.md. If you need an installable dependency, this is the wrong repository.

## How the notebooks are organised across agent frameworks

The Agents/ directory is split by framework rather than by use case, which is the more useful axis here because the frameworks differ more than the tasks do. The README lists subdirectories for OpenAI, AutoGen, llamaIndex, CrewAI, LangChain, and Anthropic.

The OpenAI folder covers two patterns: an autonomous strategy code review using LLM-as-a-judge, and a financial news bot built as an evaluator-optimizer multi-agent system, again with an LLM acting as judge. The AutoGen folder works through configuring a financial agent that fetches data and develops a momentum trading strategy, then a second notebook on collaborative task management that optimises that strategy. LlamaIndex appears with a code interpreter agent paired with Anthropic for one-click stock performance analysis, an introspective agent worker, and a multi-agent fundamental analysis workflow. CrewAI is used for multi-agent collaboration covering Apple trend analysis, sentiment extraction from news articles, and proposed trading strategies. LangGraph, filed under Agents/LangChain, walks from concept to execution and reflection for a trading strategy. Anthropic gets its own folder for a financial agentic system combining Anthropic, LlamaIndex and APIs.

That structure means the comparison you get is behavioural, not benchmarked. You read how each framework expresses the same idea, for example an evaluator loop or a data-fetching agent. The README does not publish accuracy numbers, latency or cost comparisons between them, so any judgement about which framework performs better has to come from running the notebooks yourself.

## Installing nothing: opening the first notebook

There is no install step in the README. No pip command, no environment file, no requirements.txt is described. The README points at notebook paths, so the practical route is cloning the repository and opening a notebook in Jupyter. The README gives this example path for a RAG notebook:

```bash
RAG/llamaindex/Financial_Report_Analysis_with_LlamaParse_AutoMode.ipynb
```

Because the notebooks call hosted models, expect to supply your own API credentials for whichever provider a given notebook uses. The README names OpenAI, Anthropic, LlamaIndex, AutoGen, CrewAI and LangGraph, but it does not document an environment variable name for any of them, so read the first cells of the notebook before running it. There is also a Datasets/ directory at the top level, which is where a notebook would look for local input if it does not fetch data over the network.

For a first real use, pick the notebook whose framework you already know and whose task you already understand. The RAG evaluation folder is the most self-contained starting point because it is about measuring a pipeline rather than building a strategy:

```bash
RAG/evaluation
```

The README states that this folder contains an example on how to evaluate your RAG system with DeepEval and GiskarAI. That is a bounded exercise: you bring a retrieval pipeline, and the notebook shows the measurement step. What you should see is an evaluation workflow, not a trading signal.

## The multimodal notebooks and what they are really testing

Multimodal_llms/financial_analysis is the part of the repository with the clearest research question. The README says the goal is to evaluate how effectively models such as Claude Sonnet 3.5, GPT-4o and the o1 reasoning model interpret complex charts found in financial reports. Chart interpretation is a genuinely hard case: a revenue bar chart with a broken axis, or a candlestick chart with annotations, tests whether the model reads the axis labels or pattern-matches a plausible narrative.

The limitation is that the repository frames this as evaluation but the README does not describe a scoring rubric, a ground-truth set, or a result table. So the notebooks give you a way to run the comparison, not a published verdict on which model reads charts better. If you need a defensible answer for a procurement decision, you will have to build the scoring yourself from the notebook outputs. Treat the folder as a harness template.

There is a related constraint that applies across the whole repository. Financial charts and reports are frequently licensed material. The repository includes a Datasets/ directory, but the README does not state the provenance or redistribution terms of anything inside it, and the MIT licence on the repository covers the code, not third-party documents you might feed into a notebook.

## Where this repository is the wrong tool

The README carries its own disclaimer, and it is unusually direct: the project is intended solely for educational and research purposes, it is not designed for real trading or investment use, no warranties or guarantees are provided, and the creator bears no responsibility for any financial losses. That is not boilerplate to skim past. Several notebooks develop momentum trading strategies and propose trading strategies from sentiment. Those outputs are demonstrations of an agent workflow, and the repository explicitly disclaims them as investment input.

Beyond the disclaimer, there are structural reasons not to build on it. Notebooks are not a deployment artefact: there is no packaging, no CLI, no service entry point and no test suite described in the README. There is also no version pinning documented, which matters because the frameworks involved (AutoGen, CrewAI, LlamaIndex, the OpenAI Agents SDK) change their APIs frequently, and a notebook written against one release may not run against the next. The README documents no rollback or migration guidance, because there is nothing to roll back to: you clone a snapshot and read it.

Finally, the repository is a single author's teaching material. The last push was on 2026-09-12, so it is recent, but recency of a push says nothing about whether a specific notebook still runs against the current version of the framework it imports.

## How it differs from a finance LLM benchmark or a finance-tuned model

People searching around this topic often arrive looking for something else: a benchmark leaderboard or a model fine-tuned on financial text. LLMs-in-Finance is neither. It does not train or ship a model, and it does not publish a scored leaderboard. It is a set of application notebooks that call hosted models through third-party frameworks.

The practical difference shows up the moment you have a requirement. If your question is "which model should we route financial questions to", a benchmark suite with a fixed task set and published scores answers it, and this repository does not. If your question is "how do I wire an evaluator-optimizer loop around a news feed, or attach a code interpreter to a stock analysis agent", this repository answers it with runnable notebooks, and a benchmark does not. The two are complements, and the README's Papers section, described as a curated summary of research papers on GenAI in financial applications, is the part of this repository closest to the benchmark world, though it is a reading list rather than an evaluation harness.

A second alternative is to start from the framework's own documentation and adapt its generic examples. That gives you current API coverage and official maintenance. What you lose is the finance framing: the earnings-report parsing, the chart reading, the strategy-review loops. Whether that framing is worth the risk of a notebook that lags its framework's API is the decision this repository forces on you.

## Licence, maintenance and the cost of keeping notebooks running

The repository is MIT licensed. That is permissive for the code in the notebooks: you can reuse and adapt it, subject to the usual attribution and warranty terms, and it is compatible with commercial use of the code itself. Two caveats are worth separating. First, MIT covers what the author wrote, not the model outputs a notebook produces or the third-party documents you feed in; those carry their own terms from the model provider and the document source. Second, running the notebooks costs money at the model provider, since every example calls a hosted model. The README does not estimate token usage or cost for any notebook. None of this is legal advice; check the terms of the providers and data sources you actually use.

Upgrade cost is the real ongoing expense. Because the notebooks import fast-moving frameworks, the maintenance burden sits with you, not with the repository. There are no releases listed, so there is no changelog to read before pulling. When a framework changes an API, the notebook fails at the import or the first agent call, and your fix is to read the framework's current documentation and patch the cell. Budget for that: a notebook that ran six months ago is a starting point, not a working artefact.

## Conclusion

Adopt it as a teaching and prototyping reference if you already work in notebooks and want to compare agent frameworks on finance-flavoured tasks. Do not adopt it as a production trading or investment component: the README states it is for educational and research purposes only and not designed for real trading or investment use. Before relying on any notebook, open it and confirm which API keys and paid model access it expects, since the repository documents no environment file or pinned dependency set.

## FAQ

### What is LLMs-in-Finance?

It is a Jupyter Notebook repository that explores applying generative AI to financial workflows, organised into AI Agents in Finance, RAG, Multimodals and Papers in Finance. It demonstrates frameworks including the OpenAI Agents SDK, AutoGen, LlamaIndex, CrewAI, LangGraph and Anthropic.

### Which LLM is best at finance?

The repository does not rank models. Its multimodal notebooks are described as evaluating how effectively models such as Claude Sonnet 3.5, GPT-4o and the o1 reasoning model interpret complex charts in financial reports, but the README publishes no scoring rubric or results table.

### Is LLMs-in-Finance suitable for real trading or investment use?

No. The README states the project is intended solely for educational and research purposes, is not designed for real trading or investment use, and provides no warranties or guarantees.

### How do I install LLMs-in-Finance?

There is no documented install step. The README points at notebook paths inside the repository, so the practical route is cloning it and opening the notebooks in Jupyter, supplying your own credentials for the model providers each notebook calls.

### Does LLMs-in-Finance include RAG evaluation examples?

Yes. The README states that the RAG/evaluation directory includes an example on how to evaluate your RAG system with DeepEval and GiskarAI.

## Sources

- [hananedupouy/LLMs-in-Finance on GitHub](https://github.com/hananedupouy/LLMs-in-Finance)
- [Issues](https://github.com/hananedupouy/LLMs-in-Finance/issues)
- [License: MIT](https://github.com/hananedupouy/LLMs-in-Finance/blob/main/LICENSE)
- [README](https://github.com/hananedupouy/LLMs-in-Finance/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hananedupouy-llms-in-finance
