Model or dataset
xbtlin/ai-berkshire avatar
xbtlin/ai-berkshire

AI Berkshire: a Claude Code investment research framework built on four value investors' methods

AI 时代的伯克希尔:基于 Claude Code / Codex 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built for Claude Code / Codex. 4 masters' methodologies + multi-agent adversarial analysis.

16,561 stars2,476 forksPythonMIT

At a glance

What is it?
AI Berkshire packages 20 research skills for Claude Code and Codex around Buffett, Munger, Duan Yongping and Li Lu. It forces a verdict instead of a balanced essay, and it spends tokens to get there.
Who is it for?
Adopt AI Berkshire if you already run Claude Code or Codex, you research listed companies repeatedly, and you want a fixed checklist and a forced verdict rather than a chat answer. Do not adopt it if you want a turnkey data terminal, if your research is one-off, or if you cannot absorb the token cost of multi-agent runs; the README states that deep research skills consume a lot of tokens by design.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem AI Berkshire targets: chat answers you cannot act on

Ask a general assistant whether a stock is worth buying and you get a two-sided essay that ends with a disclaimer. That output is readable and useless. The project's README makes this the central pitch: the difference is not whether an AI can analyze a company, it is whether the analysis carries decision discipline. AI Berkshire forces one of three verdicts (pass, fail, grey zone) with price bands attached, and it applies a fixed checklist so that seven companies can be compared on the same scale. The intended user is a single analyst or a small fund running repeatable research on listed companies, not someone who wants a market data terminal. The repository is Python, MIT licensed, and its last push was on 2026-04-07, so treat it as a frozen snapshot rather than a moving target.

Four masters, four agents, and the conflict between them

The architecture has three layers, described in the README as Skill, Agent and Tool. The Skill layer is 20 named entry points such as /investment-research, /earnings-review and /quality-screen. The Agent layer is where team skills like /investment-team and /earnings-team dispatch four agents in parallel, one per master perspective. Each agent searches independently, verifies data independently and returns its own score. The README's worked example is Pinduoduo: Duan Yongping scores the business model 3.7 out of 5, Buffett scores the financials 4.4, Munger scores the moat 3.5, and Li Lu scores ten-year certainty 2.0. Buffett concludes it is cheap while Li Lu concludes uncertainty disqualifies it. That disagreement is the product. A single prompt cannot manufacture four independent positions, and the project treats the friction as the anti-blind-spot mechanism. The Tool layer handles arithmetic, retrieval and report spot checks. Lightweight skills skip the agent layer entirely and call tools directly.

Installing the skills and running a first research pass

The repository is a skill collection rather than a package on an index, so setup means placing skill files where your client looks for them. The README's quick start section is the authority here; the repository layout shows skills/ for Claude Code skill definitions and codex-skills/ plus codex-prompts/ for the Codex side. Clone the repository first.

bash
git clone https://github.com/xbtlin/ai-berkshire.git
cd ai-berkshire

Before running anything expensive, check that the financial verification tool works on your interpreter. The README gives this exact invocation as its market-cap sanity check, using a price, a share count, a reported figure and a currency.

bash
python3 tools/financial_rigor.py verify-market-cap \
  --price 510 --shares 9.11e9 --reported 4.65e12 --currency HKD

The README states this returns a pass with a deviation of about 0.08 percent when the inputs agree. If the command errors, fix your Python environment before spending tokens on a research run. Then invoke a skill from your client, for example /investment-research for a full four-master analysis of one listed company, or /investment-checklist when you only want the ten-minute six-gate screen. The README notes that deep research skills are token-hungry because they run multiple rounds and cross-verification; start with the checklist skill to calibrate cost.

Where the framework is thin or the wrong tool

The README is candid about information scarcity but less so about the operational edges. An information-richness grade of A, B or C is attached to each target, and the README's own example grades Pop Mart as B with confidence intervals on derived metrics. That is honest labeling, not a fix. If your target is a small-cap with two broker notes and no English filings, the framework will produce a grey-zone verdict and a lot of caveats, and you have spent the tokens anyway. Second, the framework depends on web search being available to each agent; when sources disagree on units, the README's own Tencent example shows how easily a market cap flips between HKD billions and RMB billions. The verification tool catches arithmetic, not source selection. Third, this is not a portfolio tracker or a broker integration. There is no execution, no live pricing feed, and no position ledger beyond what the portfolio skills describe. If you want a screener that returns ranked tickers from a live database, this is the wrong layer.

How it differs from a general research assistant

The closest comparison is the built-in /deep-research orchestrator that ships inside Claude Code. The README is explicit that this one is not distributed by the repository: it comes with the client. Its flow splits a question into five search angles, extracts falsifiable claims, and puts each claim through three independent adversarial agents where two votes of refutation remove it. AI Berkshire's own skills do not replicate that; the README recommends running /deep-research first to check a single factual judgement, then running the repository's stock or industry skills on top. The practical difference is scope. /deep-research verifies facts with citations. AI Berkshire applies an investment methodology to those facts and returns a verdict with price bands. Using one without the other leaves either the numbers unverified or the methodology absent.

Maintenance, licence and the cost of upgrading

The repository is not archived, but its last push was on 2026-04-07 and the only listed release is v1.0.0 from the same date. Five months pass between that push and any review you read now, so plan for a static codebase. Upgrading means pulling the repository and re-copying skill files into your client, because the skills are markdown definitions rather than a versioned dependency. There is no package manager to resolve conflicts for you. The MIT licence permits commercial use and modification; the repository ships a LICENSE file at the top level. Two non-legal points the licence does not cover: the README's performance figures come from the author's own brokerage account and the README itself states past returns do not indicate future results, and the framework's outputs are research documents, not investment advice. If you fork it, the reports/ and data/ directories contain the author's research output, which you will want to clear before using it with client material.

Editorial conclusion

Adopt AI Berkshire if you already run Claude Code or Codex, you research listed companies repeatedly, and you want a fixed checklist and a forced verdict rather than a chat answer. Do not adopt it if you want a turnkey data terminal, if your research is one-off, or if you cannot absorb the token cost of multi-agent runs; the README states that deep research skills consume a lot of tokens by design. Verify three things first: which skills your client version can actually load, whether the reports/ and data/ directories contain the inputs you need, and whether tools/financial_rigor.py runs on your Python before you trust any computed figure in an output report.

Frequently asked questions

What is AI Berkshire and who is it for?

It is a collection of 20 investment research skills for Claude Code and Codex, built around the methods of Buffett, Munger, Duan Yongping and Li Lu. The README positions it for one person doing professional-grade research on listed companies, with a fixed checklist and a forced pass, fail or grey-zone verdict.

How do I install AI Berkshire for Claude Code?

The README's quick start treats it as a skill collection rather than an installable package: clone the repository, then place the skill definitions from skills/ where your client loads them. The codex-skills/ and codex-prompts/ directories cover the Codex side.

Does AI Berkshire verify the financial numbers in its reports?

It ships tools/financial_rigor.py, and the README shows a market-cap check that compares price times shares against a reported figure and passes at roughly 0.08 percent deviation. The README also states calculations use Python decimal.Decimal rather than float, and that key data needs at least two independent sources.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xbtlin-ai-berkshire.svg)](https://hysenlabs.com/projects/xbtlin-ai-berkshire)