Open-source project
chrisworsey55/atlas-gic avatar
chrisworsey55/atlas-gic

atlas-gic: a repository of LLM prompts treated as trainable weights

ATLAS by General Intelligence Capital — Self-improving AI trading agents using Karpathy-style autoresearch

2,187 stars392 forksPythonNOASSERTION

At a glance

What is it?
ATLAS runs a panel of trading agents whose prompts are rewritten by an autoresearch loop and scored by rolling Sharpe ratio, kept or reverted through git, with the backtest results published in the README.
Who is it for?
The interesting part of this repository is that the optimised artifact is a prompt rather than a model, which makes the whole loop legible in a diff, and the git commit or revert decision gives you an audit trail that most prompt engineering does not have. What the repository does not give you is the code path that turns a Sharpe measurement into a rewritten prompt, since the tree is four directories and the README is where the mechanism lives.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 135 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four directories, and the prompts are the weights

The tree is short enough to describe in one line: `src/`, `prompts/`, `architecture/` and `results/`, plus a LICENSE and a README. There are no releases, no package manifest, no dependency file and no contributing guide.

That sparseness is itself the clearest statement of the design. The README says the agent prompts are the weights and that the Sharpe ratio is the loss function, and it credits the idea to Karpathy's autoresearch loop with Soros's reflexivity and a swarm simulation project layered on top. No GPU is needed, in the project's phrasing, because nothing is being gradient trained.

So the unit of improvement is a text file in `prompts/`. An agent performs poorly, the system generates one targeted modification to its prompt, it runs for a defined period, and the modification is either kept as a git commit or thrown away with a revert. Each trading day is described as one iteration of that loop, and the README positions this as the reason a small virtual machine can stand in for expensive hardware.

Twenty five agents in four layers, with a chief investment officer on top

The architecture is the part of the README that is most concrete, because it names every agent and its job. Layer one is macro, with ten agents covering the central bank, geopolitics, China, the dollar, the yield curve, commodities, volatility, emerging markets, news sentiment and institutional flow. These set the regime, in the project's words: risk on or risk off.

Layer two is seven sector desks, semiconductors, energy, biotech, consumer, industrials and financials, plus a relationship mapper that tracks supply chains, ownership, analyst coverage and competitive dynamics.

Layer three is four superinvestor personas, each with a stated question. Druckenmiller asks what the big asymmetric trade is, Aschenbrenner asks who benefits from the compute capital expenditure cycle, Baker asks who has real intellectual property moats, and Ackman asks about pricing power, free cash flow and catalysts. Layer four is the decision layer: an adversarial risk officer that attacks every idea and hunts correlated risks, an alpha discovery agent that looks for names nobody else mentioned, an autonomous execution agent that converts signals into sized trades, and a CIO that synthesises the layers weighted by the agent scores.

The stated total is 25 agents debating markets daily, and a later section says the count grew from 25 to 31 through autonomous spawning.

Darwinian weights between 0.3 and 2.5, adjusted daily

The scoring mechanism is a weighting scheme rather than a rank. Each agent carries a weight with a floor of 0.3 and a ceiling of 2.5. Top quartile agents get multiplied by 1.05 each day, bottom quartile by 0.95, and the CIO uses those weights to decide how much each layer counts.

The phrasing in the README is that good agents get louder and bad agents get quieter. It is a simple exponential moving adjustment on a bounded scale, and the bounds matter more than the rate: an agent that keeps losing decays toward 0.3 and stops influencing the final call without ever being deleted, which is a gentler failure mode than removing it.

One detail in the backtest section is worth pulling out because it is the kind of thing that only shows up after running the thing. The README says the system discovered independently that its own portfolio manager was its weakest component and downweighted it before the authors diagnosed the same issue by hand. That is a plausible story about a weighted ensemble, and it is also a claim about an unreviewed measurement, so both readings are available.

The autoresearch loop itself is five steps: identify the worst agent by rolling Sharpe, generate one targeted prompt modification, run for five trading days, check whether Sharpe improved, then keep the commit or reset.

The backtest table includes the cohorts that failed

The README reports an eighteen month backtest from September 2024 to March 2026 covering 378 trading days. Of 54 prompt modifications attempted, 16 survived and 37 were reverted. The deployment phase return is stated as 22 percent over 173 days, and the best individual pick is named as a semiconductor holding bought at 152 dollars for a 128 percent gain.

Those are the author's own figures and there is no code in the tree to check them against, so they carry exactly as much weight as any other self-reported backtest.

The more informative table is the regime one, because it reports losses. Five cohorts were trained separately: bull and low volatility from 2016 to 2018 returning 7.7 percent with 35 percent of modifications kept, COVID in the first half of 2020 down 13.1 percent with none of three modifications surviving, rate tightening in 2022 and 2023 down 30.2 percent with 43 percent kept, recovery in late 2020 down 29 percent with a single modification, and euphoria in 2021 up 14.3 percent with 49 percent kept.

The lesson column is where the loop's limits become clear. Crashes move too fast for a five day feedback cycle, which is why two cohorts kept almost nothing. The rate tightening cohort produced a specific behavioural rule about not reversing during Federal Reserve weeks with a fifteen day minimum between reversals. A method that only works in calm trending markets is a much more limited thing than the deployment headline suggests.

Agents that spawn agents, and three that went extinct

The agent spawning section describes the system watching its own debates for recurring blind spots. When the same gap shows up three or more times in five days, it creates a new specialist agent at neutral weight, with no human deciding what or when.

A six month spawning test from July to December 2024 produced nine agents covering credit markets, the earnings calendar, options flow, liquidity conditions, positioning data, earnings guidance, retail sentiment and technical levels. Three went extinct by sitting at minimum weight for more than twenty days, six survived selection and reached maximum weight.

Three out of nine reaching the ceiling is a strikingly high hit rate for an autonomous process, and the failure mode here is clean: an agent that keeps losing stops mattering and eventually withers. There is no cleanup process described for extinct agents, which implies dead weight accumulates in `prompts/` unless someone prunes it.

The README also states the platform built on top of this repository launched as a commercial product with three paid tiers, a copy trading feature, an agent builder and a marketplace. The public repository is the open research artefact behind that product rather than the product itself.

Read the README's own caveats before the claims

There is an unusual amount of hedging in the README, and it clusters in the places where the numbers are weakest. The launching announcement says copy trading is up 30 percent since launch, and that a separate prediction market agent is live with a 60 percent win rate, both without a period, a sample size or a drawdown figure. A 60 percent win rate on a prediction market says nothing about profitability unless the payoff distribution is given.

The offsets are commercial. Three subscription tiers are priced in the README, from a free leaderboard tier to a 499 dollar a month tier that includes an eighteen month backtest dataset and marketplace publishing, and a discount code aimed specifically at people arriving from GitHub.

The absence of a licence identifier is worth noting separately. The project carries no standard open source licence, which means the default is all rights reserved regardless of the LICENSE file's contents, and for a repository whose entire value is its prompt files that distinction matters.

The last push was on 2026-05-27. There are no releases, so the commit history is the only record of how the architecture evolved, and with four open issues there is very little public discussion to read alongside it.

Editorial conclusion

The interesting part of this repository is that the optimised artifact is a prompt rather than a model, which makes the whole loop legible in a diff, and the git commit or revert decision gives you an audit trail that most prompt engineering does not have. What the repository does not give you is the code path that turns a Sharpe measurement into a rewritten prompt, since the tree is four directories and the README is where the mechanism lives. Read the regime table before anything else, because the negative cohorts are the honest part of the record and they say the loop needs time it did not have in a crash. If you want to reproduce this, `prompts/` and `architecture/` are the two places to start, and treat every return figure in the README as the author's own report.

Frequently asked questions

What is ATLAS AI trading?

It is a panel of LLM agents that analyse markets across four layers, from macro regime through sector desks and investor personas to a final decision layer. Each agent's prompt is rewritten based on how it performs, and agents are weighted by rolling Sharpe ratio, with the highest weighted layers having more influence on the final call.

How does ATLAS improve its trading agents?

The system identifies the worst performing agent by rolling Sharpe ratio, generates one targeted modification to that agent's prompt, runs it for five trading days, and then keeps the change as a git commit or resets it. The prompts themselves are treated as the parameters being optimised, so no GPU training is involved.

Does the atlas-gic repository include the code or a licence?

The repository tree holds four directories, `src/`, `prompts/`, `architecture/` and `results/`, plus a README and a LICENSE file, with no releases and no package manifest. No standard open source licence identifier is attached to the project, so check the LICENSE file before reusing anything from it.

Official sources

  1. chrisworsey55/atlas-gic on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chrisworsey55-atlas-gic.svg)](https://hysenlabs.com/projects/chrisworsey55-atlas-gic)