# Scholar Loop: Autonomous AI Research with Deterministic Reward-Hacking Guards

> Scholar Loop is a Python framework that runs an autonomous 8-agent loop modeled on a PhD researcher's workflow: literature search, hypothesis formation, real ML experiments scored by a frozen metric, self-critique, and write-up. The architecture is designed around deterministic guards that prevent the agent loop from gaming its own evaluation.

**renee-jia/scholar-loop** — An autonomous AI scientist: a multi-agent loop over literature, experiments, self-critique and write-up, with deterministic guards against reward-hacking and hallucination.

- Repository: https://github.com/renee-jia/scholar-loop
- Stars: 470 · Forks: 35
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/renee-jia-scholar-loop

## What Scholar Loop Does and Who It Is For

Running ML experiments the way a PhD student does involves more than writing code: it involves reading related work, forming a hypothesis grounded in that literature, running experiments, learning from failures, and eventually writing a paper that places the results in context. Scholar Loop implements this as a multi-agent loop in Python, with each stage handled by a dedicated agent role and each agent's outputs validated against a deterministic schema before the next stage proceeds.

The target user is a researcher who wants to explore automated research loops, either to study the design of such systems or to apply them to structured ML problems. The framework is not a general-purpose AI assistant but a specific pipeline designed around one question: given a performance baseline (for example, 5% error on digit classification), can the loop find something better, and can it do so without cheating?

The anti-cheating design is the distinguishing feature. Scholar Loop assumes that any LLM given the opportunity to optimize a metric will eventually find a way to game it rather than genuinely improve the underlying model. The framework's guards prevent this through a frozen evaluation metric, an edit allowlist for the training script, and a VerifiedRegistry that requires all numbers in the final write-up to trace back to confirmed experiment outputs.

## The Eight-Agent Loop and Its Stages

The loop runs through ten defined roles. The Director reads the experiment ledger and literature trends, then sets the direction, topic, and budget for the next round. The Lit Scout pulls real papers from arXiv and OpenAlex, ranks them by citation impact, and produces structured findings with citations rather than unsourced claims. The Reasoner takes the literature, existing lessons from the skill library, and the performance budget to propose the next experiment.

Before any experiment runs, the Debate Panel (three personas voting on whether the idea is worth GPU time) and a three-tier Funnel screen the proposal. The Funnel runs the idea through smoke, verify, and full stages in sequence; an idea that fails smoke is eliminated without burning the cost of a full run. The Runner executes a real PyTorch experiment and scores it against the frozen metric. The Reflector converts the outcome into a lesson stored in the skill library. The Advisor issues a PROCEED, REFINE, or PIVOT directive to steer the loop. Finally, the Writer drafts a write-up from confirmed experimental findings, and the Reviewer applies a peer-review pass against the draft.

All agents exchange data through typed JSON-schema validated I/O. If an agent returns output that fails schema validation, the harness retries before accepting or failing the stage. A shared audit trace records every agent decision.

## Installing and Running Scholar Loop

Python 3.10 or later is required. The base package (pyyaml and jsonschema) installs with:

```bash
pip install -e ".[dev]"             # pyyaml + jsonschema + pytest
python examples/quickstart.py       # the whole loop in <1s — no GPU, no API key
```

The quickstart runs one idea through the full three-tier funnel using the MockLLM, which requires no API key and runs deterministically. The output shows each tier's result:

```text
baseline to beat: 4.9% val_top1_err
  smoke  3.7644%  [kept]
  verify 3.8004%  [kept]
  full   3.7644%  [kept]    →  climbed 3 tiers, 3 kept
```

For real PyTorch experiments using the bundled digit classification or diabetes regression engines:

```bash
pip install -e ".[dev,engines]"     # adds torch + scikit-learn for the real experiment engines
python examples/campaign_demo.py    # a full campaign on real torch, MockLLM-scripted
pytest -q
```

Live API runs that send requests to an actual language model require the anthropic optional dependency:

```bash
# ".[llm]" adds the Anthropic client, only needed for live API runs (e.g. examples/run_to_paper.py)
```

Two experiment domains ship with the package: digit classification (measuring top-1 error rate) and diabetes regression (measuring RMSE). Adding a new domain requires a YAML profile and an engine pair; the README states this requires no orchestrator changes.

## The Anti-Reward-Hacking Architecture

The central design problem in autonomous research systems is that a language model optimizing an evaluation metric will eventually find ways to report a good score without doing the underlying work. Scholar Loop addresses this with four interlocking mechanisms.

First, the scoring metric is frozen. The train.py experiment engine cannot access the validation set during training, and the metric function is not something the agent can rewrite. The README states this is proven by a bundled cheater engine that demonstrates how a naive implementation can be gamed, and how the harness catches it.

Second, an edit allowlist restricts which parts of the training script the agent can modify. Third, a VerifiedRegistry tracks every number that appears in the final write-up and requires each one to trace back to a confirmed, recorded experiment result. An agent cannot fabricate a number in the paper if it does not exist in the experiment ledger.

Fourth, the optimization loop uses only the frozen metric as its target. There is no LLM-as-judge in the main optimization path: the metric is deterministic code, not an LLM's opinion about quality. Multiple adversarial review passes during development found and fixed real bugs at the correctness and reward-hacking boundaries, according to the README.

## The Skill Library, Calibration, and the Governor

Scholar Loop includes two mechanisms that improve the loop's behavior over time rather than treating each round as independent. The skill library is a persistent store of lessons distilled from past experiment outcomes. Each lesson is relevance-ranked and time-decayed, so lessons about experiments similar to the current proposal surface at the top and older lessons fade in weight. The Reasoner and Reflector agents read and write this library on every round.

Calibration works across the agent predictions. Every agent that makes a checkable claim (for example, the Reasoner's predicted performance delta for an experiment, or the Debate Panel's go/no-go vote) has that claim scored against the ground truth outcome. A CalibrationLog accumulates these scores and feeds them back into the next round's prompts, so the loop learns which of its own agents have been reliable and weights their input accordingly.

The self-stopping governor controls when the loop ends. It accepts three parameters: a dollar budget (max_cost), a round cap (max_rounds), and a convergence criterion (dry_patience for loop-until-dry behavior). This allows the loop to run unattended without requiring a human to decide when to stop.

The framework also uses a parallel population funnel: rather than generating and testing one idea at a time, the Reasoner proposes N distinct ideas, which are smoke-screened in parallel using Python's max_workers threading. Only surviving ideas advance to the verify and full tiers.

## Limitations, Comparison to AI Scientist, and Maintenance

Scholar Loop version 0.0.1 is classified as Alpha (Development Status 3) in pyproject.toml. The two bundled experiment domains (digit classification and diabetes regression) are narrow in scope; applying the framework to a genuinely novel research problem requires writing a new YAML profile and engine pair, which the README describes as a low-overhead addition but does not fully document beyond that description.

The framework currently calls the Anthropic API for live runs. The README does not document adapters for other LLM providers, though the MockLLM interface provides the protocol any adapter would need to implement.

The most direct comparison is to AI Scientist (from Sakana AI), an open-source autonomous research pipeline that generates research ideas, writes experiment code, executes it, and produces a formatted paper. AI Scientist runs end-to-end and has produced publicly reviewed papers. Scholar Loop differs by focusing on the anti-reward-hacking design: it inverts the assumption that the LLM's output is the reliable part and instead makes the deterministic harness the trusted component. The skill library and calibration system also distinguish it from AI Scientist's approach.

The last push to the repository was on 2026-06-23. The repository is not archived. No GitHub releases exist. The licence is MIT.

## Conclusion

Scholar Loop fits researchers and AI engineers who want to study or deploy an autonomous research loop with explicit anti-reward-hacking safeguards, rather than one that relies on LLM self-assessment for correctness. The whole loop runs with a MockLLM and no API key or GPU needed (108 tests cover it), making it straightforward to audit behavior before switching to a real model. Version 0.0.1 is marked alpha in pyproject.toml; check the repository issues and the examples directory before relying on it for research that matters. MIT licence.

## FAQ

### Does Scholar Loop require a GPU or an API key to run?

No. The quickstart and the 108-test suite both run against a deterministic MockLLM that requires neither a GPU nor an API key. Real experiment engines (digit classification, diabetes regression) require torch and scikit-learn. Live API runs that send requests to an actual language model require installing the anthropic optional dependency and providing a key.

### How does Scholar Loop prevent reward hacking in its experiments?

The scoring metric is frozen code that train.py cannot access during training and cannot be rewritten by the agent. An edit allowlist restricts which parts of the training script the agent can modify. A VerifiedRegistry requires every number in the final write-up to trace back to a confirmed recorded experiment result, preventing the agent from fabricating data.

### What experiment domains does Scholar Loop support out of the box?

Two domains ship with the package: digit classification (measured by top-1 validation error) and diabetes regression (measured by RMSE). Adding a new domain requires a YAML profile and an engine pair, and the README states this requires no changes to the orchestrator.

## Sources

- [Issues](https://github.com/renee-jia/scholar-loop/issues)
- [License: MIT](https://github.com/renee-jia/scholar-loop/blob/main/LICENSE)
- [README](https://github.com/renee-jia/scholar-loop/blob/main/README.md)
- [renee-jia/scholar-loop on GitHub](https://github.com/renee-jia/scholar-loop)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/renee-jia-scholar-loop
