# AgentEval: a loopback proxy that grades its own capture, with the judge key missing from the sample file

> AgentEval is a Rust proxy that sits between an agent and an LLM API, splits the captured traffic into sessions, scores four weighted dimensions, runs ten diagnostic rules, and lets a read-only agent review the target project's configuration. The sample environment file and the documented setup disagree on the upstream endpoint and omit the judge credentials, which is where a first run breaks.

**canwhite/AgentEval** — The agent responsible for conducting the agent evaluation

- Repository: https://github.com/canwhite/AgentEval
- Stars: 452 · Forks: 4
- Language: Rust
- License: not declared
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/canwhite-agenteval

## The description calls it an agent, the build is a proxy

The repository description says this is the agent responsible for conducting the agent evaluation, while the page itself describes a transparent HTTP proxy with a dashboard attached. The manifest settles it: the package is named AgentEval at version 0.1.0 on edition 2021, and the dependency list is server plumbing rather than agent machinery. axum 0.8.9 serves the proxy and the dashboard, reqwest 0.13.4 with streaming and JSON features forwards upstream, tokio runs with the full feature set, and serde, serde_json, dotenvy, chrono, futures, bytes, http-body, and http-body-util fill in the rest. No model client, no tool-calling library, no vector store. The flow is short: traffic arrives at the proxy on 127.0.0.1 port 57633, raw requests go to logs/{stem}.jsonl, session detection runs on the fly, and a structured view is written to logs/{session}.view.json.

## .env.example sends upstream to DeepSeek and lists no judge credentials

Two different setups live in this repository, and the one you are told to copy is not the one shipped. The documented .env block names an edge function endpoint as the upstream and adds five variables on top of it:

```bash
# Upstream LLM API / 上游 LLM 地址
AGENTEVAL_UPSTREAM=https://api.edgefn.net

# Proxy port / 代理监听端口
AGENTEVAL_PORT=57633

# Log directory / 日志目录
AGENTEVAL_LOG_DIR=./logs

# Judge LLM for grading, diagnosis summary, and probing / 评测 LLM（评分+诊断总结+探针共用）
AGENTEVAL_JUDGE_API_BASE=https://api.deepseek.com
AGENTEVAL_JUDGE_MODEL=deepseek-chat
AGENTEVAL_JUDGE_API_KEY=sk-xxx

# Source project directory for probe / 被探针审查的 agent 项目路径
PROBE_SOURCE_PROJECT_DIR=/path/to/your/agent/project
```

The sample file .env.example points AGENTEVAL_UPSTREAM at api.deepseek.com instead, carries no judge variables at all, comments the log directory out on the grounds that it defaults to a path under the home directory rather than ./.logs, and introduces AGENTEVAL_VERBOSE, a flag that prints request bodies and appears nowhere in the quick start.

## Session boundaries come from message shrinkage or two idle minutes

A session ends in one of three ways, and all three are mechanical. A new conversation shrinks the message array back to the system prompt plus a new question, and when the common prefix between the old and new arrays is 1 or less the old session is sealed, graded in the background, and a new one starts. Two idle minutes with no new request does the same thing. Process shutdown flushes the last session with a synchronous grade instead. The rule keys on message history rather than on time alone, which has two consequences worth knowing before you point an agent at it. An agent that sends each request without carrying history looks like a brand new conversation every time, so it produces one session per request. And any pause longer than two minutes, including a long tool call or a slow model, closes the session and splits the work in two, each half graded on its own. A sealed session is meant to be read rather than summarised: the detail view expands turn by turn into user input, reasoning, text, tool calls, and results, with severity, category, detail, and evidence on each issue and confidence, root cause, recommendation, and evidence on each finding.

## Just over half the score comes from the judge, and the dashboard rescales it

Four weighted dimensions produce one overall score between 0 and 1, and the weights decide where the trust goes. task_completion carries 0.35 and comes from the LLM judge. tool_efficiency carries 0.30 and is rule-based, counting tool errors, duplicate calls, and call patterns. response_quality carries 0.20 from the judge, measuring accuracy, conciseness, and substance. performance carries 0.15 from rules, covering token efficiency, latency, and turn count. The two judge dimensions add to 0.55 of the total, and they read from the single model named in AGENTEVAL_JUDGE_MODEL, which the same configuration also uses for the diagnosis summary and the probe. When that model is unavailable the tool falls back to rule-based estimates, so the largest single dimension changes source and nobody is told which path produced the number. The dashboard then puts the 0 to 1 result on a ten point display, since its filter bar splits into below six, six to eight, and above eight with low scores first.

## The CLI documents diagnose and probe but never grading

Two operations are available from the terminal, and a third exists only in the browser:

```bash
# Run diagnosis on a session / 对某个 session 运行诊断
cargo run -- diagnose <session_id> [--format terminal|json]

# Run probe on a session (requires prior diagnosis) / 对某个 session 运行探针
cargo run -- probe <session_id>
```

Diagnosis accepts a session id and an output format of terminal or json. Probe takes a session id and nothing else, because it is documented as requiring a prior diagnosis, which makes the two-step order a constraint rather than a suggestion. There is no documented command for grading a session from the terminal, even though grading is the first of the three panels in the detail view and the one the session list offers a button for. So a scripted workflow has to trigger the browser path, or wait for the background grade that fires when a session is sealed.

## Ten rules in four categories, and a data check in a category of one

Diagnosis is a pure rule engine with no model in the loop, running ten checks across four categories, after which a model writes a two or three sentence summary and that step is skipped when there is no API key. The tool category holds four rules: result_missing, result_error, duplicate_3plus, and result_empty, covering broken tool chains, retry loops, and silent failures. The prompt category holds two, bloat and context_overflow, for overlong system prompts and orphaned tool call ids. The token category holds three, empty_response, waste, and excessive_input. The view category holds exactly one, mismatch, and it is the only check in the set that looks at the capture rather than the conversation, comparing the turn count against the number of entries in the JSONL file. One rule guarding the recorder means a recorder bug that inflates the count is caught by the tool built on top of it.

## list_dir is the only probe tool without a stated cap

The probe is a model agent carrying four read-only file tools that enters the target project's directory and reviews its configuration files, and every recommendation lands in the report and is never applied. Three of the four tools have published limits: read_file truncates at 1MB, grep stops at 1000 lines, and glob returns at most 2000 entries. list_dir has no limit stated at all. The guard rails around the run are tighter than the tool caps: paths are sandboxed so .. is rejected and the canonical form is verified again, three identical calls trigger an injected warning, the run stops after 30 steps, and the model gets 300 seconds. Loop detection only catches identical calls, so varied reads of the same tree are bounded by the step limit instead, and one slow step can consume the entire timeout. The files it looks for are named in the text, CLAUDE.md among them, even though the quick start also points OpenAI SDK users at the proxy.

## A bugs directory, no tests, no workflow, and no license file

The root of the repository holds .env.example, .gitignore, Cargo.lock, Cargo.toml, README.md, architecture-summary.md, and four directories: bugs, docs, and src. There is no test directory, no CI configuration directory, and no Dockerfile, which is worth noting for a tool whose entire purpose is measurement. A committed bugs directory in a 0.1.0 package suggests the known defects are tracked in the tree rather than in an issue tracker. The manifest declares no license field and no LICENSE file is listed, so the terms are not stated anywhere in what is published here. The output file listing also stops partway: after {stem}.jsonl for raw traffic and {stem}_{N}.view.json for the structured session view, the next line is a single opening brace and the remaining file names do not appear.

## Conclusion

AgentEval is a measurement tool built for a single operator on one machine, and it should be read that way. The proxy binds to 127.0.0.1, raw traffic lands in plaintext JSONL under the log directory, and the upstream is whatever AGENTEVAL_UPSTREAM names, so point it only at an endpoint you already trust with your prompts. Before trusting any score, check that a judge key is set: without one, grading falls back to rule-based estimates on the two dimensions that carry 0.45 of the weight, the diagnosis summary disappears, and the probe has nothing to summarise. Nothing here is auto-applied, and the repository ships no tests and no license.

## FAQ

### What does AgentEval need in its environment to grade sessions?

The README sets AGENTEVAL_UPSTREAM, AGENTEVAL_PORT, AGENTEVAL_LOG_DIR, the three AGENTEVAL_JUDGE variables, and PROBE_SOURCE_PROJECT_DIR. The shipped .env.example sets only the upstream, the port, and commented log and verbose lines, with no judge credentials at all.

### How does AgentEval decide that a new session has started?

A conversation normally grows its message array turn by turn, and a new one shrinks back to the system prompt plus a new question. When the common prefix length between the two is 1 or less, the old session is sealed and graded in the background. Two idle minutes and process shutdown also end a session.

### What happens to an AgentEval score when the judge model is unavailable?

Grading falls back to rule-based estimates. The two dimensions carried by the LLM judge are task_completion at 0.35 and response_quality at 0.20, together 0.55 of the weight, and the diagnosis summary is skipped when there is no API key.

### Can the AgentEval probe change my agent's files?

No. The probe carries read-only file tools with no write or execute capability, every recommendation is written to the report, and nothing is ever applied automatically. Paths are sandboxed with .. rejected plus a canonicalize check, three identical calls trigger a warning, and the run is capped at 30 steps and 300 seconds.

## Sources

- [canwhite/AgentEval on GitHub](https://github.com/canwhite/AgentEval)
- [Issues](https://github.com/canwhite/AgentEval/issues)
- [README](https://github.com/canwhite/AgentEval/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/canwhite-agenteval
