Model or dataset
agentscope-ai/OpenJudge avatar
agentscope-ai/OpenJudge

OpenJudge: a Python grader library for scoring AI agents and turning grades into rewards

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

860 stars74 forksPythonApache-2.0

At a glance

What is it?
OpenJudge (package name py-openjudge) ships 50+ built-in graders, rubric generation, and a Streamlit UI for evaluating agents, text, code and multimodal output. This review covers what it does, how to install and run a first grader, and where the framework stops being the right tool.
Who is it for?
Adopt OpenJudge if you already have a Python evaluation loop and need ready-made graders for agent trajectories, tool selection, memory or multimodal coherence, or if you want to convert grading output into reward signals for fine-tuning. Do not adopt it as a hosted evaluation service: the repository is a library plus a Streamlit UI, and the online playground at openjudge.me/app is a separate product surface.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap OpenJudge fills between ad hoc scripts and a full eval platform

Teams building AI agents usually start evaluation with a handful of assertions and a scoring prompt pasted into a notebook. That works until the criteria change, at which point every script has to be edited in parallel. OpenJudge's answer is a library of named graders plus a workflow the README states directly: collect test data, define graders, run evaluation at scale, analyze weaknesses, iterate. The intended users are developers evaluating AI applications such as agents or chatbots, and the second audience is teams that want grading output as reward signals for fine-tuning. The project is written in Python, requires Python 3.10 or newer according to the PyPI badge, and is licensed Apache-2.0.

What is inside the grader taxonomy: general, agent and multimodal

The README groups the built-in graders into three families. The general set covers semantic relevance, text similarity, code syntax validation and JSON structure matching. The agent set is lifecycle-oriented rather than outcome-only: tool selection accuracy, memory or context preservation, plan feasibility and trajectory quality. The multimodal set handles image-text coherence, text-to-image generation quality and image helpfulness. That agent grouping is the part worth noting. Most evaluation libraries score a final answer; OpenJudge scores intermediate decisions such as which tool was called and whether the plan was feasible. The README also states that every grader ships with benchmark datasets and pytest integration, with datasets hosted at huggingface.co/datasets/agentscope-ai/OpenJudge. The repository layout supports that claim in structure: there are tests/, pytest.ini and a .pre-commit-config.yaml at the top level.

How grading results become reward signals

The framework's second purpose is turning grades into rewards for fine-tuning and optimization. The README states this as a capability rather than documenting the pipeline in the excerpt, so treat the mechanics as under-specified in the top-level documentation and check the docs site at agentscope-ai.github.io/OpenJudge for the reward path. What is clear from the taxonomy is the design intent: because graders cover trajectory, memory and tool use, not just final answers, the reward signal can be shaped at intermediate steps. That is a different proposition from outcome-only scoring, and it is the main reason to consider this project over a generic LLM-as-judge wrapper.

Installing py-openjudge and running a first grader

The PyPI package name is py-openjudge, not openjudge, so the install command uses the distribution name while the import is likely under the openjudge/ directory. The README does not print a pip line in the excerpt, so install from PyPI using the package name shown in the badge. Python 3.10 or newer is required.

bash
pip install py-openjudge

If you prefer not to install anything, the README points to an online playground at openjudge.me/app where you can test built-in graders and build custom rubrics in the browser. For a local visual interface, the project ships a Streamlit UI; the README gives the run command as:

bash
streamlit run ui/app.py

The repository also includes a docker/ directory and a cookbooks/ directory. The README references cookbooks/skills_evaluation/README.md for the skill-grader examples, which is the most concrete starting point for a first real use. The README does not document the exact Python call signature for invoking a grader in the excerpt, so the cookbooks are the place to copy a working example rather than guessing at import paths.

Rubric generation: zero-shot versus data-driven

Two paths exist when no grader fits. Zero-shot generation takes a task description and optional sample queries, and an LLM produces evaluation rubrics you can use as a grader. That suits prototyping before you have labeled data. Data-driven generation uses a component the README calls GraderGenerator to summarize rubrics from annotated examples and emit an LLM-based grader. The trade-off is straightforward: zero-shot gets you moving but the criteria are unvalidated against your domain, while data-driven output inherits whatever biases exist in your annotations. Neither path removes the need to check that the generated rubric actually separates good outputs from bad ones on held-out cases.

Where OpenJudge is the wrong choice

OpenJudge is a library, not a hosted evaluation service with dashboards, alerting or a scheduler. If your requirement is continuous production monitoring with an operations console, this is not that. The online playground is a separate surface from the pip package, and the README does not describe a self-hosted server mode for the graders. Second, the grader library is heavily LLM-based, so cost and latency scale with evaluation volume in a way that deterministic assertion suites do not. Third, the project's own classifier in pyproject.toml reads Development Status :: 4 - Beta, which is the maintainers' own label. The last push to the default branch was on 2026-09-07, and the most recent tagged release is v0.2.2 from 2026-02-12, so there is a seven-month gap between the newest tag and the newest commit. Verify which graders exist in your installed version rather than trusting the feature list.

Alternatives and the actual difference in approach

The related-search data around this project consistently pairs it with AgentScope, and the repository sits under the agentscope-ai organization, so the natural comparison is a general agent framework's own evaluation utilities. The difference: an agent framework's evaluation helpers are typically scoped to that framework's message and trajectory format, while OpenJudge positions itself as a standalone grader library with a taxonomy spanning text, code, math, multimodal and agent lifecycle. If your agents already run inside one framework and you only need to score that framework's outputs, the built-in helpers are less integration work. If you need the same grader to score outputs from more than one harness, OpenJudge's separation of graders from the runtime is the reason to pick it. The related searches also surface EvalScope and ModelScope-Agent as adjacent tooling; the README does not make a direct comparison, so evaluate those on their own documentation.

Licence, maintenance and upgrade cost

The project is Apache-2.0, declared in pyproject.toml as license = "Apache-2.0". That permits commercial use and modification with the usual notice and patent-grant terms; it is not a copyleft licence, so it does not force you to publish your own evaluation code. This is not legal advice, and if you redistribute the graders or their benchmark datasets, read the licence text and the dataset terms on Hugging Face separately. On upgrade cost: the version is 0.x, the classifier says Beta, and the gap between the v0.2.2 tag and the September 2026 commits means main carries work that is not in a release. If you pin py-openjudge, expect to re-validate custom graders against each minor bump, because grader behaviour is the API surface you depend on. The repository includes .pre-commit-config.yaml, .flake8 and pytest.ini, so contributions are expected to pass lint and tests, which is a signal about internal discipline but says nothing about your integration cost.

Editorial conclusion

Adopt OpenJudge if you already have a Python evaluation loop and need ready-made graders for agent trajectories, tool selection, memory or multimodal coherence, or if you want to convert grading output into reward signals for fine-tuning. Do not adopt it as a hosted evaluation service: the repository is a library plus a Streamlit UI, and the online playground at openjudge.me/app is a separate product surface. Before committing, verify two things against the repository itself: the exact grader names available in your installed version, and whether the graders you need are covered by the benchmark datasets on Hugging Face, since the README describes validation per grader rather than uniform coverage across all 50+.

Frequently asked questions

What is OpenJudge and what does it evaluate?

OpenJudge is an open-source Python evaluation framework for AI applications such as agents and chatbots. It provides built-in graders across general text quality, agent lifecycle behaviour and multimodal output, and can convert grading results into reward signals for fine-tuning.

How do I install OpenJudge?

Install the PyPI distribution named py-openjudge with pip, on Python 3.10 or newer. The README also offers an online playground at openjudge.me/app if you want to try graders without installing anything.

Does OpenJudge have a user interface?

Yes. The project ships a Streamlit-based UI for grader testing and Auto Arena, which the README says can be run locally with streamlit run ui/app.py. There is also a hosted playground at openjudge.me/app.

Can OpenJudge generate graders automatically?

The README describes two paths. Zero-shot rubrics generation takes a task description and optional sample queries and has an LLM produce rubrics. Data-driven generation uses the GraderGenerator to summarize rubrics from annotated data and emit an LLM-based grader.

Official sources

  1. agentscope-ai/OpenJudge on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/agentscope-ai-openjudge.svg)](https://hysenlabs.com/projects/agentscope-ai-openjudge)