# Dingo: rule, LLM and agent checks for AI training data and RAG output

> MigoXLab's dingo-python evaluates datasets, model responses and RAG pipelines with built-in rules, LLM judges and optional hallucination models. It is strongest when you need field-level quality gates, and weakest when you expect the README to tell you how to roll a bad batch back.

**MigoXLab/dingo** — Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool

- Repository: https://github.com/MigoXLab/dingo
- Website: https://dingo.openxlab.org.cn/
- Stars: 757 · Forks: 77
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/migoxlab-dingo

## What Dingo evaluates, and who ends up using it

Dingo is positioned in its README as a comprehensive AI data, model and application quality evaluation tool for ML practitioners, data engineers and AI researchers. The scope it claims is wider than a single dataset linter: pre-training data, fine-tuning datasets and production AI systems, including RAG pipelines, are all listed as targets.

The concrete problem is that quality rules differ per column. A table with an isbn field and a title field should not be checked the same way. The README calls this multi-field evaluation and gives exactly that example, ISBN validation for isbn and text quality for title, applied in parallel. That is the design centre of the project. If your data is one big blob of text, you get less out of it than if your data is structured records where each field has its own notion of correct.

The second problem it targets is judge selection. Heuristic rules are cheap and deterministic; LLM judges catch things rules cannot express. Dingo ships both and lets you combine them, which is the honest answer to the fact that neither alone is sufficient. The topics list on the repository also points at hallucination detection, which is a different task from data cleaning: there you are scoring a generated answer against evidence, not filtering rows.

## How the rule, LLM and agent layers fit together

The repository layout separates model code into a dingo/ package with model/rule and model/llm subpackages. The README example imports RuleSpecialCharacter from dingo.model.rule.rule_common and LLMTextQualityV4 from dingo.model.llm.text_quality.llm_text_quality_v4, which confirms that rules and LLM evaluators are distinct classes you select explicitly rather than one pipeline that decides for you.

Data enters through a Data object from dingo.io.input. In the README example it is constructed with data_id, prompt and content. So the unit of evaluation is a record with named fields, and evaluators read those fields. Configuration for the LLM path comes from EvaluatorLLMArgs in dingo.config.input_args, which is the object you populate with model and connection settings before calling the evaluator.

On the application side, the README states that RAG system assessment covers retrieval and generation quality with five academic-backed metrics, and that a retrieval extra exists for benchmark evaluation via MTEB and pytrec-eval-terrier. That is a separate dependency path from the core install, which tells you retrieval scoring is not part of the default footprint.

Execution is described as local for iteration or Spark for billion-scale datasets. The README does not document how the Spark path is configured, so treat that as something to confirm in the code before you plan around it. The same caution applies to the agent layer: the topics list agent-as-a-judge and the setup.py exposes an agent extra, but the README excerpt does not show an agent example.

## Install Dingo from PyPI and run one rule plus one LLM check

Installation is a plain pip install of the dingo-python distribution. The README lists four variants, and the difference is which optional dependencies come with it. The core package includes rule evaluation, LLM evaluation, the MCP server and datasource support.

```bash
pip install dingo-python
```

If you want the hallucination detection model, which the README says requires transformers and torch, install the hhem extra instead:

```bash
pip install "dingo-python[hhem]"
```

There is also a retrieval extra for MTEB and pytrec-eval-terrier, and an all extra that the README describes as HHEM plus Agent plus Retrieval. The setup.py adds two more extras that the README excerpt does not advertise: litellm, pinned as litellm>=1.80.0,<1.87.0, and lmdeploy, which setup.py keeps isolated on purpose because it requires transformers>=4.56 and pulls in a large dependency set. If you plan to use lmdeploy as an inference backend, that is a separate install step, not part of all.

A first real use is the chat data example from the README. It builds a Data record and selects an LLM evaluator class. Note the content string, which deliberately contains a replacement character and a stray caret, the kind of noise you would want flagged.

```python
from dingo.config.input_args import EvaluatorLLMArgs
from dingo.io.input import Data
from dingo.model.llm.text_quality.llm_text_quality_v4 import LLMTextQualityV4
from dingo.model.rule.rule_common import RuleSpecialCharacter

data = Data(
    data_id='123',
    prompt="hello, introduce the world",
    content="\ufffdI am 8 years old. ^I love apple because:"
)
```

After defining the record, the README example continues into an llm() function that calls LLMTextQualityV4.dynamic_config, and the excerpt is truncated at that point. What you should expect from a run is a quality report; the README describes rich reporting with GUI visualization and field-level insights. The precise report schema is not in the excerpt, so read the output of your first run before wiring it into anything downstream.

## Where Dingo gets in your way

The README is a product page more than a manual. It tells you the extras exist, but not what happens when an LLM judge call fails mid-dataset, whether partial results are written, or how to resume. There is no documented rollback story for a batch you have already evaluated and acted on. For a tool whose output decides which training rows survive, that is a real gap, and it is the first thing I would look for in the code.

The evaluation cost is the second constraint. Rules are cheap. LLM judges are not, and the README gives no guidance on batching, concurrency or rate limiting for the LLM path. If you point LLMTextQualityV4 at a large dataset without thinking about throughput, you are paying per record with no documented throttle.

The dependency story is genuinely awkward. setup.py isolates lmdeploy because it hard-requires transformers>=4.56 and drags in heavy packages. The hhem extra needs transformers and torch. The retrieval extra needs MTEB and pytrec-eval-terrier. Installing all gets you everything at once, which is convenient until two of those constraints collide in your existing environment. A comment in setup.py notes that the hallucination model moved from Vectara HHEM to a standard T5 MiniCheck, which removed a transformers<4.49 ceiling and the conflict with lmdeploy. That history is a useful signal: heavy optional dependencies have caused version friction here before.

Finally, if you want a web UI, access control with JWT or Google OAuth, or a RESTful API, the README sends you to the separate SaaS edition and an application form with a stated 1 to 5 business day review. The open source package is the evaluation engine, not the platform.

## OpenCompass and Dingo solve different halves of the problem

OpenCompass is the natural comparison, since dingo's topics list includes it. The difference is what is being measured. OpenCompass evaluates models against benchmark suites: you have a model, you have standard tasks, and you want a comparable score. Dingo evaluates the data and the application: you have a dataset or a RAG system, and you want to know whether the inputs and outputs meet quality rules you define.

That distinction matters for adoption. If your question is which of two models is better on reasoning benchmarks, Dingo is the wrong tool and its RAG metrics will not answer it. If your question is whether the corpus you are about to fine-tune on contains the kind of junk your rules describe, or whether your retrieval plus generation stack is returning grounded answers, benchmark harnesses do not address that at all.

The overlap is the LLM-as-a-judge machinery, which both categories of tool now use. Dingo's version is configurable per field and combinable with deterministic rules, which is the part that distinguishes it from a benchmark runner that treats every sample identically.

## Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-11, so the project is being worked on. Release cadence visible in the release list is roughly quarterly: v2.3.0 and v2.4.0 both landed on 2026-05-29, and v2.5.0 on 2026-08-04. The setup.py version string matches 2.5.0, so the packaged version tracks the release tag.

Upgrade cost is dominated by the optional extras, not by Dingo's own code. The litellm extra is pinned to a range, litellm>=1.80.0,<1.87.0, which means a future litellm release outside that window will not be picked up until Dingo widens the pin. The lmdeploy extra is deliberately excluded from all, so a version bump there cannot break a default install, but it also means you carry that dependency decision yourself. The setup.py comment about the HHEM to MiniCheck switch shows what an upgrade can cost you: a model swap changes what your hallucination scores mean, even when the API surface stays the same.

The licence is Apache-2.0, which is permissive and includes a patent grant. That is the licence of the open source package; the SaaS edition is a separate product with its own access process, and the README does not state its terms. If you are deploying Dingo inside a commercial product, read the LICENSE file in the repository rather than relying on the badge, and check whether any optional dependency you enable carries a different licence. This is not legal advice.

## Conclusion

Adopt Dingo if you already have a Python data pipeline and want rule and LLM checks applied per field before training or serving, and if a JSON quality report per run fits your workflow. Do not adopt it if you need a hosted dashboard, role-based access or a REST API out of the box, since the README puts those features in the separate SaaS edition. Before committing, verify two things in your own environment: which extras you actually need (hhem, retrieval, agent, litellm, lmdeploy or all), because each pulls a different dependency set, and whether the report format produced by your configured evaluator is the one your downstream tooling can parse.

## FAQ

### How do I install Dingo?

Install the core package with pip install dingo-python. The README also lists dingo-python[hhem] for hallucination detection, dingo-python[retrieval] for MTEB and pytrec-eval-terrier, and dingo-python[all] for HHEM plus Agent plus Retrieval.

### What is Dingo used for?

The README describes it as a comprehensive AI data, model and application quality evaluation tool for training data, fine-tuning datasets and production AI systems. It combines built-in heuristic rules with LLM-based assessment and RAG metrics.

### Does Dingo require an LLM to run?

No. The core package includes rule evaluation alongside LLM evaluation, and the README example imports a rule class, RuleSpecialCharacter, next to the LLM evaluator. Rules are the heuristic half of the hybrid approach, so a rules-only run does not need a model endpoint.

### Can Dingo evaluate retrieval quality in a RAG system?

Yes. The README states that RAG system assessment covers retrieval and generation quality with five academic-backed metrics, and a separate retrieval extra installs MTEB and pytrec-eval-terrier for benchmark evaluation.

### Do I need the SaaS version to get a web UI?

According to the README, the Web UI, JWT and Google OAuth access control, visual reports and RESTful API are features of the separate SaaS edition, obtained through an application form with a stated 1 to 5 business day review. The open source package is installed from PyPI.

## Sources

- [License: Apache-2.0](https://github.com/MigoXLab/dingo/blob/main/LICENSE)
- [MigoXLab/dingo on GitHub](https://github.com/MigoXLab/dingo)
- [Project website](https://dingo.openxlab.org.cn/)
- [README](https://github.com/MigoXLab/dingo/blob/main/README.md)
- [Releases](https://github.com/MigoXLab/dingo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/migoxlab-dingo
