# PyOD 3: 61 Detectors, ADEngine Routing, and an Agentic Layer for Python Anomaly Detection

> PyOD is a BSD-2-Clause Python library for outlier detection across tabular, time series, graph, text, image and audio data. Version 3 adds ADEngine orchestration and an od-expert skill for Claude Code and Codex, while the classic fit/predict API stays as it was.

**yzhao062/pyod** — A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.

- Repository: https://github.com/yzhao062/pyod
- Website: https://pyod.dev/
- Stars: 10,020 · Forks: 1,507
- Language: Python
- License: BSD-2-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yzhao062-pyod

## What PyOD solves, and who ends up using it

Anomaly detection in Python has a long tail of incompatible implementations. One paper ships a script that expects pre-normalized features. Another expects a contamination rate. A third returns scores where higher means normal, not anomalous. Assembling a comparison across a handful of methods becomes a data-plumbing project before it becomes a modelling project.

PyOD addresses that by putting detectors behind one interface. The README shows the shape: construct a detector, call fit, read decision_scores_ on the training data, and call decision_function on new data. That contract is the reason the library is widely used, and it is the part the project says stays unchanged in version 3.

The audience is narrower than the tagline suggests. If you have a labelled dataset and a classifier that already works, you do not need this. PyOD fits the case where labels are absent or extremely sparse: fraud review queues, sensor fault screening, novelty detection in embeddings, and research comparisons where several detectors must run over the same matrix. The pyproject classifiers list financial and insurance, scientific research, and information technology as intended audiences, which matches that reading.

## Three layers: classic API, ADEngine, and agentic investigation

The README describes three entry points rather than one. Layer 1 is the classic API: you choose the detector and call fit and predict yourself. Layer 2 is ADEngine, which the project describes as choosing, comparing, and assessing detectors automatically. Layer 3 routes a natural-language request through the od-expert skill or the MCP server into an ADEngine workflow.

The important detail is that layers 2 and 3 are not separate products. Both are described as powered by ADEngine, the lifecycle orchestration core, and the pure-Python path to it is an import: from pyod.utils.ad_engine import ADEngine. The agentic layer is a front end over the same machinery, which means you can adopt ADEngine without installing any agent tooling.

The MCP server is the more conventional integration point. It is started with python -m pyod.mcp_server and exposes ten stateless tools, split across knowledge queries (list_detectors, explain_detector, compare_detectors, get_benchmarks) and planning calls (profile_data, plan_detection, and others the truncated README cuts off). Stateless matters here: each call stands alone, so the agent holds the conversation state, not the server. That keeps the server simple but pushes the burden of remembering which dataset was profiled onto the caller.

## Installing PyOD and running a first detector

The core library is a single pip install and is required for every activation path. The README gives no virtual environment step and no version pin, so the command is the whole story.

```bash
pip install pyod
```

After that, the README's five-line example is the fastest way to confirm the install works. It uses IForest, the isolation forest detector, and the two attributes are worth reading carefully: decision_scores_ holds training scores, and decision_function returns scores for new rows.

```python
from pyod.models.iforest import IForest
clf = IForest()
clf.fit(X_train)
y_train_scores = clf.decision_scores_
y_test_scores = clf.decision_function(X_test)
```

The README does not state the orientation of those scores, so check the detector's documentation before treating a high value as anomalous. If you want the library to pick the detector instead, the pure-Python path skips the agent tooling entirely.

```python
from pyod.utils.ad_engine import ADEngine
```

Before going further, run the built-in diagnostic. The README says it reports version, detector counts, and the install state of each activation path, and that it detects a Claude Code or Codex install and recommends the matching command.

```bash
pyod info
```

If you do want the agent path, the commands differ by tool. Claude Code installs user-globally, while Codex is project-local because it has no user-global directory, which the README states explicitly.

```bash
pyod install skill
pyod install skill --project
```

The MCP route needs an optional extra before the server will start.

```bash
pip install pyod[mcp]
pyod mcp serve
```

## Where PyOD is the wrong tool

The dependency list is the first constraint. requirements.txt pins joblib, matplotlib, numpy, numba, scipy and scikit-learn. Numba is not optional in that list, and it is the kind of dependency that constrains which Python and NumPy builds you can combine. In a locked environment, adding PyOD can force a wider upgrade than the detector itself justifies.

Deep detectors sit behind extras. The pyproject defines torch, suod, xgboost, combo, pythresh, embedding, openai and huggingface as optional dependency groups, so a plain pip install pyod does not give you the deep-learning or embedding-backed detectors. That is sensible packaging, but it means the detector count in the headline is not the count available after the base install.

Streaming is the harder mismatch. The API is fit-then-score, and the README's examples are batch shaped. Nothing in the README describes an incremental update path for a detector that has already been fitted, so a pipeline that must score an unbounded stream without refitting is not what this library is built for. Refitting per window is possible, but it is your loop, not a library feature.

The agentic layer is also new ground. The README states that Claude Code and Codex can drive ADEngine through the od-expert skill, and that MCP-compatible agents can query detector knowledge. It does not describe what happens when a plan is wrong, how a failed investigation is surfaced, or what the skill does with ambiguous requests. Treat the agent layer as a convenience over ADEngine, not as a substitute for knowing which detector family suits your data.

## PyOD against scikit-learn's own outlier detectors

The obvious alternative is scikit-learn, which already ships IsolationForest, LocalOutlierFactor, OneClassSVM and EllipticEnvelope, and which PyOD depends on anyway. The difference is breadth and framing. Scikit-learn offers a small set of general-purpose detectors inside a general-purpose machine learning library; PyOD offers 61 detectors, per the pyproject description, organized around anomaly detection as the primary problem, including graph, text, image and audio modalities that scikit-learn does not target.

The second difference is routing. Scikit-learn leaves model selection to you. PyOD's ADEngine is explicitly positioned as choosing, comparing and assessing detectors automatically, and the README ties that routing to benchmark results from ADBench, TSB-AD, BOND and NLP-ADBench. Whether automated routing beats a well-chosen isolation forest on your data is an empirical question the README does not answer for you.

A third option is PyOD's own ecosystem extras. The suod, combo and pythresh groups point at separate packages for acceleration, combination and thresholding, which suggests that if you only need one detector with a tuned threshold, the smaller packages may be a lighter fit than the full library.

## Maintenance, licensing, and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-08. Releases have been frequent: v3.6.5 on 2026-08-17, v3.6.4 on 2026-08-02, and v3.6.3 on 2026-08-01. Patch releases at that cadence usually mean bug fixes rather than interface churn, but the README does document one compatibility detail: the legacy pyod-install-skill command from v3.0.0 still works as an alias for pyod install skill. That is a deliberate concession to existing scripts, and it is the kind of thing to check before pinning a version.

The licence is BSD-2-Clause, declared both in pyproject and in the LICENSE file at the repository root. That is a permissive licence with minimal conditions, but it is not the same as the BSD-3-Clause variant that adds a non-endorsement clause. If your legal review assumes three-clause BSD by default, the two-clause text is what applies here. This is a description of the licence identifier, not legal advice.

Upgrade cost is mostly dependency cost. requires-python is >=3.9 and the classifiers run through 3.13, so the supported interpreter range is wide. The optional groups are where version pressure accumulates: torch>=2.0, sentence-transformers>=5.0.0 and openai>=1.0 each carry their own upgrade cadence, and those are the pins most likely to move under you.

## Conclusion

Adopt PyOD if you need many detectors behind one fit/predict API and want benchmark-backed routing instead of hand-picked models; the classic layer carries no agent dependency. Skip it if your data is a streaming, unbounded signal where per-batch refitting is the wrong shape, or if you need a maintained graphical interface. Before committing, run pyod info to confirm the version and which activation paths are installed, and check the pyod install skill target directory, because Codex has no user-global skills directory.

## FAQ

### What are the three types of anomaly detection?

PyOD's documentation and keywords frame detection around supervised, unsupervised and semi-supervised settings, alongside related concepts such as novelty detection and out-of-distribution detection. The library's core examples are unsupervised, since the classic API fits on X_train without labels.

### What is anomaly detection used for?

The pyproject classifiers list financial and insurance, scientific research, information technology, and education as intended audiences, and the keywords include fraud detection. PyOD itself covers tabular, time series, graph, text, image and audio data.

### How can I detect anomalies in Python?

Install the library with pip install pyod, then construct a detector such as IForest, call fit on your training data, and read decision_scores_ or call decision_function on new rows. The README presents that as the five-line starting point.

### Which model is best for anomaly detection?

PyOD does not name a single best detector. The README positions ADEngine as choosing, comparing and assessing detectors automatically, and the project ties its routing to benchmark results from ADBench, TSB-AD, BOND and NLP-ADBench. The right choice depends on your data modality and whether you have labels.

## Sources

- [License: BSD-2-Clause](https://github.com/yzhao062/pyod/blob/master/LICENSE)
- [Project website](https://pyod.dev/)
- [README](https://github.com/yzhao062/pyod/blob/master/README.md)
- [Releases](https://github.com/yzhao062/pyod/releases)
- [yzhao062/pyod on GitHub](https://github.com/yzhao062/pyod)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yzhao062-pyod
