# Nine new AgentDoG checkpoints and a search prefix that no longer matches them

> AgentDoG 1.5 ships nine classifiers across three task labels, a runtime guardrail for OpenClaw agents and four arXiv reports, but the page describing all of it stops partway through the safety taxonomy and gives no install path.

**AI45Lab/AgentDoG** — A Diagnostic Guardrail Framework for AI Agent Safety and Security

- Repository: https://github.com/AI45Lab/AgentDoG
- Stars: 700 · Forks: 35
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ai45lab-agentdog

## The documented search prefix misses every 1.5 checkpoint

The opening instruction says to search checkpoints with names starting with `AgentDoG-`. Every 1.0 row in the two model tables matches that shape: `AgentDoG-Qwen3-4B`, `AgentDoG-FG-Llama3.1-8B`, `AgentDoG-Qwen2.5-7B`, and the rest put a hyphen right after the family name. None of the nine 1.5 rows do. They are written `AgentDoG1.5-...`, with the version number attached straight to the family name and no separator, so the single line the page offers for locating weights returns nothing for the current generation. A reader who follows it lands on the six older checkpoints or on an empty result. No `huggingface-cli` command appears anywhere, no organisation path is spelled out as a prerequisite, and the two hosts are not kept in step: `AI45Research` on Hugging Face and `Shanghai_AI_Laboratory` on ModelScope carry different link targets for the same row. The two collection links in the header are the only pointers that still carry a version marker, one named for 1.5 and one with no version at all, so the paths that do work are the ones the search instruction does not describe.

## Three task labels, and only one of them is a single model

The 1.5 table has a Task column with three values. Unified safety diagnosis appears exactly once, on `AgentDoG1.5-Unified-Qwen3.5-4B`, a 4B model built on Qwen3.5-4B. Coarse-grained moderation appears four times, spanning 0.8B, 2B and 4B on Qwen3.5 and 8B on Llama3.1. Fine-grained diagnosis appears four times with the same size spread and the same four base models. The `FG-` fragment in the checkpoint name is the only thing separating the fine-grained rows from the coarse ones, so the naming convention carries the taxonomy that the Task column also carries. The 1.0 tables have no Unified row at all and nothing below 4B: the smallest 1.0 weight is 4B and the largest is 8B. The lightest pair in the entire release is therefore the two 0.8B checkpoints, and both of them exist only in 1.5. The 1.0 tables name their base models explicitly, Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct and Llama3.1-8B-Instruct, while the 1.5 table gives bare family names such as Qwen3.5-4B with no instruct or revision suffix, so which exact upstream checkpoint a 1.5 row started from has to be settled by the reader.

## Two Hugging Face slugs double the L in Llama

In the 1.5 table, the coarse-grained 8B row points at `huggingface.co/AI45Research/AgentDoG1.5-LLamma-3.1-8B`, and the fine-grained 8B row points at `AgentDoG1.5-FG-LLamma-3.1-8B`. The display name in the very same row reads Llama3.1 with two Ls. Repository identifiers are compared as strings, so copying the name printed in the Model column does not land on the address printed beside it. Two more rows mix cases: `AgentDoG1.5-Qwen3.5-0.8b` and `AgentDoG1.5-FG-Qwen3.5-0.8b` end in a lowercase b, while their ModelScope links in the same rows end in an uppercase B. The 1.0 table leans the other way, because all six of its ModelScope cells are the same collection page rather than a per-model address, so a reader who changes hosting between the two versions also changes how much precision the links give them. Both 8B rows in 1.5 sit on Llama3.1 as well, which makes the doubled letter the only place in either table where a link target and its own label disagree outright.

## The safety taxonomy heading opens and stops at the letter A

The last heading on the page is `Safety Taxonomy`. Underneath it there is a single character, `A`, and nothing more. Everything substantive about what the labels mean, how many risk types the revised taxonomy carries, and how the added Codex and OpenClaw risk categories are defined is absent. The introduction says the three-dimensional taxonomy was revised and supplemented, the news entries announce two benchmark releases beyond the original ATBench, and four links point at arXiv PDFs, yet none of that content is restated on the page. The one word that does appear under the heading is enough to confirm the section was meant to carry a lettered enumeration, because the rest of the page numbers its sections only by version. If you need the label set that `AgentDoG1.5-Unified-Qwen3.5-4B` was trained to emit, this file will not hand it to you. `docs/`, `figures/` and `prompts/` sit in the repository root and are not mentioned outside the file list, so whether the taxonomy is written down in the repository at all is something a reader has to go check.

## The runtime guardrail has no configuration surface described

One of the four introduction bullets claims a practical runtime guardrail built on AgentDoG 1.5 for real-world OpenClaw agent deployment, with online safety monitoring and intervention in deployed agentic workflows. Nothing after that bullet describes it. There is no policy file, no threshold, no port, no latency figure, no statement of what happens when a trajectory is judged unsafe, and no word on whether the monitor can stop a tool call or only annotate the run. `Online Agentic Guardrail/` sits at the repository root and the page never names a file inside it. The same bullet set puts a number on the training side, around 1k samples, and claims performance comparable to frontier open and closed models, but no results table exists on the page to hold either claim up. A third bullet claims a dedicated agentic training environment that keeps a standard 8-core machine above 10,000 concurrent agentic environments, which is a throughput figure for a training setup that is never shown, named or configured. Neither the 1.0 page nor the 1.5 page says how the two runtime pieces, the guardrail and the training environment, share a codebase with the weight files they produce.

## The single runnable example calls someone else's moderation API

`examples/` contains three markdown walkthroughs, `getting_started_v1.md`, `getting_started_v1_5.md` and `readme_v1.md`, one sample input named `trajectory_sample.json`, one Python file, and a committed `__pycache__/`. The Python file is `run_openai_moderation.py`. Its name points at a third-party moderation endpoint rather than at any AgentDoG checkpoint, which is an odd single example to ship in a repository whose only artifacts are its own classifiers. No command to run it appears on the page, no dependency list sits next to it, and the sample trajectory is never described field by field, so the input contract for the models can only be read by opening that file. The compiled cache directory is committed rather than ignored. The walkthrough names also drift: three files exist for a project with two released generations, and the newest one carries a 1_5 suffix while the version elsewhere in the repository is always written as a decimal point.

## Three top-level directories carry spaces, and a .gitmodules names nothing

The repository root lists `.DS_Store`, `.gitmodules`, `Agentic Safety Training/`, `AgenticXAI/`, `Online Agentic Guardrail/`, `README.md`, `colleague-skill`, `docs/`, `examples/`, `figures/` and `prompts/`. Three of those names contain spaces, so every path into them has to be quoted, and not one of the three is explained in the prose. `colleague-skill` is a bare directory with no description and no mention outside the file list, and `AgenticXAI/` is never named in the text either. `.gitmodules` says submodules are declared, but the page does not say which repositories they point at, which branch each tracks, or whether an ordinary shallow clone leaves those directories empty. `.DS_Store` is committed at the root, which is the file macOS Finder uses to store folder view settings. Between them the eleven entries describe four code trees, one guardrail, one skill folder and three documentation directories, and the page that lists them explains none of the eleven.

## Four papers, no releases, and no license anywhere in the tree

The repository has no GitHub releases at all, so there is nothing to pin, no tag to check out and no package to install. The deliverables are weight files on Hugging Face and ModelScope, reached through a project page hosted at `ai45lab.github.io/AgentDoG/`. The root carries no LICENSE file and the GitHub license field is empty, so no terms are stated for the code or for the trained weights, and a reader cannot settle that question from the paper links either. The newest commit on the default branch is dated 2026-06-08. Two issues are open against 699 stars and 35 forks. The four news entries run from 2026/01/26 to 2026/05/28 and cover two model generations plus two benchmark releases, while the documentation on the page itself stops at the model tables. The gaps between those entries are uneven as well, four months from 1.0 to 1.5, which is worth knowing before anyone plans a migration path between the two generations, since nothing on the page maps a 1.0 label onto a 1.5 one.

## Conclusion

What this repository hands you is weight files and four papers, not a running safety layer. Anyone planning to put AgentDoG 1.5 in front of a live agent has to source the checkpoints from the hosting site themselves, work out the trajectory input shape from a sample file, and decide the intervention policy on their own, because none of those are settled here. As a reference set for trajectory-level safety labels the 1.5 tables are cheap to try at 0.8B. Before relying on any of it, confirm the exact repository slug you actually resolved, since the address the table gives for the 8B models does not match the name printed beside it, and confirm the terms yourself, because neither the repository root nor its GitHub license field states them.

## FAQ

### What does AgentDoG judge in an agent run?

It classifies whole trajectories rather than single prompts. The 1.5 model table splits the task into unified safety diagnosis, coarse-grained moderation and fine-grained diagnosis, with the fine-grained checkpoints marked by an FG- prefix in their names.

### What is the smallest AgentDoG model and what is it built on?

Two 0.8B checkpoints exist: AgentDoG1.5-Qwen3.5-0.8B for coarse-grained moderation and AgentDoG1.5-FG-Qwen3.5-0.8B for fine-grained diagnosis, both fine-tuned from Qwen3.5-0.8B. The 1.0 tables stop at 4B.

### How do I install AgentDoG?

The repository gives no install command, no requirements file and no Python version requirement. It points at the Hugging Face and ModelScope organizations and says checkpoints are the things to search for, under names starting with AgentDoG-.

### Is AgentDoG released under a license?

No LICENSE file appears in the repository root listing and the GitHub license field is empty, so terms for the code and the weights are not stated there. What is linked instead are two technical reports and two benchmark papers on arXiv.

### Can the AgentDoG guardrail block a running agent?

The introduction claims online safety monitoring and intervention for OpenClaw deployments, but it does not say whether the guardrail can halt a tool call or only label a trajectory. The Online Agentic Guardrail directory is listed at the root with no description of its contents.

## Sources

- [AI45Lab/AgentDoG on GitHub](https://github.com/AI45Lab/AgentDoG)
- [Issues](https://github.com/AI45Lab/AgentDoG/issues)
- [README](https://github.com/AI45Lab/AgentDoG/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ai45lab-agentdog
