# Tencent AICGSecEval (A.S.E): a repository-level benchmark for AI-generated code security

> A.S.E evaluates the security of code that LLMs and coding agents produce inside a real project context, using static and dynamic analysis. It is heavy, Docker-backed and aimed at teams that already have a model endpoint and a machine to spare.

**Tencent/AICGSecEval** — A.S.E (AICGSecEval) is a repository-level AI-generated code security evaluation benchmark developed by Tencent Wukong Code Security Team.

- Repository: https://github.com/Tencent/AICGSecEval
- Website: https://aicgseceval.tencent.com
- Stars: 662 · Forks: 113
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tencent-aicgseceval

## The problem A.S.E targets: generated code that compiles but is not safe

Most coding benchmarks stop at correctness. A.S.E (AICGSecEval) asks a different question: when a model writes code inside an existing repository, does it introduce a known vulnerability class? The README frames the benchmark as simulating real development workflows, with tasks derived from real GitHub projects and authoritative CVE patches. That derivation matters. A task is not a blank function signature; it carries the surrounding project so the model has to work with existing imports, call sites and conventions. The audience is narrow and specific: model vendors, security researchers and platform teams who need a repeatable number for AI-assisted programming rather than an anecdote about one bad completion. The repository is published by the Tencent Wukong Code Security Team and is not archived; the last push was on 2026-05-25.

## Dataset composition: 29 CWE types and the OWASP and CWE lists

The 2.0 release notes describe a dataset expansion to key risks from the OWASP Top 10 and CWE Top 25, covering 29 CWE vulnerability types across C/C++, PHP, Java, Python and JavaScript. The data lives in the data/ directory, and the quick start points at ./data/data_v2.json. Two validation scripts sit at the repository root, validateV1Data.py and validateV2Data.py, which suggests the dataset schema changed between the v1 and v2 generations. If you plan to extend the benchmark with your own cases, start from those validators rather than from the JSON alone, because they encode what the harness expects. The language spread is also a constraint worth noting: Go, Rust, TypeScript and C# are not in the list given by the release notes, so a team whose stack is mostly Go will be measuring a smaller slice of its real surface.

## How the evaluation pipeline works, from retrieval to scan

The top-level layout shows the pipeline as separate stages rather than one monolith. run_data_retrieval_bm25.py and run_data_retrieval_claude_code.py handle context retrieval, run_code_generation_llm.py and run_code_generation_agent.py produce the code, run_security_scan.py and run_security_scan_static.py scan it, and run_evaluate.py aggregates. invoke.py is the launcher that ties these together. The retrieval step is the part that makes this repository-level rather than snippet-level: the README states that the framework automatically extracts project-level code context to simulate realistic AI programming scenarios. BM25 appears as a retrieval method, and pyserini==0.44.0 in requirements.txt is consistent with a lexical retrieval index. The scan stage is hybrid by design. The 2.0 notes describe a dynamic scheme based on test cases and vulnerability PoCs layered on static analysis, with the stated goal of balancing detection breadth against verification precision. That is a real trade-off being managed, not a slogan: static analysis finds more candidates and produces more false positives, while a PoC that actually triggers the flaw is harder to argue with but only exists for vulnerability classes someone wrote a trigger for.

## Installing A.S.E and running a first LLM evaluation

The README states system requirements of at least 16GB of memory, 100GB of disk, Python 3.11 or newer, and Docker 27 or newer. The disk figure is the one people underestimate, since Docker images, cloned repositories and generated outputs all land on the same volume. Dependencies install from the pinned requirements file, which includes openai, gitpython, docker, transformers and pyserini.

```bash
pip install -r requirements.txt
```

The launcher takes either --llm or --agent as the mode selector. Running it with -h prints the available options, and the README notes that unrecognized arguments are forwarded to the agent module, which is how agent-specific flags get through.

```bash
python3 invoke.py -h
```

A basic LLM run needs a model name, a base URL, an API key, a batch identifier, the dataset path and an output directory. The README gives this shape, including a github_token that falls back to anonymous cloning when omitted.

```bash
python3 invoke.py \
  --llm \
  --model_name gpt-4o-2024-11-20 \
  --base_url https://api.openai.com/v1/ \
  --api_key sk-xxxxxx \
  --batch_id v1.0 \
  --dataset_path ./data/data_v2.json \
  --output_dir ./outputs
```

The README warns that a full run can take a long time and that --max_workers controls concurrency. It also states that automatic checkpoint recovery is supported, so an interrupted run can be restarted with the same command and resume from the last checkpoint. Expect the first run to be dominated by repository cloning and Docker image pulls rather than by model latency.

## Evaluating coding agents, not just raw model completions

The 2.0 notes list support for agentic programming tools as an evaluation target upgrade. The agent path is selected with --agent and --agent_name, and the README walks through Claude Code as an example, passing a Claude API URL, key and model name. The important design detail is the argument forwarding: the launcher does not need to know every agent flag, so agent modules parse their own options. That keeps the launcher stable while agents change, but it also means the option surface is not fully discoverable from invoke.py -h alone. If you add an agent, you are responsible for documenting its flags somewhere the launcher does not cover. The README also cautions that different agents may require distinct configurations, including model parameters, credentials or APIs, which is an honest admission that agent evaluation is less standardized than LLM evaluation here.

## Where A.S.E is the wrong tool

Three cases stand out. First, if your question is whether a single pull request is safe, A.S.E is the wrong instrument: it is a benchmark that clones repositories, generates code and scans it, and the resource floor of 16GB memory, 100GB disk and Docker 27 reflects that. Second, if you need coverage of a language outside C/C++, PHP, Java, Python and JavaScript, the 2.0 release notes do not claim it. Third, if you have no model endpoint or agent credentials, there is nothing to measure, because the benchmark generates code rather than analysing code you already have. There is also a maintenance dimension: the last push was on 2026-05-25, and the README does not document a rollback procedure for a partially completed run beyond rerunning the command for checkpoint recovery. Treat the harness as a research artifact you pin to a specific release, not as a service.

## Alternatives and how their approach differs

The obvious alternative is running a static analyser such as Semgrep or CodeQL directly over model output. The difference is scope and intent. A scanner answers "what patterns are in this file"; A.S.E answers "did this model, given this repository context, produce code that a PoC can trigger". The second question requires the retrieval step, the generation step and the dynamic verification step, which is why it needs Docker and a large disk. The cost of that extra machinery is speed and setup burden. A second alternative is a correctness-oriented benchmark such as HumanEval-style suites; those are cheap to run and tell you nothing about CWE coverage. A third is building an internal harness around your own CVE patches. That gives you stack-specific relevance, but you inherit the work of writing validators, PoC triggers and aggregation, which is precisely what run_evaluate.py and the validateV2Data.py scripts already do here. The related searches for this project point at AI-Infra-Guard and the Zhuque AI Detection Assistant, which are Tencent security tooling in adjacent territory, not drop-in substitutes for a repository-level generation benchmark.

## Licence, upgrade cost and what to check before pinning

The repository carries a LICENSE file at the root, but the GitHub licence field reports NOASSERTION, meaning the platform could not map the file to a recognized identifier. Read LICENSE directly before any redistribution or commercial use; nothing in the README substitutes for it, and this is not legal advice. On upgrade cost, the release history shows v2.0.0 in November 2025, v2.0.1 in December 2025, and a separate report-v1.0 tag for the AI-generated code security risk report. The dataset schema moved between v1 and v2, which is why two validators exist. If you have stored results keyed to a batch_id, a dataset change will invalidate comparisons across that boundary. Pinning requirements.txt matters too: it fixes pyserini, openai, transformers and docker at exact versions, so an unpinned environment can drift away from what the harness was written against. The README states the project intends to keep evolving through issues and pull requests, and the last push on 2026-05-25 is consistent with that, but there is no stated release cadence to plan around.

## Conclusion

Adopt A.S.E if you evaluate coding models or agents and can give it a 16GB machine, 100GB of disk, Docker 27, Python 3.11 and an API key, because the static plus dynamic split is what separates it from single-tool scanners. Do not adopt it if you want a fast lint pass over one file, or if you cannot supply a model endpoint, since the benchmark only measures generated code. Before committing, read LICENSE, run python3 invoke.py -h to see the real option surface, and confirm the dataset path and batch_id you intend to use.

## FAQ

### What is Tencent AICGSecEval (A.S.E)?

It is a repository-level benchmark from the Tencent Wukong Code Security Team for evaluating the security of AI-generated code. It builds generation tasks from real GitHub projects and CVE patches, then assesses the output with static and dynamic analysis.

### What are the system requirements for running A.S.E?

The README lists at least 16GB of memory, 100GB of disk, Python 3.11 or newer, and Docker 27 or newer. A full evaluation can take a long time, and --max_workers controls concurrency.

### How do I install and start an evaluation with A.S.E?

Install dependencies with pip install -r requirements.txt, then run python3 invoke.py with either --llm or --agent. An LLM run needs --model_name, --base_url, --api_key, --batch_id, --dataset_path and --output_dir.

### Which vulnerability types and languages does A.S.E cover?

The 2.0 release notes state coverage of key risks from the OWASP Top 10 and CWE Top 25, spanning 29 CWE vulnerability types across C/C++, PHP, Java, Python and JavaScript.

### Does A.S.E support evaluating coding agents?

Yes. The 2.0 release notes list support for agentic programming tools, selected with --agent and --agent_name. The README shows Claude Code as an example, and unrecognized arguments are forwarded to the agent module for parsing.

## Sources

- [Issues](https://github.com/Tencent/AICGSecEval/issues)
- [Project website](https://aicgseceval.tencent.com)
- [README](https://github.com/Tencent/AICGSecEval/blob/master/README.md)
- [Releases](https://github.com/Tencent/AICGSecEval/releases)
- [Tencent/AICGSecEval on GitHub](https://github.com/Tencent/AICGSecEval)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tencent-aicgseceval
