# InferenceX: SemiAnalysis's Continuous Inference Benchmark Platform

> InferenceX is an Apache-2.0 Python benchmark harness from SemiAnalysis that re-runs inference frameworks against a fixed hardware fleet so results do not go stale between releases. This article covers what it measures, how to run it, and where it stops being the right tool.

**SemiAnalysisAI/InferenceX** — Open Source Continuous Inference Benchmark Research Platform, Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon TPUv6e/v7/Trainium2/3 | , Kimi K2.7-Code MiniMax M3 DeepSeekv4 GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 TPUv6e/v7/Trainium2/3.

- Repository: https://github.com/SemiAnalysisAI/InferenceX
- Website: https://inferencex.com/
- Stars: 1,763 · Forks: 305
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/semianalysisai-inferencex

## What InferenceX measures that a one-off benchmark cannot

The problem InferenceX targets is stated plainly in its README: a benchmark run at a fixed point in time goes stale, because inference frameworks such as SGLang, vLLM, TensorRT-LLM, CUDA and ROCm ship kernel-level and scheduling improvements in releases that can be days apart. A number measured in March does not describe the same software in August. InferenceX re-runs the same benchmark against the same hardware fleet continuously, so the delta between two dates reflects software progress rather than a change in test conditions.

The audience follows from that. The README names operators of large token-serving fleets and the ML community around PyTorch Foundation, vLLM, SGLang and Tri Dao as the parties it is trusted by. In practice that means two kinds of reader: infrastructure teams deciding which framework version to promote onto a production cluster, and framework maintainers who want an external signal that a kernel change actually moved the pareto frontier. It is not aimed at someone profiling a single prompt on a laptop.

## How the benchmark is structured: runners, configs and the model matrix

The repository layout separates the moving parts. Top-level entries include benchmarks/, configs/, runners/, utils/, experimental/, infx/ and golden_al_distribution/, alongside a MODELS.md and a perf-changelog.yaml. That split suggests the intended flow: a configuration selects a model and framework, a runner under runners/ executes it against the target hardware, and results feed the changelog and the public dashboard rather than living only on someone's machine.

The model coverage is deliberately tied to architecture rather than to a single checkpoint. The news entries state that Kimi K2.5 was added because it shares an architecture with Kimi K2.7-Code, GLM5 because it shares an architecture with GLM5.1, and MiniMax M2.5 because it shares an architecture with MiniMax M2.7. That is a pragmatic choice: benchmarking one representative per architecture keeps the continuous run tractable while still covering the newer releases. The trade-off is that a model-specific optimisation which does not apply to its architecture sibling will not show up until the sibling itself is added to the matrix.

The hardware table is the other half of the design. GB300 NVL72, GB200 NVL72, MI355X, B300, B200, MI325X, MI300X, H200 and H100 are listed as supported, while TPUv7x Ironwood Ghostfish, MI455 UALoE72, Vera Rubin NVL72 and Rubin NVL8 are marked as coming soon. Several further entries are anonymised as Chip #1 from Hardware Vendor #1 and so on, which tells you the roadmap includes parts that cannot yet be named. If your accelerator is in the coming-soon block, the platform does not cover you yet.

## Installing InferenceX and running a first benchmark

The README does not give a pip install line or a quickstart command. It points readers to CONTRIBUTING.md for the PR review flow and merge process, and to docs/index.md as the map for architecture, configuration, workflow, eval, runner and troubleshooting references. That is where a new user should start, because the configuration keys and runner entry points live in those documents rather than in the top-level README.

What the repository does tell you is the shape of the working tree, so you can orient before reading the docs. Cloning the default branch gives you the layout below.

```bash
git clone https://github.com/SemiAnalysisAI/InferenceX.git
cd InferenceX
ls benchmarks configs runners utils
```

The four directories are the ones a first run touches: configs holds the benchmark definitions, runners executes them, benchmarks holds the workload definitions, and utils holds shared helpers. The README does not document the contents of each, so treat the docs/ tree as the authority.

Because the README is silent on installation steps, the honest statement is that this project expects you to read docs/index.md and CONTRIBUTING.md before running anything. Anyone who needs a single copy-paste command to get started will not find one here. What the README does provide is the review checklist at docs/PR_REVIEW_CHECKLIST.md, which is the closest thing to a definition of a correct run, and it is worth reading even if you never open a pull request.

## The hardware requirement is the real adoption barrier

InferenceX is not a tool you run on a workstation to compare two framework versions on your own model. The supported SKU list is datacenter-class: NVL72 racks, MI355X, B300, B200, H200, H100. The acknowledgements section makes the dependency explicit, crediting AMD for MI355X and CDNA3 GPUs, NVIDIA and Oracle for GB200 NVL72 access through OCI and B200 GPUs, and Crusoe, CoreWeave, Nebius, TensorWave, Oracle and TogetherAI for compute resources. That is the operational reality of the project: the benchmark runs on donated or sponsored capacity, not on hardware most teams own.

This is the case where InferenceX is the wrong tool. If your question is which of two quantisation schemes is faster on a single A100, or how a new attention kernel behaves on a consumer card, InferenceX will not answer it and its supported-hardware table says so. The second limitation is provenance. The README carries an explicit warning that only SemiAnalysisAI/InferenceX contains the official result, that forks and other repositories are unofficial, and that the benchmark setup and quality of machines and clouds in unofficial repositories may differ, leading to subpar benchmarking. A chart copied from a fork can therefore look authoritative while resting on different hardware. The README also states that forks may not remove this disclaimer, which is a licence-adjacent condition worth reading alongside the Apache-2.0 text.

## InferenceX compared with MLPerf

MLPerf is the obvious alternative and the difference is in the cadence, not the ambition. MLPerf is a periodic submission process: vendors prepare, tune and submit results against a fixed benchmark definition, and the results are published as rounds. That structure produces comparable, audited numbers at a point in time, which is exactly what a procurement decision needs.

InferenceX inverts the trade-off. Its stated purpose is to capture progress in near real time and provide a live indicator of inference performance progress, with results published continuously on the dashboard rather than in submission rounds. The cost of that cadence is that the benchmark definition itself can move: the project has gone from InferenceMAX v1 in October 2025 to InferenceX v2 in February 2026, added GB300 NVL72 in February 2026, and added AgentX, described as a fully open source Apache 2.0 realistic one-million-token-plus long-context multi-turn benchmark, in August 2026. A number from a v1 run and a number from a v2 run are not necessarily the same measurement. MLPerf gives you a stable ruler that is slow to update; InferenceX gives you a fast ruler that occasionally changes length. Which one you want depends on whether you are buying hardware or tracking software.

## Licence, maintenance and what an upgrade costs

The repository is licensed Apache-2.0, and the LICENSE file is at the top level. For most users that means you can read, modify and redistribute the benchmark code, including in a commercial internal tool. Two caveats belong here rather than in legal advice. First, the README asserts a trademark claim through the InferenceX name and states that unofficial forks must be explicitly labelled as unofficial and may not remove the disclaimer. Apache-2.0 grants copyright and patent rights; it does not grant trademark rights, so the labelling requirement is a separate matter from the licence. Second, if you publish numbers derived from the code, the disclaimer condition applies to you.

On maintenance, the last push to the default branch was on 2026-08-19, and the most recent release is tilert-v0.1.5.post2-inferencex.1, dated the same day and described as a TileRT 0.1.5.post2 InferenceX queue backport. The repository is not archived. The news list shows a steady cadence of model additions through 2026, which is consistent with the project's own framing of continuous benchmarking. The upgrade cost is not in the code but in the interpretation: because the platform tracks upstream frameworks and adds models as architectures appear, a result you cited three months ago may no longer describe the current stack. Anyone embedding an InferenceX number in an internal document should record the date and the framework version alongside it.

## Conclusion

Adopt InferenceX if you already operate a multi-GPU or multi-vendor fleet and want a repeatable, continuously re-run comparison of vLLM, SGLang and TensorRT-LLM on the same hardware, under Apache-2.0 and with a public dashboard you can point colleagues at. Do not adopt it as a single-machine latency profiler, and do not adopt it if you cannot supply the hardware the officially supported SKU table lists. Before trusting any number, verify three things: that the run came from SemiAnalysisAI/InferenceX rather than a fork, that the framework and model versions match what you plan to deploy, and that the hardware SKU appears in the officially supported table. The README's own warning that unofficial forks may be benchmarked on different machines is the boundary to take seriously, because a subpar setup produces a subpar number that looks identical in a chart.

## FAQ

### What is InferenceX and who is it for?

InferenceX, formerly InferenceMAX, is an Apache-2.0 inference performance research platform that continuously benchmarks popular open source inference frameworks on a fixed set of hardware. It is aimed at operators of large token-serving fleets and at the ML community around frameworks such as vLLM, SGLang and TensorRT-LLM.

### Which hardware does InferenceX officially support?

The README lists GB300 NVL72, GB200 NVL72, MI355X, B300, B200, MI325X, MI300X, H200 and H100 as supported. TPUv7x Ironwood Ghostfish, MI455 UALoE72, Vera Rubin NVL72 and Rubin NVL8 are marked as coming soon.

### How do I install InferenceX?

The README does not give installation or quickstart commands. It directs readers to CONTRIBUTING.md for the review and merge process and to docs/index.md for the architecture, configuration, workflow, eval, runner and troubleshooting references.

### Are results from an InferenceX fork official?

No. The README states that only the SemiAnalysisAI/InferenceX repository contains the official result and that all other forks and repositories are unofficial, because the benchmark setup and the quality of machines and clouds may differ.

### What is AgentX in InferenceX?

AgentX is described in the news list as the world's first fully open source Apache 2.0 realistic one-million-token-plus long context, multi turn benchmark, with results on the InferenceX dashboard.

## Sources

- [Official documentation](https://inferencex.com/)
- [Official README](https://github.com/SemiAnalysisAI/InferenceX#readme)
- [Project repository](https://github.com/SemiAnalysisAI/InferenceX)
- [Release notes](https://github.com/SemiAnalysisAI/InferenceX/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/semianalysisai-inferencex
