Model or dataset
microsoft/vidur avatar
microsoft/vidur

Vidur: Simulate LLM Inference Clusters Before You Buy the GPUs

Accurate, large-scale, and extensible simulator for LLM inference Systems

686 stars131 forksPythonMIT

At a glance

What is it?
Vidur is a high-fidelity LLM inference system simulator from Microsoft Research that models serving performance across different hardware configurations, scheduling algorithms, and workloads without requiring continuous GPU access. It was presented at MLSys 2024 and supports tensor parallelism, pipeline parallelism, and several pre-profiled models including Llama 2, Llama 3, and Qwen-72B.
Who is it for?
Vidur is a fit for ML infrastructure teams doing capacity planning before provisioning hardware, and for researchers testing scheduling algorithms without needing GPU time. The supported model and hardware matrix is limited to what has been profiled (seven models across four hardware configurations), so teams running models outside that list need to complete a profiling step first.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Vidur Simulates and Who Uses It

Vidur addresses a specific cost problem in LLM infrastructure: evaluating serving configurations requires running actual inference, which consumes GPU time before you know whether the configuration is optimal. Vidur decouples that evaluation from actual hardware by simulating the inference cluster's behavior using pre-computed execution time profiles. The simulation covers time-to-first-token (TTFT), time per output token (TPOT), request end-to-end time, and batch size dynamics under configurable workloads.

The README identifies three use cases: studying system performance under different workloads and hardware, capacity planning to find the best deployment configuration per dollar, and testing new research ideas like scheduling algorithms or speculative decoding. The profiling step that makes this possible requires real GPUs once per model-hardware pair, but the simulation afterward does not. The README notes that profiling instructions are in docs/profiling.md.

Supported Models and Hardware Configurations

Vidur ships pre-computed profiles for seven models across four hardware configurations. The hardware configurations are A100 80GB DGX nodes (8 GPUs, fully NVLink-connected), H100 DGX nodes (same topology), 4xA100 80GB pairwise NVLink nodes, and 8xA40 pairwise NVLink nodes.

The supported model list includes `meta-llama/Meta-Llama-3-8B`, `meta-llama/Meta-Llama-3-70B`, `meta-llama/Llama-2-7b-hf`, `codellama/CodeLlama-34b-Instruct-hf`, `meta-llama/Llama-2-70b-hf`, `internlm/internlm-20b`, and `Qwen/Qwen-72B`. H100 DGX and 8xA40 profiles exist for Llama-2-7b, CodeLlama-34b, Llama-2-70b, InternLM-20b, and Qwen-72B but not for Llama-3-8B or Llama-3-70B.

All models cap at a 4k context length except Llama-3-8B and Llama-3-70B, which support 16k context by passing additional CLI parameters for the random forest execution time predictor. Pipeline parallelism is available for all models; the PP dimension must divide the number of layers evenly. In DGX nodes, tensor parallelism dimensions of 1, 2, 4, and 8 are supported. In 4x pairwise NVLink nodes, TP4 is less efficient than TP4 on DGX because the inter-pair interconnect is slower than NVLink.

Setting Up and Running the Simulator

The recommended setup method uses mamba, which resolves the conda environment faster than conda itself:

sh
mamba env create -p ./env -f ./environment.yml
mamba env update -f environment-dev.yml

A standard venv alternative is available for environments without conda. Create and activate the venv, then install dependencies:

sh
python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

Run the simulator from the repository root:

sh
python -m vidur.main

For a specific configuration, pass parameters on the command line. This example runs Llama-3-8B on a single A100 with synthetic workload:

sh
python -m vidur.main  \
--replica_config_device a100 \
--replica_config_model_name meta-llama/Meta-Llama-3-8B \
--cluster_config_num_replicas 1 \
--replica_config_tensor_parallel_size 1 \
--replica_config_num_pipeline_stages 1 \
--request_generator_config_type synthetic \
--synthetic_request_generator_config_num_requests 512

The full parameter list is available via `python -m vidur.main -h`.

Simulation Output: Chrome Traces and wandb Metrics

Each simulation run writes its output to `simulator_output/<TIMESTAMP>/`. Metrics are logged to wandb and a local copy is stored in the same directory. The README points to docs/metrics.md for a description of all logged metrics.

Vidur exports a Chrome trace for every simulation. Open it at `chrome://tracing/` or `edge://tracing/` and load the trace file to inspect scheduling events and batch composition over time. This trace format is useful for identifying where requests wait in the queue versus where they stall during decode.

wandb integration is required by default but can be disabled. To opt out, either set `WANDB_MODE=disabled` in the shell before running, or clear the `wandb_project` and `wandb_group` fields in `vidur/config/default.yml` and remove the corresponding CLI parameters. The README shows both methods.

Canary Branch: Prefix Caching, Routing Policies, and Reduced Memory

The README documents a canary branch that contains several improvements not yet merged to main: prefix caching support, different routing policies, and reduced memory requirements for the simulator itself. The README warns that the canary branch has sharp edges under active resolution.

This matters because prefix caching is a significant optimization in production inference systems, and its absence in the main branch means simulations comparing cached versus uncached workloads are not supported without switching to canary. Teams who need that feature accept the risk of encountering instability.

Code formatting uses black, isort, and flake8, all configured in the Makefile. Running `make format` applies black and isort together; `make lint` checks without modifying files.

Limitations: Fixed Profiles, No New Model Support Out of the Box

Vidur's accuracy depends on pre-computed profiles specific to each model-hardware combination. A model or hardware configuration not in the shipped list requires running the profiling procedure described in docs/profiling.md on actual hardware before it can be simulated. This is a one-time cost per model-hardware pair, but it means Vidur is not immediately useful for evaluating a new model without initial GPU access.

A comparable alternative is vLLM's benchmark tooling, which runs real inference to measure throughput and latency. vLLM measures actual system performance rather than simulating it, which gives precise numbers for supported configurations but requires continuous GPU access for each measurement. Vidur's advantage is that once profiles exist for a configuration, any scheduling policy or workload variation can be tested in minutes without hardware, making it faster for systematic capacity studies.

The repository version is 0.0.1 in setup.py, indicating it is still early-stage software with no stable API guarantees.

Contributions, License, and Maintenance

Vidur is licensed under MIT. The last push was on 2026-08-24. The project is attributed to the MSR-India Systems Group and the Systems for AI Lab at Georgia Tech in setup.py.

Contributions require signing a Contributor License Agreement (CLA) managed by Microsoft's CLA bot, which checks pull requests automatically. The project has adopted Microsoft's Open Source Code of Conduct. The code requires Python 3.10 or later as declared in setup.py, and runtime dependencies include numpy, pandas, scikit-learn, wandb, plotly_express, matplotlib, seaborn, and ddsketch, all listed in requirements.txt.

The MLSys 2024 paper and the associated talk at mlsys.org are linked in the README for readers who want the full technical description of the profiling methodology and the simulation engine's accuracy validation. The README refers to the paper as the authoritative source for the design rationale behind the random forest execution time predictor that drives the simulation.

Editorial conclusion

Vidur is a fit for ML infrastructure teams doing capacity planning before provisioning hardware, and for researchers testing scheduling algorithms without needing GPU time. The supported model and hardware matrix is limited to what has been profiled (seven models across four hardware configurations), so teams running models outside that list need to complete a profiling step first. Check the docs/profiling.md file before planning a simulation to confirm whether your target model and hardware combination already has pre-computed profiles.

Frequently asked questions

What is Vidur AI?

Vidur is an LLM inference system simulator from Microsoft Research. It models the performance of serving large language models under different hardware configurations, parallelism settings, and workload patterns without requiring active GPU access after an initial profiling step.

Does Vidur require GPUs to run simulations?

No. After a one-time profiling step on real hardware to generate execution time profiles for a model-hardware pair, Vidur runs simulations without GPUs. The repository ships pre-computed profiles for seven models across four hardware configurations.

Which models does Vidur support out of the box?

Vidur ships profiles for Llama-3-8B, Llama-3-70B, Llama-2-7b, CodeLlama-34b, Llama-2-70b, InternLM-20b, and Qwen-72B. Models outside this list require a profiling run on actual hardware before they can be simulated.

What output does Vidur produce from a simulation?

Vidur writes metrics to wandb and stores a local copy in the simulator_output directory. It also exports a Chrome trace for each simulation that can be loaded at chrome://tracing to inspect scheduling and batching behavior over time.

Official sources

  1. Issues
  2. License: MIT
  3. microsoft/vidur on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-vidur.svg)](https://hysenlabs.com/projects/microsoft-vidur)