CTranslate2: A Custom Runtime for Transformer Inference on CPU and GPU
Fast inference engine for Transformer models
At a glance
- What is it?
- CTranslate2 is a C++ and Python inference engine for Transformer models, with model conversion, quantization down to INT8 and INT4, and runtime dispatch across MKL, oneDNN, OpenBLAS, Ruy and Apple Accelerate. It is a good fit when you control the model format and want CPU throughput; it is the wrong tool if you need a training framework or a generic graph executor.
- Who is it for?
- Adopt CTranslate2 if you serve a supported Transformer family on CPU, or on GPU with a fixed model, and you are willing to add a conversion step to your pipeline. Do not adopt it if you need training, arbitrary model architectures, or a graph format you can hand-edit.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What CTranslate2 replaces in a serving stack
The README describes CTranslate2 as a C++ and Python library for efficient inference with Transformer models, built around a custom runtime rather than a general deep learning framework. That distinction is the whole point. PyTorch and TensorFlow carry an autograd engine, a training loop, and a graph abstraction that must handle any model a researcher writes. CTranslate2 carries none of that. It supports a fixed list of encoder-decoder, decoder-only and encoder-only architectures, and in exchange it can fuse layers, remove padding, reorder batches, cache keys and values, and serialize weights at reduced precision.
The target user is someone deploying a model that already exists, not someone experimenting with a new one. If you have an OpenNMT-py, OpenNMT-tf, Fairseq, Marian, OPUS-MT or Transformers checkpoint and you need more tokens per second per core, the conversion path is the product. If you are still choosing an architecture, or you need gradients, this is the wrong layer of the stack.
How the runtime gets its speed: conversion, quantization, dispatch
The data flow has two stages. First, a conversion step rewrites a framework checkpoint into CTranslate2's own on-disk format. Second, the runtime loads that format and executes it. The README states that compatible models should first be converted into an optimized model format, and that converters exist for OpenNMT-py, OpenNMT-tf, Fairseq, Marian, OPUS-MT and Transformers. The conversion is where quantization is baked in: model serialization and computation support FP16, BF16, INT16, INT8 and AWQ (INT4).
At execution time the library does not assume a single backend. One binary can include multiple backends, for example Intel MKL and oneDNN, plus multiple instruction set architectures such as AVX and AVX2, and selects among them at runtime based on CPU information. That is why a wheel can run on an older x86-64 machine and a newer one without a rebuild. On the GPU side, the same runtime covers CUDA, and the README notes that tensor parallelism lets a very large model be split across multiple GPUs, with a pointer to docs/parallel.md for the environment setup.
The optimizations listed are concrete and mechanical: layer fusion, padding removal, batch reordering, in-place operations, and caching. None of them change the model's math in principle, but quantization does change numerics, and the README's own benchmark table pairs throughput with a BLEU column for exactly that reason.
Installing CTranslate2 and running a first translation
The README gives a single install command for the Python package. The same package covers conversion and inference, so there is no separate CLI to install for the common path.
pip install ctranslate2After that, the README's usage example is four lines. A Translator is constructed from a path to a converted translation model, and translate_batch takes tokenized input. A Generator is constructed from a converted generation model, and generate_batch takes start tokens.
import ctranslate2
translator = ctranslate2.Translator(translation_model_path)
translator.translate_batch(tokens)
generator = ctranslate2.Generator(generation_model_path)
generator.generate_batch(start_tokens)The important detail is that translation_model_path points at a directory produced by a converter, not at a Hugging Face hub id and not at a .pt or .bin checkpoint. The README does not show the conversion call for any framework; it links to per-framework guides under opennmt.net/CTranslate2/guides for OpenNMT-py, OpenNMT-tf, Fairseq, Marian, OPUS-MT and Transformers. Read the guide for your framework before writing the inference code, because the conversion argument names live there and not in the README.
For AMD hardware there is a second path. The README states that specific Python wheels for AMD ROCm GPUs are provided on the releases page, so a plain pip install is not the ROCm route. The repository also ships a docker/ directory and a .dockerignore at the top level, though the README does not document a published image name.
Where CTranslate2 is the wrong choice
The supported model list is a hard boundary, not a starting point. The README enumerates encoder-decoder families (Transformer base/big, M2M-100, NLLB, BART, mBART, Pegasus, T5, Whisper, T5Gemma, T5Gemma2, MADLAD-400), decoder-only families (GPT-2, GPT-J, GPT-NeoX, OPT, BLOOM, MPT, Llama, Mistral, Gemma, CodeGen, GPTBigCode, Falcon, Qwen2) and encoder-only families (BERT, DistilBERT, XLM-RoBERTa). A custom attention variant, a mixture-of-experts routing scheme, or a multimodal encoder outside that list has no conversion path documented in the README.
The second boundary is the format itself. Because the runtime consumes its own serialized format, you cannot inspect or patch the graph the way you would with an ONNX file. Debugging a numerical discrepancy means comparing against the original framework, not reading the graph.
The README also flags its own edges: the project is described as production-oriented with backward compatibility guarantees, but it also includes experimental features related to model compression and inference acceleration. Treat anything under that experimental label as something to pin and re-verify on upgrade. And the benchmark disclaimer is worth quoting in spirit: the results are stated to be valid only for the configuration used during that benchmark, so absolute and relative performance may change with different settings. Do not lift those tokens-per-second figures into a capacity plan for your own hardware.
CTranslate2 compared with llama.cpp and ONNX Runtime
The closest alternative in the related searches is llama.cpp, and the difference is architectural. llama.cpp is a C/C++ inference engine built around the GGUF format, with its own quantization schemes and a focus on running decoder-only LLMs locally. CTranslate2 is broader in model coverage (it also handles encoder-decoder translation and encoder-only encoders such as BERT and XLM-RoBERTa) but narrower in format flexibility. If your workload is a single Llama or Mistral model on a laptop, llama.cpp's tooling is built for that case. If your workload is a translation pipeline plus an embedding encoder plus a Whisper transcription model, CTranslate2's single runtime covers all three.
Against ONNX Runtime, the split is different again. ONNX Runtime executes a portable graph format that many exporters target, and the graph is a first-class artifact you can inspect and transform. CTranslate2 trades that portability for a runtime that knows the operations of specific Transformer families and can therefore fuse and reorder them. ONNX gives you reach; CTranslate2 gives you a runtime tuned for a smaller set of shapes.
For Whisper specifically, the searches point at faster-whisper and whisper.cpp. The README does not discuss faster-whisper, so the only verifiable statement here is that Whisper appears in CTranslate2's supported encoder-decoder list, which is why a Whisper front end would build on this runtime at all.
Maintenance, release cadence and the MIT licence
The repository is not archived. The last push was on 2026-08-31, and the most recent release, v4.8.2, is dated the same day, following v4.8.1 on 2026-07-03 and v4.8.0 on 2026-06-06. That is a roughly monthly patch cadence across the three most recent releases, which is a reasonable signal for a project you would put in a serving path.
The README states that the project comes with backward compatibility guarantees and links to a versioning page. That is a commitment about the API surface, not about model numerics. On upgrade, the thing most likely to move is output on a quantized model, so a regression check on your own evaluation set is the practical cost of each version bump, not a code migration.
The licence is MIT, listed at the repository root and in the metadata. MIT is permissive: it allows commercial use and modification with the copyright notice retained. That is a statement about the licence text, not legal advice, and it says nothing about the licences of the model weights you convert. Model licences are separate and are not covered by the repository's LICENSE file.
Editorial conclusion
Adopt CTranslate2 if you serve a supported Transformer family on CPU, or on GPU with a fixed model, and you are willing to add a conversion step to your pipeline. Do not adopt it if you need training, arbitrary model architectures, or a graph format you can hand-edit. Before committing, verify that your exact architecture appears in the supported model list, check which CPU backend your wheel resolves to with ctranslate2.get_supported_compute_types, and confirm that the converted model reproduces your reference output within your own tolerance.
Frequently asked questions
How do I install CTranslate2?
The README gives one command, pip install ctranslate2, which installs the Python package used for both model conversion and inference. If you have an AMD ROCm GPU, the README states that specific Python wheels are provided on the releases page instead.
What is CTranslate2?
It is a C++ and Python library for efficient inference with Transformer models, built as a custom runtime that applies weights quantization, layer fusion and batch reordering. It supports a fixed list of encoder-decoder, decoder-only and encoder-only model families.
How does CTranslate2 compare with llama.cpp?
llama.cpp is not discussed in the README, so the comparison has to be made on scope. CTranslate2 covers encoder-decoder translation models, decoder-only models and encoder-only models such as BERT and XLM-RoBERTa, while its own serialized format is the only format its runtime loads.
How does CTranslate2 compare with ONNX Runtime?
ONNX Runtime executes a portable graph format that many exporters target and that you can inspect. CTranslate2 instead converts models into its own optimized format and runs a runtime specialized for the Transformer families it lists, which is what allows the fusion and reordering optimizations.
Does CTranslate2 support Whisper?
Yes. Whisper is listed among the supported encoder-decoder model types in the README. The README does not document faster-whisper, so consult that project separately if you plan to use it as a front end.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/opennmt-ctranslate2)