apache/mahout: the repository name, the package name and the subject have all diverged
Apache Mahout - an environment for quickly creating scalable, performant machine learning applications.
At a glance
- What is it?
- The Apache Mahout repository now contains Qumat, a Python quantum computing library with a Rust GPU data plane, published to PyPI as qumat with release tags named after it. The project description still says machine learning, the classifiers say physics, and the build tolerates a missing CUDA toolkit.
- Who is it for?
- Adopt this if you want one circuit API over Qiskit, Cirq and Braket, or a GPU state-preparation path that exchanges tensors with PyTorch through DLPack, and treat the repository name as an accident you will have to explain to your team. Do not come here for Mahout the scalable machine learning environment, because that is the stale description at the top of the README and none of the package, the tags or the classifiers match it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The repository is mahout, the package is qumat, the subject is quantum
Start with the names, because they are the first thing anyone will trip over. The repository is apache/mahout. The package on PyPI is qumat, and the PyPI badges in the README point at the qumat project rather than anything named Mahout. The release tags are named mahout-qumat-0.5.0 and mahout-qumat-0.6.0, so the tag naming carries both identities. The primary language recorded for the repository is Rust. And the actual content is a quantum computing library. Meanwhile the project description at the top of the README still says the goal is to build an environment for quickly creating scalable, performant machine learning applications, which is the historical positioning of a machine learning library that once ran on Spark. The metadata disagrees with itself in a way that is worth mapping precisely, because each field points somewhere different. The description says machine learning. The package description in the project metadata says a library for composing quantum machine learning. The topic classifiers say scientific and engineering software, physics, and artificial intelligence, in that order. The code is a circuit abstraction and a GPU data plane. The practical cost of this is discoverability in both directions. Someone looking for the Mahout they remember will arrive and find a quantum library, and someone looking for quantum tooling has no reason to search for Mahout. If you are evaluating this, evaluate the code and ignore the description line entirely.
Two extras, because the data plane is a native extension with CUDA
The quick start is two lines, and the difference between them is the whole architecture. The plain install is
pip install qumatand the version with the data plane is
pip install qumat[qdp]That extra exists for one reason: QDP is described as encoding classical data into quantum states using GPU-accelerated kernels, which means a compiled native extension with a CUDA dependency, and that cannot be a hard requirement of a package that also runs a pure-Python circuit layer. The split is the right packaging decision, and it also defines the boundary of what you get. Install without the extra and you have the circuit abstraction over your chosen backend and nothing else. The data plane lives in a separate import namespace, qumat.qdp, so the failure when you forgot the extra appears at import time rather than at install time, which is a small mercy. What the extra pulls in is visible in the tree structure rather than in the README: there is a qdp directory and a qumat directory side by side, with examples split the same way into examples/qdp and examples/qumat. A Rust toolchain, a uv lock file and a group of dev dependencies in the sync command all belong to the data plane, while the circuit layer is ordinary Python. The project metadata also lists Rust as a programming language classifier, which is the packaging-level admission that this is a two-language package.
A CUDA stub means the Python tests pass on a machine with no GPU
The Makefile is where the honest engineering notes are, and it documents the build's tolerance for missing hardware in plain language. Two detections run at configure time. One asks whether an NVIDIA driver is present by testing for the nvidia-smi command and listing devices, and one asks whether the CUDA compiler is on the path. The comments attached to them explain the consequence, and the Python one is the important one. The comment says that detecting a driver is sufficient for the Python test target, because qdp-core can build against a CUDA stub fallback when the toolkit is absent, and that in that case the Rust extension simply is not GPU-functional at runtime. So your Python tests run and pass on a laptop with no NVIDIA hardware at all, and the data plane does not work there. That is a deliberate design so that contributors without a GPU can still work on the library, and it is also a trap for continuous integration, where a green Python test run tells you nothing about the path you are shipping. The Rust target has the opposite requirement. Its comment says the qdp-core integration tests link against the CUDA runtime from the toolkit, and it makes a point worth repeating: the copy of that runtime bundled inside a PyTorch wheel does not satisfy the link, so the toolkit itself must be installed and on the path. The Rust tests also depend on a coverage target that installs cargo-llvm-cov if it is missing, which is a Rust-native coverage path rather than a Python one.
Amplitude encoding fixes your data length at two to the qubit count
The data plane example is four lines, and the third line contains the constraint that shapes every use of the library:
engine = qdp.QdpEngine(device_id=0)
qtensor = engine.encode([1.0, 2.0, 3.0, 4.0], num_qubits=2, encoding_method="amplitude")Four values are encoded into two qubits. That is not a coincidence, it is the definition of amplitude encoding, where the state of n qubits has exactly two to the power of n amplitudes. The method is also a parameter rather than the only option, which implies other schemes exist, and the device_id parameter is how you select a GPU. The engineering consequence is specific and immediate: your data length is not a free choice, it is a power of two determined by the qubit count you ask for. That is a hard physical limit of state preparation rather than an implementation choice, so it cannot be worked around by a different backend. It also means a batch of samples cannot be a ragged tensor, and it means the tensor you hand the encoder is immediately reinterpreted as a fixed-length amplitude vector, which is worth knowing before you build a preprocessing pipeline that assumes arbitrary shapes. The zero-copy claim sits underneath this. Tensor exchange with PyTorch, NumPy and TensorFlow is described as happening through DLPack without overhead, which matters because the encode call sits on the critical path of every batch. A serialization round trip there would dominate the cost of the operation itself.
Three backends and a simulator option, which is not the same as hardware
The circuit abstraction is the older and simpler half. You construct an object with a configuration dictionary naming a backend and its options, create a circuit of a given size, apply gates, and execute. The README's example uses Qiskit as the backend with a simulator type option set to an Aer simulator, then applies a Hadamard gate to the first qubit and a controlled-not from the first to the second. The supported backends are Qiskit, Cirq and Amazon Braket, with the promise written once and executed anywhere. Here is the honest caveat. A circuit that runs on an Aer simulator tells you that your code is syntactically and semantically valid against one provider's model, and that is worth having, but it does not tell you the circuit behaves the same on real hardware, does not tell you the other two backends accept the same gate set, and does not tell you anything about noise, connectivity constraints or queue behaviour. Three further things follow from the backend list. Braket is a paid cloud service, so one of the three targets costs money to exercise. The simulator option is exposed in the configuration, which means the default path most developers take is the one that requires no hardware at all. And the class in the example is named QuMat with capitals while the package is qumat in lowercase, a small remnant of the rename, which is worth knowing if you are reading older examples.
Python 3.13 is excluded, and the container is for AMD while the build detects NVIDIA
Two details in the build configuration will stop you if you do not see them coming. The first is the Python range. The project metadata declares support for 3.10 and later but below 3.13, and the classifiers list 3.10, 3.11 and 3.12 only. So 3.13 is excluded, and no reason is given. For a scientific package that wants to attract contributors, an upper bound with no explanation is the kind of thing that gets noticed, and if you are on 3.13 today you will find pip refusing the install with a message that looks like a bug rather than a decision. The second is the GPU story, which is inconsistent between two files. The topics list CUDA, and the Makefile tests for an NVIDIA driver with nvidia-smi and for the NVIDIA compiler nvcc, and its comments talk about the NVIDIA CUDA toolkit and nvidia-cuda-toolkit. But the only container file in the repository is named Dockerfile.qdp-amd, for AMD hardware. So the build tooling assumes one vendor and the packaged container targets another, and there is no NVIDIA container visible in the top-level listing. If you are on an NVIDIA machine you are probably fine, because the build detects it; if you are on AMD you have a container and a build that does not know about you. Around the edges there is a healthy amount of process, with a pre-commit configuration, a devcontainer, a link checker configuration, a backport tool configuration, a doap metadata file for package listings and an examples tree.
What this repository is for now, and who should not come looking
Judged on its current contents, this is a two-part project: a circuit abstraction that lets you write a circuit once against a chosen backend, and an optional GPU data plane that encodes classical tensors into quantum states with no copy between the tensor libraries you already use. That is a coherent piece of software with a specific audience, and the audience is people doing quantum machine learning experiments who want to avoid writing three circuit implementations. The version line supports that reading. The package is at a development version of 0.7.0, the last two releases were 0.5.0 in February 2026 and 0.6.0 in May 2026, the project status is declared as beta, and the last push was on 2026-09-24. The boundaries are equally clear. This is not a general machine learning library and it is not a replacement for whatever people mean when they say Mahout. It is not a simulator, and running on Aer is not evidence of hardware behaviour. It is not a place to learn quantum computing from, because the documentation is a set of gate reference pages and an example, not a tutorial. And for anything where you need a stable, widely adopted dependency with a long compatibility record, a 0.x package whose repository name does not match its package name is a poor foundation. Use it if you are experimenting with quantum encoders, and expect to read the source when the documentation runs out.
Editorial conclusion
Adopt this if you want one circuit API over Qiskit, Cirq and Braket, or a GPU state-preparation path that exchanges tensors with PyTorch through DLPack, and treat the repository name as an accident you will have to explain to your team. Do not come here for Mahout the scalable machine learning environment, because that is the stale description at the top of the README and none of the package, the tags or the classifiers match it. Three things to check before you install. Python is pinned to 3.10 through 3.12, so 3.13 is excluded, and there is no reason given. The qdp extra pulls a native Rust extension, and the Makefile notes that a CUDA stub lets the Python tests pass on a machine with no GPU, so a green run does not mean the data plane works. And the amplitude encoding example encodes four values into two qubits, which means your data length is fixed by the qubit count, so size the encode call against that limit before you design a batch shape around it.
Frequently asked questions
How do I install Qumat from the apache/mahout repository?
Run pip install qumat for the circuit abstraction, or pip install qumat with the qdp extra for the GPU data plane. The data plane is a native Rust extension with a CUDA dependency, which is why it is a separate extra rather than a hard requirement.
What quantum computing backends does Qumat support?
Qiskit, Cirq and Amazon Braket, behind one API, with a configuration naming the backend and its options. The README example uses Qiskit with an Aer simulator as the simulator type.
Does Qumat run on a machine without an NVIDIA GPU?
The Python test target does. The Makefile notes that qdp-core can build against a CUDA stub fallback when the toolkit is absent, in which case the Rust extension is not GPU-functional at runtime. The Rust tests instead require the CUDA toolkit on the path, and the copy inside a PyTorch wheel does not satisfy that link.
What does amplitude encoding constrain in Qumat QDP?
The data length. The README example encodes four values with two qubits, because the state of n qubits holds two to the power of n amplitudes. Your batch shape is therefore fixed by the qubit count you request, and it cannot be ragged.
Which Python versions does Qumat support?
The project metadata requires Python 3.10 and later but below 3.13, with classifiers for 3.10, 3.11 and 3.12. Python 3.13 is therefore excluded, and the repository gives no reason.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-mahout)