# Magnetron: a C machine learning runtime with a Python API and no runtime dependencies

> Magnetron is a from-scratch ML runtime in C with a small Python interface. It is built for engineers who want to read the tensor system, dispatch layer and autograd engine rather than call into a large framework. The CUDA backend is documented as incomplete.

**MarioSieg/magnetron** — A zero-dependency ML framework in C with a modern Python API for full control over execution and memory.

- Repository: https://github.com/MarioSieg/magnetron
- Stars: 707 · Forks: 39
- Language: C
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mariosieg-magnetron

## What Magnetron solves, and who it is actually for

Most machine learning work happens behind an abstraction wall. You write Python, the framework decides memory layout, kernel selection and graph construction, and when something is slow or wrong you have to reason about a stack you did not build. Magnetron takes the opposite position. It is a machine learning runtime written in C with its own tensor system, operator set, autograd engine and execution model, exposed through a small Python interface. The README states the goal directly: keep the stack small enough to understand and hackable, but powerful enough to run real models.

The intended audience is narrow and the README names it. Magnetron is for developers who want full control over execution and memory, and for people who want a clean base for experimentation. If your job is training a large model on a deadline, this is not the tool. If your job is understanding why a matmul dispatches to a particular SIMD path, or prototyping a new memory layout, the distance between your idea and the running code is much shorter here than in a layered framework.

The project is explicit that it does not compete with PyTorch on ecosystem or feature count. That framing matters when you evaluate it. Magnetron is not a drop-in replacement with a smaller install size. It is a different axis: inspectable core, explicit execution, minimal dependencies, kernels that are meant to be modified.

## How the runtime is put together: tensors, views, dispatch and autograd

The architecture is described as a single cohesive runtime rather than a set of loosely coupled libraries. Four pieces carry most of the weight.

The tensor system owns dtype, shape, strides and memory. It supports a view system with a view solver, which the README says enables slicing, reshaping and broadcasting semantics similar to PyTorch while staying explicit. That is the part worth reading first if you care about memory: a view solver means shape transformations are computed rather than hardcoded per operation, and it is the mechanism that lets broadcasting stay predictable.

Execution is eager with a dynamic autograd graph. The graph is reverse-mode and is constructed per forward pass, then traversed during backward. Nothing is traced ahead of time and nothing is cached between steps by the framework, so what you see in a forward pass is what backward walks.

Above that sits a central dispatch layer that maps high-level operations to architecture-specific kernels. The CPU backend is a multi-dispatch design with compile-time optimized kernels across Intel, AMD Zen1 to Zen5 and ARM. At runtime, CPUID-based detection picks the kernel path. The supported instruction sets listed in the README are SSE 1 through 4, AVX, AVX2, FMA, AVX-512, AVX-512-BF16, AVX-512-FP16, F16C and ARM NEON, combined with multithreaded execution.

The CUDA backend is the gap. The README says the kernel layer is implemented, while memory management, the execution pipeline and integration are still being completed. Treat any GPU claim as unfinished until you check the current state of that directory yourself.

## Installing Magnetron and running a first tensor

The README says Magnetron is available on PyPI and asks you to be inside a Python virtual environment. The published package carries no runtime dependencies, which matches the `dependencies = []` entry in `pyproject.toml`.

```bash
pip install magnetron
```

If you use uv, the README gives the equivalent form.

```bash
uv pip install magnetron
```

For local development, clone with submodules because the repository has an `extern/` directory and a `.gitmodules` file, then install from the checkout with verbose output.

```bash
git clone --recursive https://github.com/MarioSieg/magnetron
cd magnetron
uv pip install . -v
```

Building from source goes through scikit-build-core, which drives CMake. The backend and feature switches live under `[tool.scikit-build.cmake.define]` in `pyproject.toml`: `MAGNETRON_ENABLE_BACKEND_CPU`, `MAGNETRON_ENABLE_BACKEND_CUDA`, `MAGNETRON_BUILD_PYTHON_BINDINGS`, `MAGNETRON_DEBUG`, `MAGNETRON_BUILD_TESTS` and `MAGNETRON_BUILD_BENCHMARKS`. The Python bindings flag is marked as required for the Python package, so if you turn it off you get the C library only.

The README's quick start builds an XOR problem. Inputs and targets are constructed as tensors, and the model is an `nn.Sequential` of `nn.Linear` layers.

```python
from magnetron import Tensor, nn, optim

x = Tensor([[0.0, 0.0], [0.0, 1.0], [1.0, 0.0], [1.0, 1.0]])
y = Tensor([[0.0], [1.0], [1.0], [0.0]])

model = nn.Sequential(
    nn.Linear(2, 2),
    nn.Ta
```

Note the last line of that snippet. The README text available here is truncated mid-expression, ending at `nn.Ta`, so the full model definition and the training loop are not reproduced above. Read the `examples/xor/` directory for the complete version rather than guessing at the activation name. The same caution applies to the operator surface generally: the README points to `docs/Magnetron-Cheatsheet.md` for the full operator, dtype and semantics reference, and that file is the authority, not the README summary.

## Where Magnetron is the wrong tool

The clearest limitation is stated by the project itself: the CUDA backend is in progress. The kernel layer exists, but memory management, the execution pipeline and the integration are described as actively being completed. If your workload needs GPU training today, this is the wrong tool, and no amount of CPU kernel tuning changes that.

The second limitation is the operator catalog. The README describes it as compact but expressive, covering elementwise operations, reductions, tensor transformations, neural building blocks and type casting. That is a deliberate trade. A model that needs an operator outside that set has no fallback path through a large external library, because the point of the project is that there is no large external library. You either write the kernel or you do not run the model.

The third is portability of the optimized path. The CPU backend is built around compile-time kernels for specific microarchitectures with runtime CPUID selection. That is a strength on the hardware it targets and a question mark everywhere else. The README lists Intel, AMD Zen1 to Zen5 and ARM, but it does not say what happens on an unlisted microarchitecture beyond falling back through the dispatch layer.

Finally, the maintenance picture is worth stating plainly. The repository is not archived, and the last push was on 2026-09-06, with v0.2.0 released the same day. That is recent, but the version numbers are still in the 0.x range and the project describes itself as a research-oriented runtime. Expect the API surface to move.

## How it differs from PyTorch in approach

The README includes a comparison table, and the honest reading of it is that the two projects optimize for different things rather than one being a faster version of the other.

PyTorch is a large layered system with an implicit, abstracted execution model and a heavy runtime. Magnetron is a small inspectable core with explicit execution and minimal dependencies. In PyTorch, when you want to know which kernel ran, you go through profilers and dispatch tracing. In Magnetron, the dispatch layer maps operations to kernels directly and the CPU backend selects a path by CPUID, so the answer is closer to the surface.

The deeper difference is what you can change. The README says Magnetron makes it easy to modify kernels and harder to reason about the backend in PyTorch, and that Magnetron is good for research and systems work while PyTorch is good for production and scale. That is a fair split. If your problem is model architecture, PyTorch removes obstacles. If your problem is the runtime itself, PyTorch is the obstacle.

Serialization is where the design shows most clearly. Magnetron uses a native `.mag` format designed for zero-copy, memory-mapped loading, with conversion tools to import weights from external formats. That is a startup-time and memory-footprint decision, not an ecosystem decision, and it is the kind of choice a framework optimizing for compatibility would not make.

## Licence and the cost of keeping up

The repository ships a `LICENSE` file and `pyproject.toml` points the project metadata at it with `license = { file = "LICENSE" }`. The GitHub metadata reports the licence as NOASSERTION, meaning the platform could not classify it automatically. Those two facts together mean you should open the `LICENSE` file and read it before you depend on the project commercially. Nothing here is legal advice, and the classification string is not a substitute for the text.

On upgrade cost, the available information supports a limited but useful observation. Releases are frequent and close together: v0.1.8 on 2026-07-29, v0.1.9 on 2026-08-22, and v0.2.0 on 2026-09-06. Three releases in roughly six weeks, all in the 0.x line, is the pattern of a project still settling its interfaces. The Python package version tracks the release, currently 0.2.0, and requires Python 3.10 or newer.

What you would actually be maintaining is not just the dependency. The optional dependency groups show the split: `examples` pulls in tokenizers, rich, matplotlib and huggingface_hub, while `dev` pulls in torch, numpy, transformers, tiktoken, matplotlib, pytest, pytest-xdist, pyinstrument, line_profiler, rich, tokenizers, huggingface_hub and pandas. The core package has none of these. So the runtime is cheap to install, but the moment you touch the example models or the conversion tooling you are managing a second, much heavier environment alongside it.

## Conclusion

Adopt Magnetron if you need to read and modify the kernels, tensor views and autograd graph of the model you run, and if a CPU-first, C-core runtime with a Python front end matches your work. Do not adopt it if you need a finished GPU backend or a broad operator catalog, because the README describes the CUDA backend as in progress and the project states it is not competing with PyTorch on feature count. Before committing, verify that your target CPU microarchitecture appears in the dispatch list, that the `.mag` conversion path covers the weight formats you have, and that the operator set in the cheat sheet contains every op your model needs.

## FAQ

### How do I install Magnetron?

The README says Magnetron is on PyPI and asks you to be inside a Python virtual environment first. You run `pip install magnetron`, or `uv pip install magnetron` if you use uv. For a local build, clone the repository with `--recursive` and run `uv pip install . -v` from the checkout.

### How do I use Magnetron in Python?

Import `Tensor`, `nn` and `optim` from the `magnetron` package, build input and target tensors, and compose layers with `nn.Sequential`. The README's quick start uses an XOR problem with two `nn.Linear` layers and an `optim` optimizer. The README text available here is truncated mid-model definition, so read `examples/xor/` for the complete training loop.

### How does Magnetron work internally?

It uses its own tensor system with dtype, shape, strides and memory, plus a view solver for slicing, reshaping and broadcasting. Execution is eager with a dynamic reverse-mode autograd graph built per forward pass. A central dispatch layer maps operations to architecture-specific kernels, and the CPU backend selects a path at runtime using CPUID.

### What can you do with Magnetron?

The README lists end-to-end demos under `examples/`: Qwen3 transformer inference in bfloat16 with tokenizer integration, `.mag` weights, a CLI chat and an HTTP streaming API; GPT-2 causal language model inference with KV cache and token streaming; a convolutional autoencoder with a training loop; a 1D linear regression with SGD; and a minimal XOR MLP demonstrating autograd and optimization.

### How do I work with Magnetron's operators?

The README says the operator set covers elementwise operations, reductions, tensor transformations, neural building blocks such as matmul, softmax and layernorm, and type casting with memory views. It points to `docs/Magnetron-Cheatsheet.md` for the full reference of operators, data types and semantics.

## Sources

- [Issues](https://github.com/MarioSieg/magnetron/issues)
- [MarioSieg/magnetron on GitHub](https://github.com/MarioSieg/magnetron)
- [README](https://github.com/MarioSieg/magnetron/blob/master/README.md)
- [Releases](https://github.com/MarioSieg/magnetron/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mariosieg-magnetron
