# Arraymancer finds BLAS and LAPACK on your PATH, and its newest tag is five years older than its last commit

> Arraymancer is a Nim tensor library with CPU, Cuda and OpenCL ambitions whose behaviour is set almost entirely by compile flags. The interesting decisions are the two deprecated switches, the AVX512 portability trade, and the gap between release tags and repository activity.

**mratsim/Arraymancer** — A fast, ergonomic and portable tensor library in Nim with a deep learning focus for CPU, GPU and embedded devices via OpenMP, Cuda and OpenCL backends

- Repository: https://github.com/mratsim/Arraymancer
- Website: https://mratsim.github.io/Arraymancer/
- Stars: 1,409 · Forks: 101
- Language: Nim
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mratsim-arraymancer

## BLAS and LAPACK are found by searching your path, not by configuration

By default Arraymancer does not ask which linear algebra library to use. Leave -d:blas unset and the build goes looking on your path for something like blas.so, blas.dll or libopenblas.dll, and -d:lapack searches the same way for lapack.so, lapack.dll or libopenblas.dll. The two flags exist for the case where that search picks the wrong library, and both take a name rather than a switch: -d:blas=blaslibname and -d:lapack=lapacklibname. The instruction is to set them only when you have a reason to want one specific library. The bindings that connect Nim to these libraries are separate projects, nimblas and nimlapack. Library paths can be retuned after installation in nim.cfg, which is where you would point at OpenBLAS, MKL or Cuda. The stated defaults work on Mac and Linux. Windows is the case that needs extra work: download libopenblas.dll or another BLAS or LAPACK DLL and copy it into a folder on your path or into the compilation output folder.

## Two deprecated flags survive, and the old MKL one changed your threading

The flag list still accepts -d:mkl and -d:openblas, and marks both as deprecated. Their replacements are pairs rather than single switches: -d:blas=mkl with -d:lapack=mkl, and -d:blas=openblas with -d:lapack=openblas. The guidance is to reach for the pairs only when you intend to force MKL or OpenBLAS instead of letting the build search for whatever BLAS and LAPACK it finds. One behaviour is worth knowing before you migrate. The old -d:mkl flag implies -d:openmp, so it silently turned multithreaded compilation on as a side effect. The replacement pair carries no such implication, which means a project moving from -d:mkl to -d:blas=mkl and -d:lapack=mkl has to pass -d:openmp itself to keep the threading it already had. Plain -d:openmp remains a valueless switch on its own, and -d:cudnn likewise implies -d:cuda, so asking for CuDNN pulls in Cuda support whether or not you also named it.

## -d:avx512 is a deployment decision, not a machine detection

Setting -d:avx512 does not select an instruction set by itself. What it does is hand -mavx512dq to gcc or clang, and the text is explicit that without that flag the resulting binary does not use AVX512 even on a CPU that supports it. The cost is stated in the same breath: a binary built this way is incompatible with CPUs that do not support AVX512. That makes the flag a claim about every machine the binary will run on, decided at build time and not adjustable afterwards, which is a different question from whether the machine you are compiling on happens to have the instructions. The discussion behind the flag lives in the comments of issue 505, and the text dates that discussion from v0.7.9. That version number is later than the newest tag in the release list, so the reasoning behind the flag was written after the last released version. Reading the flag documentation means reading guidance that shipped untagged.

## -d:danger removes the bounds check that guards every slice you write

Two supported flags change what happens when a program misbehaves rather than what it computes. -d:release is the ordinary Nim release mode, which drops stacktraces and debugging information. -d:danger removes runtime checks, and array bound checking is the one named explicitly. Tensor indexing in this library is range based, as in foo[1..2, 3..4], so the bounds check is the mechanism sitting between a wrong range and memory the program should not touch. Removing it is a deliberate trade of a diagnostic for throughput on a hot loop. The two flags interact badly if combined carelessly: -d:release has already removed the stacktraces, so a build that passes both has neither the check that would have caught the bad index nor the trace that would have told you where it happened. The usual pairing is -d:release for anything shipped and -d:danger only for a profiled section that needs it.

## The first sample builds a Vandermonde matrix out of nested sequences

The opening example constructs its tensor from ordinary Nim sequences and converts at the end, which shows where the boundary between seq and Tensor sits:

```Nim
import math, arraymancer

const
    x = @[1, 2, 3, 4, 5]
    y = @[1, 2, 3, 4, 5]

var
    vandermonde = newSeq[seq[int]]()
    row: seq[int]

for i, xx in x:
    row = newSeq[int]()
    vandermonde.add(row)
    for j, yy in y:
        vandermonde[i].add(xx^yy)

let foo = vandermonde.toTensor()
```

The nested loop fills row i of a five by five grid with xx raised to the power of yy, and toTensor converts it. Echoing the result prints the shape and the backend alongside the data, in the form Tensor[system.int] of shape "[5, 5]" on backend "Cpu", which confirms that a tensor created from plain integers lands on the CPU backend. The last line of the sample slices it with foo[1..2, 3..4], two ranges in one bracket, the same syntax that -d:danger stops checking.

## The same tensor reshaped two ways concatenates into two different shapes

The reshaping sample takes one flat sequence and splits it twice before joining, and the two joins disagree about what they produce:

```Nim
import arraymancer, sequtils

let a = toSeq(1..4).toTensor.reshape(2,2)

let b = toSeq(5..8).toTensor.reshape(2,2)

let c = toSeq(11..16).toTensor
let c0 = c.reshape(3,2)
let c1 = c.reshape(2,3)

echo concat(a,b,c0, axis = 0)
# Tensor[system.int] of shape "[7, 2]" on backend "Cpu"
```

Six values are reshaped into three rows of two and also into two rows of three. Concatenating along axis = 0 stacks the row counts and prints a tensor of shape "[7, 2]", while the same call with c1 and axis = 1 prints shape "[2, 7]", where the column counts stack instead. Nothing about the data changed between the two runs, only the shape c was given and which axis was named. Both results keep the backend label, so the operation is arithmetic rather than a transfer between devices.

## Broadcasting is a 4 by 1 tensor added to a 1 by 3 tensor

The broadcasting sample is three lines of setup and one operator:

```Nim
import arraymancer

let j = [0, 10, 20, 30].toTensor.reshape(4,1)
let k = [0, 1, 2].toTensor.reshape(1,3)

echo j +. k
# Tensor[system.int] of shape "[4, 3]" on backend "Cpu"
```

Neither operand has the shape of the result. j is four rows by one column and k is one row by three columns, and adding them with the dotted operator yields four rows by three columns, printed with values running 0, 1, 2 on the first row through 30, 31, 32 on the last. The axis argument is left out entirely, which is the point of the operation: the singleton dimension of each side is the one that gets stretched. The sample carries an attribution to Scipy, so the behaviour it demonstrates is the broadcasting rule familiar from that library rather than an Arraymancer invention, and the CPU backend label confirms it happens without a Cuda or OpenCL build.

## The examples run from XOR to Shakespeare, and the tags stopped in 2021

The examples directory moves from a perceptron written from scratch to a text generator, and the file names show the path: ex01_xor_perceptron_from_scratch.nim, ex02_handwritten_digits_recognition.nim, ex04_fizzbuzz_interview_cheatsheet.nim, ex05_sequence_classification_GRU.nim, ex06_shakespeare_generator.nim and ex07_save_load_model.nim. Example 3 exists twice, as ex03_simple_two_layers.nim and ex03_pytorch_simple_two_layers.py, so the same two-layer network is available in the project's own idiom and in the framework it is imitating. Example 6 reads from shakespeare_input.txt and pride_and_prejudice.txt, and the text it generates was trained for 45 minutes on a laptop CPU to produce 4000 characters. Against that, the release list stops: v0.7.0 shipped on 2021-07-04, before v0.6.1 in 2020 and v0.6.0 in January 2020, while the last push to master is dated 2026-05-26. Commits have continued for years without a new tag.

## Conclusion

Use Arraymancer if you want a Nim ndarray with autograd and can live with a build-time backend decision. Before committing, check two things: whether your deployment machines all support AVX512 if you intend to pass -d:avx512, and whether the OpenMP thread count you had under the old -d:mkl flag is still set after you migrate to -d:blas=mkl. Treat the tag list with care, since the newest release predates the last commit by years.

## FAQ

### What does Arraymancer need on Windows before it will build?

A BLAS or LAPACK DLL on disk. The defaults are said to work on Mac and Linux, while Windows needs libopenblas.dll or another BLAS or LAPACK DLL copied into a folder on your path or into the compilation output folder, with paths retunable in nim.cfg.

### How do I force Arraymancer to use OpenBLAS instead of the library it finds?

Pass -d:blas=openblas together with -d:lapack=openblas. The single -d:openblas flag still works but is marked deprecated, and the same deprecation applies to -d:mkl, whose replacement is -d:blas=mkl with -d:lapack=mkl.

### Does Arraymancer need a GPU to be useful?

No. The ndarray component works on its own without the machine learning component, and the Cuda and CuDNN switches are opt-in at compile time. -d:openmp gives multithreaded compilation on the CPU.

### Where is the Arraymancer tutorial and the rest of the documentation?

The tutorial lives at the project documentation site, with the first-steps page at mratsim.github.io/Arraymancer/tuto.first_steps.html. The repository also carries a docs directory, a changelog, a Contributors file and a benchmarks directory beside src and tests.

### How current is the released version of Arraymancer?

The newest tag is v0.7.0 from 2021-07-04, preceded by v0.6.1 in November 2020 and v0.6.0 in January 2020. The last push to the master branch is dated 2026-05-26, so commits continued well after the final tag.

## Sources

- [License: Apache-2.0](https://github.com/mratsim/Arraymancer/blob/master/LICENSE)
- [mratsim/Arraymancer on GitHub](https://github.com/mratsim/Arraymancer)
- [Project website](https://mratsim.github.io/Arraymancer/)
- [README](https://github.com/mratsim/Arraymancer/blob/master/README.md)
- [Releases](https://github.com/mratsim/Arraymancer/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mratsim-arraymancer
