OpenBLAS: a BLAS and LAPACK library you build for your own CPU
OpenBLAS is an optimized BLAS library based on GotoBLAS2 1.13 BSD version.
At a glance
- What is it?
- OpenBLAS is an optimized BLAS library derived from GotoBLAS2 1.13, shipped with LAPACK and LAPACKE. It solves dense linear algebra performance on CPUs, and it expects you to choose a target and a thread count rather than trusting defaults.
- Who is it for?
- Adopt OpenBLAS if you ship C, Fortran or Python code that spends its time in dense matrix kernels and you are willing to pin TARGET and NUM_THREADS at build time. Do not adopt it if you need a vendor-tuned math library with a support contract, or if you cannot rebuild when you move to a different CPU.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OpenBLAS actually replaces
BLAS is a set of subroutine specifications for basic linear algebra: matrix-vector products, matrix-matrix products, rank updates, triangular solves. LAPACK sits above it and implements factorizations, eigenvalue problems and least squares on top of those kernels. The netlib reference implementations define the interfaces but are written for clarity, not speed. OpenBLAS is the optimized replacement: the README describes it as an optimized BLAS library based on GotoBLAS2 1.13 BSD version, and the LAPACK package comes included. The audience is anyone whose program spends its time inside dgemm, sgemm or a Cholesky factorization: numerical simulation code, statistics and machine learning stacks, R and Python scientific packages that call into a compiled backend. The project's own topics list is blas, lapack, lapacke, and the repository ships cblas.h at the top level, so the CBLAS interface is part of the delivered surface, not an add-on. If your workload is sparse, graph-based, or dominated by memory-bound elementwise operations, the BLAS interface is the wrong level of abstraction and OpenBLAS will not help you.
How the kernels and threading are organized
The build is organized around per-architecture makefiles. The top-level Makefile includes Makefile.system and then descends into BLASDIRS, which the file defines as interface, driver/level2, driver/level3 and driver/others. When DYNAMIC_ARCH is not 1, kernel is added to that list, meaning architecture-specific kernels are compiled into the library. When DYNAMIC_ARCH is 1, that directory is skipped and the library carries multiple kernel variants selected at runtime instead. That is the central design trade-off: a single-target build produces a smaller, more focused binary that only runs well on the CPU it was built for, while a dynamic-arch build is portable across a family of CPUs at the cost of size. The Makefile also shows how LAPACK is treated: if NO_LAPACK is not 1, lapack is appended to SUBDIRS, and if NO_FORTRAN is 1 the build exports C_LAPACK, which is the machine-translated LAPACK the README mentions as the fallback when no Fortran compiler is available. Threading is a build-time property, not a runtime afterthought. The README warns that after building, make install needs all the command line options repeated, because settings such as the supported maximum number of threads are derived from the build host by default. That is a real operational hazard: a library built on a 64-core machine and installed on a 4-core laptop will default to a thread count the laptop cannot use.
Installing OpenBLAS from source and from a package manager
The README points to binary packages for Windows x86, x86_64 and Windows arm64, hosted on sourceforge.net and in the GitHub Releases section, and notes that OpenBLAS is packaged for many package managers, with details in the installation section of the docs. For source builds you need GNU Make or CMake, a C compiler such as GCC or Clang, and optionally a Fortran compiler for LAPACK. The simplest build auto-detects the CPU and then installs. Run this from the top of the source tree, and expect the build to print a summary block listing OS, architecture, binary width, C compiler and version when it finishes:
make
make installTo install somewhere other than the default prefix, the README's optional section covers PREFIX; the same rule applies, so any option you gave to make must be repeated on the install line.
Pinning the target and the thread count
Automatic detection is convenient on the machine you build on and wrong everywhere else. To set a specific target CPU the README gives make TARGET=xxx with make TARGET=NEHALEM as the example, and it states that the full target list is in TargetList.txt. Consult that file rather than guessing at a name. A build for a modern CPU of the same architecture, for instance TARGET=SKYLAKEX on a HASWELL host, can use CROSS=1 to suppress the automatic test run at the end of the build:
make TARGET=NEHALEM
make TARGET=NEHALEM installThe thread count deserves the same treatment. The README is explicit that settings like the supported maximum number of threads are derived from the build host by default, so if you build on a large machine and deploy elsewhere, pass the value you want at build time and repeat it at install time. The search phrase openblas num threads and the question about openblas_num_threads both point at this same knob: it is a build parameter, not something you tune after the fact without rebuilding. A debug build is available with make DEBUG=1 when you need to chase a numerical or linkage problem.
Cross-compiling and the platforms with dedicated makefiles
Cross compilation is a first-class path, not an afterthought. The README instructs you to set CC and FC to the cross toolchains and, when using make, to set HOSTCC to your host C compiler, and it states that the target must be specified explicitly when cross compiling. Its example for a 64-bit MIPS router board looks like this:
make BINARY=64 CC=mipsisa64r6el-linux-gnuabi64-gcc FC=mipsisa64r6el-linux-gnuabi64-gfortran HOSTCC=gcc TARGET=P6600The repository layout backs this up: Makefile.alpha, Makefile.arm, Makefile.arm64, Makefile.csky, Makefile.e2k, Makefile.ia64, Makefile.loongarch64, Makefile.mips, Makefile.mips64, Makefile.power, Makefile.riscv64, Makefile.sparc, Makefile.wasm, Makefile.x86, Makefile.x86_64 and Makefile.zarch sit at the top level. The breadth here is unusual for a numerical library and is the main reason OpenBLAS turns up on embedded and non-x86 systems. The cost is that each architecture family has its own makefile to keep working, and the README does not document a rollback path if a build on a new target produces wrong results, which is why the test targets in the Makefile matter.
OpenBLAS versus MKL, BLIS and a reference BLAS
The comparison people search for most is openblas vs mkl. Intel MKL is a vendor library: closed source, tuned by Intel, distributed with a licence that restricts use on non-Intel hardware, and paired with a support channel. OpenBLAS is BSD-3-Clause, and the repository also carries GotoBLAS_00License.txt and the GotoBLAS readme and FAQ files, evidence of the GotoBLAS2 1.13 lineage the README names. The practical difference is who owns the performance work: with MKL you take Intel's tuning and its terms, with OpenBLAS you take a portable codebase you can rebuild for your target. BLIS is the closer architectural alternative. It is also an open-source BLAS implementation, but it is organized around a framework of configurable building blocks rather than OpenBLAS's per-architecture kernel makefiles, so the porting effort and the tuning surface differ. Against the netlib reference BLAS, the difference is simply that the reference exists to define behaviour; OpenBLAS exists to run fast. The README links to netlib for the routine documentation and for the LAPACK reference, which is the right place to look up what a routine computes regardless of which implementation you link.
Licence, upgrade cost and the questions to settle before adopting
The project is BSD-3-Clause, a permissive licence, and the GotoBLAS lineage files are kept in the tree. This is not legal advice, but the practical consequence for most teams is that linking OpenBLAS into a proprietary binary is the kind of use a permissive licence is designed to allow, subject to the licence text itself, which you should read rather than take from an article. Upgrade cost is where OpenBLAS differs from a typical dependency. Because TARGET and thread settings are baked in at build time, an upgrade is a rebuild, and the rebuild must reproduce your original options or you silently change behaviour. The README's warning about repeating make options at install time applies just as much to a version bump. The release cadence visible in the repository is frequent enough that a pinned version is the sane default for production. The first thing to verify on your own hardware is not a benchmark number someone else published: it is that your chosen TARGET appears in TargetList.txt, that make install received the same flags as make, and that the maximum thread count in the installed library matches the machine it will run on.
Editorial conclusion
Adopt OpenBLAS if you ship C, Fortran or Python code that spends its time in dense matrix kernels and you are willing to pin TARGET and NUM_THREADS at build time. Do not adopt it if you need a vendor-tuned math library with a support contract, or if you cannot rebuild when you move to a different CPU. Before anything else, check TargetList.txt for your microarchitecture and read the note in the README that make install must repeat the options you gave to make.
Frequently asked questions
What is OpenBLAS used for?
It provides optimized implementations of the BLAS routines and includes the higher-level LAPACK package, so it is used as the compiled numerical backend for programs that do dense linear algebra. The README describes it as an optimized BLAS library based on GotoBLAS2 1.13 BSD version.
How do I install OpenBLAS?
You can use a binary package (official ones exist for Windows x86, x86_64 and Windows arm64) or a package manager, and the README points to the installation section of the docs for the packaged options. From source, install GNU Make or CMake and a C compiler such as GCC or Clang, then build and install.
Is MKL faster than OpenBLAS?
The README and the linked documentation do not contain benchmark results for either library, so no comparison can be made here. What can be said is that MKL is a vendor library tied to Intel's tuning and licence terms, while OpenBLAS is BSD-3-Clause and is built for a target you choose.
What are the differences between BLAS and LAPACK?
BLAS covers the basic linear algebra subprograms, the low-level kernels such as matrix products and triangular solves. LAPACK is the higher-level package built on top of them, and it comes included with OpenBLAS; the README points to netlib for the reference documentation of both.
What is openblas_num_threads?
It relates to the maximum number of threads the built library supports. The README states that settings such as the supported maximum number of threads are derived from the build host by default, so the value is fixed by the build and install options rather than set freely at runtime.
Is OpenBLAS multithreaded?
Yes, but the thread support is a build-time property. The README warns that settings such as the supported maximum number of threads are derived from the build host by default, so you should pass the value you want when building and repeat those options when running make install.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openmathlib-openblas)