Library / SDK
simdjson/simdjson avatar
simdjson/simdjson

simdjson: a SIMD JSON parser for C++ that validates while it parses

Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks

24,310 stars1,291 forksC++Apache-2.0

At a glance

What is it?
simdjson targets C++ services that parse JSON on the hot path, using SIMD instructions and runtime CPU dispatch to reach gigabytes per second. The single-header build makes it easy to try; the on-demand API and the strict validation rules are where the real adoption cost sits.
Who is it for?
Adopt simdjson if your service parses JSON on the hot path and you can build with g++ 7 or clang++ 6 on a 64-bit target; the single-header build makes a pilot cheap. Do not adopt it if you need a DOM you can freely mutate, if your input is untrusted and you have not budgeted for the strict number and UTF-8 rules, or if a small scripting-language binding already covers your workload.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem simdjson solves, and who actually needs it

JSON is everywhere on the Internet, and servers spend a lot of time parsing it. The README states that simdjson parses JSON 4x faster than RapidJSON and 25x faster than JSON for Modern C++. Those numbers come from the project's own comparison against what it calls commonly used production-grade parsers, so treat them as the project's claim rather than a neutral measurement. The intended reader is a C++ engineer who owns a service that ingests JSON continuously: an analytics engine, a database front end, a log pipeline, or a runtime that embeds a JSON parser. The README's real-world usage list includes Node.js, ClickHouse, Meta Velox, Milvus, Apache Doris, StarRocks and Dgraph, which is the kind of company this library keeps.

If you parse a few small configuration files at startup, none of this applies to you. The library's value shows up when parsing is a measurable share of CPU time, and the README explicitly frames the problem that way: parsing gigabytes per second on commodity processors, millions of documents per second on a single core. Everything else in the design follows from that goal, including the parts that make it less convenient than a general-purpose JSON library.

How the SIMD approach and runtime dispatch work

The README describes the mechanism in one line: the library uses commonly available SIMD instructions and microparallel algorithms. Instead of walking the document byte by byte and branching on each character, the parser processes blocks of input with vector instructions, which is why the project claims three-quarters fewer instructions than RapidJSON. The same machinery is exposed for adjacent jobs: minifying JSON at 6 GB/s, validating UTF-8 at 13 GB/s, and NDJSON at 3.5 GB/s.

A second mechanism matters more for deployment. The library selects a CPU-tailored parser at runtime, with no configuration needed. The README points to doc/implementation-selection.md for the details. In practice this means you compile once and the binary picks the best available instruction set for the machine it lands on, rather than shipping separate builds per microarchitecture. The documentation also states that the library offers full UTF-8 validation and exact number parsing, so the speed does not come from skipping validation.

The parsing model is worth reading carefully before you commit. The quickstart uses ondemand::parser and ondemand::document, and the repository documentation index lists doc/basics.md for the API overview and doc/performance.md for tuning. The on-demand design defers work until you actually read a field, which is where much of the throughput comes from when you only touch part of a document. It also means the document is tied to the input buffer, a constraint that shapes how you write code around it.

Installing simdjson and parsing your first document

The README gives a single-header path that needs no package manager. The prerequisites are g++ version 7 or better, or clang++ version 6 or better, on a 64-bit system with a command-line shell such as Linux, macOS or freeBSD. Visual Studio and Xcode are supported but need different steps. The README notes that clang++ users may need to specify the C++ version, for example -std=c++17, because clang++ tends to default on C++98.

Fetch the two source files and a sample document with wget:

bash
wget https://raw.githubusercontent.com/simdjson/simdjson/master/singleheader/simdjson.h https://raw.githubusercontent.com/simdjson/simdjson/master/singleheader/simdjson.cpp https://raw.githubusercontent.com/simdjson/simdjson/master/jsonexamples/twitter.json

Then write the quickstart program exactly as the README presents it. It loads the file into a padded_string, iterates with an ondemand::parser, and reads a nested field:

cpp
#include <iostream>
#include "simdjson.h"
using namespace simdjson;
int main(void) {
    ondemand::parser parser;
    padded_string json = padded_string::load("twitter.json");
    ondemand::document tweets = parser.iterate(json);
    std::cout << uint64_t(tweets["search_metadata"]["count"]) << " results." << std::endl;
}

Compile it together with the implementation file and run it:

bash
c++ -o quickstart quickstart.cpp simdjson.cpp
./quickstart

The README states the output is 100 results. If you see that line, the SIMD path compiled and the runtime dispatch picked an implementation for your CPU. The README also warns that if you plan to use simdjson in a product, you should work from one of the releases rather than the master branch, so the wget URLs above are a convenience for a first look, not a distribution strategy.

Where simdjson is the wrong choice

The strictness is the first trap. The library performs full JSON and UTF-8 validation and exact number parsing, and the README presents that as a feature. It is one, until your input is not actually valid JSON. A parser that quietly coerces a malformed number or a broken UTF-8 sequence will let a pipeline keep running; simdjson will not. If your data source is a third-party feed with known encoding defects, you are choosing between cleaning the feed and handling parse errors on the hot path.

The second constraint is the parsing model. The quickstart's ondemand::document is built from a padded_string that must outlive it. The on-demand API is designed for reading fields, not for building a mutable tree you pass around your application. If your code wants to parse once, transform the structure, and serialize it back, you are working against the design. The README does document a builder in doc/builder.md for writing JSON strings, so the project is aware of the other half of the job, but the reading side is deliberately not a general-purpose DOM.

The third case is simply scale. If you parse a small request body per HTTP call and the rest of the handler dominates, a faster parser buys you nothing measurable. The README's own framing is gigabytes per second and millions of documents per second; below that, the cost of a new dependency and its build requirements outweighs the gain. And if your project is not C++, the library is not for you directly. The README links to bindings and ports, and there are related searches for simdjson python, simdjson rust, simdjson java, simdjson golang, simdjson nodejs and simdjson c#, which tells you people look for those, but the repository you are reading is the C++ implementation.

simdjson vs RapidJSON and the other parsers people compare it to

The README makes one direct comparison: simdjson parses 4x faster than RapidJSON and uses three-quarters fewer instructions. RapidJSON is the reference point because it is a mature, widely used C++ parser with a DOM-style API. The difference in approach is structural. RapidJSON gives you a document object you can walk and modify; simdjson gives you an on-demand reader over a buffer you already hold. If your code is written around a mutable document, switching to simdjson is a rewrite of the parsing layer, not a drop-in replacement.

The same logic applies to the other comparisons people search for. Against boost json, the question is whether you want a Boost dependency and its conventions or a standalone pair of files. Against serde_json or simdjson vs jackson, you are comparing across languages, and the honest answer is that those libraries solve the problem in their own runtimes with their own idioms; simdjson's advantage only exists inside a C++ process. Against protobuf, the comparison is category error: protobuf is a schema-driven binary format, and simdjson parses text JSON. The README does not mention protobuf, Jackson, serde_json, boost json or serde, so any detailed claim about those comparisons would be invented. What the README does support is the RapidJSON number and the instruction-count claim, and those are the ones worth quoting.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-20. Releases are frequent: v4.6.11 on 2026-09-05, v4.6.10 on 2026-09-03, and v4.6.9 on 2026-08-27. That cadence cuts both ways. You get fixes quickly, and you also get a moving target if you track master. The README's own advice is to work from one of the releases in a product, which is the right default here. Pin a tag, and read the release notes before moving, because a parser that sits under your data path is not a dependency you upgrade casually.

The upgrade cost is mostly build-side. The single-header pair means there is no link-time dependency to manage, but it also means the simdjson.cpp file is compiled into your build, so a version bump touches your compile times. The runtime CPU dispatch is the part that keeps upgrades cheap: you do not maintain per-architecture builds. On the licence, the repository carries Apache-2.0, and the top-level entries include both LICENSE and LICENSE-MIT. The README's badge section references a license and a licensemit badge, which is consistent with an Apache-2.0 project that also offers an MIT option for the single-header files. If licence terms matter to your legal review, read the actual LICENSE and LICENSE-MIT files rather than inferring from badges; this is not legal advice, and the files are the authority.

Editorial conclusion

Adopt simdjson if your service parses JSON on the hot path and you can build with g++ 7 or clang++ 6 on a 64-bit target; the single-header build makes a pilot cheap. Do not adopt it if you need a DOM you can freely mutate, if your input is untrusted and you have not budgeted for the strict number and UTF-8 rules, or if a small scripting-language binding already covers your workload. Verify first that your compiler version matches the prerequisites, that you are building from a release rather than master, and that the on-demand API's document-lifetime rules fit how your code holds parsed values.

Frequently asked questions

How does simdjson work?

It uses commonly available SIMD instructions and microparallel algorithms to process JSON in blocks rather than byte by byte, and it selects a CPU-tailored parser at runtime with no configuration. The README states it performs full JSON and UTF-8 validation and exact number parsing.

How is simdjson so fast?

The README attributes the speed to SIMD instructions and microparallel algorithms, and states the library uses three-quarters fewer instructions than RapidJSON. It also offers full UTF-8 validation and exact number parsing, so the gain is not from skipping validation.

What are the key differences between RapidJSON and simdjson?

The README states simdjson parses 4x faster than RapidJSON and uses three-quarters fewer instructions. The API model differs too: simdjson's quickstart uses an ondemand::document read from a padded_string buffer, while RapidJSON is a DOM-style parser. The README does not document a migration path from one to the other.

What is the fastest C++ JSON library?

The README claims simdjson parses 4x faster than RapidJSON and 25x faster than JSON for Modern C++, and says it is, to the project's knowledge, the first fully-validating JSON parser to run at gigabytes per second on commodity processors. That is the project's own claim, not an independent benchmark.

How do I install simdjson?

The README's quickstart pulls simdjson.h and simdjson.cpp from the singleheader directory along with a sample JSON file, then compiles them together with your program using c++ -o quickstart quickstart.cpp simdjson.cpp. Prerequisites are g++ version 7 or better or clang++ version 6 or better on a 64-bit system.

What is simdjson?

It is a C++ JSON parsing library that uses SIMD instructions and runtime CPU dispatch, distributed as a single .h and .cpp pair. The README lists Node.js, ClickHouse, Meta Velox, Milvus, Apache Doris and StarRocks among its real-world users.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. simdjson/simdjson on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/simdjson-simdjson.svg)](https://hysenlabs.com/projects/simdjson-simdjson)