kostya/benchmarks: A Cross-Language Benchmark Suite Built Around Ordinary Code
Some benchmarks of different languages
At a glance
- What is it?
- kostya/benchmarks compares C, Rust, Go, the JVM languages, Julia and more on five small test cases, scoring time, memory and CPU energy. Its rules about writing code the way an average developer would are what separate it from synthetic microbenchmarks.
- Who is it for?
- Adopt kostya/benchmarks if you need a defensible, reproducible comparison of language runtimes on small idiomatic workloads, or if you are choosing a language for a CLI tool where the whole process is short-lived. Do not adopt it for throughput or concurrency planning: the test cases are single-threaded and tiny, and the README does not document any rollback path for a bad measurement run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 75 days ago.
- What is it written in?
- Mainly Makefile, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the suite measures and the rules it imposes on submissions
This repository is a set of five benchmark programs, each reimplemented in many programming languages, plus the Makefile and analysis scripts that build and run them. It exists for people who have to pick a language or runtime and want numbers that reflect ordinary code rather than hand-tuned kernels. The README states the criteria plainly: implementations are written "as the average software developer would write them", algorithms follow public sources, libraries are used as their tutorials and examples describe, and data structures are idiomatic. Variants are allowed when a reference implementation exists. All final binaries are releases, optimized where possible, because the README notes that debug performance varies too much depending on the compiler.
That constraint is the whole point. A language comparison where one entry has been vectorized by an expert and another has not tells you about the authors, not the languages. The trade-off is that a suite built from idiomatic code will not reveal what a runtime can do at its ceiling, and the README does not claim otherwise.
The five test cases: Brainfuck, Base64, Json, Matmul and Primes
The workload set is deliberately small. Brainfuck runs two interpreters, bench.b and mandel.b, over two code samples. The README describes two modes: verbose, which prints output immediately, and quiet, which accumulates output through a Fletcher-16 checksum and prints it after the benchmark. That quiet mode matters because it removes I/O from the measurement, and the checksum keeps the output verifiable. Base64, Json, Matmul and Primes cover string encoding, parsing, floating-point arithmetic and integer work respectively.
None of these test concurrency, network behavior, or long-running memory pressure. Each is a short, self-contained program. The README also states that loading required data and code self-testing are excluded from the measured time, so what you see is the benchmark body, not startup plumbing. If your real workload is a request handler that lives for days, this suite measures a different shape of problem.
How time, memory and energy are recorded
Three quantities are reported. Time is the benchmark execution itself. Memory is reported as base plus increase, where base is the RSS before the benchmark and increase is the peak RSS growth during it. This split matters for runtimes with large fixed footprints: the Brainfuck table shows Scala at 66.03 MiB base plus 196.44 MiB increase, while C/gcc sits at 1.38 MiB base with no measurable increase. A single RSS number would hide that distinction.
Energy is the third column: PP0 (cores) plus PP1 (uncores such as GPU) plus DRAM, read through the powercap interface. The README is explicit that only Intel CPUs are supported for this. On AMD or ARM hardware the energy column is not available, and the README does not describe a fallback. All values are presented as median plus or minus median absolute deviation, which is a sensible choice for noisy short runs. The README records an update dated 2026-07-19, and the repository's last push was on 2026-07-19.
Installing and running your first benchmark
The README offers two paths: Docker, or manual execution with prerequisites installed. Docker is the shorter route because the toolchains for a dozen languages are the expensive part. The repository has a docker/ directory and a run.sh at the top level.
A typical Docker invocation clones the repository and runs the script inside the container:
git clone https://github.com/kostya/benchmarks.git
cd benchmarks
./run.shThe README's Using Docker section is where the exact image and flags belong; check it before running, since the script's defaults are not restated in the README's overview. For manual execution, the README lists prerequisites per language, and the Makefile drives the builds. Individual test cases have their own directories, so you can build one language in one case rather than everything:
cd brainfuck
makeAfter a run you should see the median and median absolute deviation for time, memory and energy, in the format the README describes. If you set the QUIET environment variable for the Brainfuck cases, output is replaced by a Fletcher-16 checksum printed at the end, which is the signal that quiet mode is active.
Where the suite misleads you
The most concrete limitation is hardware scope for energy. The README says energy measurement currently only supports Intel CPUs via powercap. On anything else, one of the three reported dimensions is simply absent, and there is no documented substitute. A second limitation is scale: these are small, single-threaded programs. A runtime with fast startup and a runtime with excellent steady-state throughput are not distinguished by a workload that finishes in about a second. The Brainfuck table shows this clearly: C# (Staged)/.NET Core leads at 0.373 s, ahead of V/gcc at 0.985 s and Rust at 1.008 s, but that ordering says more about staged compilation of a tiny interpreter than about general runtime speed.
The third limitation is interpretive. The README defines the coding criteria but does not publish a scoring formula that combines time, memory and energy into one ranking. You have to decide which column matters for your case. And because implementations are meant to look like ordinary developer code, a poor result may reflect an unidiomatic port rather than a slow language; the README's variant rule is an attempt to control for this, not a guarantee.
How it differs from jit-benchmarks and the benchmarks game
The same author maintains jit-benchmarks and crystal-benchmarks-game, both linked from the README, along with LangArena. The benchmarks game family takes a different approach: its programs are typically written and tuned by enthusiasts, and the emphasis falls on peak achievable performance. That produces useful upper bounds but makes it hard to reason about what a normal codebase will do. This suite inverts the priority, accepting lower absolute numbers in exchange for implementations that a working developer would recognize.
jit-benchmarks, by name, concentrates on just-in-time compilation behavior, which is a narrower question than the one here. If your interest is specifically in how a JIT warms up, that repository is the closer fit. If your interest is in whether a language is fast enough for a small tool, this one is. The two answer different questions, and the README treats them as companion projects rather than substitutes.
Maintenance, contribution and licence
The repository is not archived, and the last push was on 2026-07-19, so it is current as of that date. There are no releases: the README carries an UPDATE line with a date instead of a version number, which means the results table is the artifact, not a tagged package. That has an upgrade implication. If you cite a number, cite it with the update date, because the table is edited in place and there is no changelog of past tables in the repository.
Contributing is structured. The README has a Makefile guide covering binary executables, compiled artifacts and scripting languages, plus separate notes on README update and Docker image update. Adding a language means touching the Makefile in the documented way and regenerating the README table, which the README explains. The project is MIT licensed, which permits reuse and modification with attribution; the LICENSE file is at the repository root. This is a description of the licence terms, not legal advice, and anyone embedding the code in a product should read the LICENSE text directly.
Editorial conclusion
Adopt kostya/benchmarks if you need a defensible, reproducible comparison of language runtimes on small idiomatic workloads, or if you are choosing a language for a CLI tool where the whole process is short-lived. Do not adopt it for throughput or concurrency planning: the test cases are single-threaded and tiny, and the README does not document any rollback path for a bad measurement run. Before trusting a number, verify the CPU matches the Intel powercap requirement for the energy column, and check the Makefile target for the language you care about actually builds against released binaries rather than debug builds.
Frequently asked questions
What is kostya/benchmarks and what is it used for?
It is a collection of five benchmark programs (Brainfuck, Base64, Json, Matmul and Primes) implemented in many languages, used to compare execution time, memory consumption and CPU energy consumption across runtimes. The README states the implementations are written the way an average software developer would write them, so the results reflect ordinary code rather than hand-tuned kernels.
How do I set up and run kostya/benchmarks?
The README gives two routes: Docker, or manual execution with the per-language prerequisites installed. The Docker path uses the repository's docker/ directory and the top-level run.sh, while manual runs go through the Makefile in each test case directory such as brainfuck/ or json/.
What does the energy column in kostya/benchmarks measure, and does it work on my CPU?
Energy is the CPU package consumption during the benchmark: PP0 for cores, PP1 for uncores such as the GPU, and DRAM, read through the powercap interface. The README states that currently only Intel CPUs are supported for this measurement.
Does kostya/benchmarks measure concurrency or multi-threaded performance?
The README does not describe the test cases as concurrent or multi-threaded workloads. They are small, self-contained programs, and the reported values are wall-clock time, RSS memory and CPU energy for a single benchmark execution.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kostya-benchmarks)