# llama3.java: Running Llama 3 Models from a Single Java File

> llama3.java implements Llama 3, 3.1, and 3.2 inference in a single Java file with no runtime dependencies, targeting the JVM for compiler research and practical LLM experimentation.

**mukel/llama3.java** — Llama 3+ inference in pure Java

- Repository: https://github.com/mukel/llama3.java
- Stars: 818 · Forks: 96
- Language: Java
- License: MIT
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/mukel-llama3-java

## What llama3.java Does and Who It Is For

llama3.java is a self-contained implementation of Llama 3 family inference in a single .java file. The README describes its primary purpose as twofold: educational value (demonstrating that LLM inference is achievable in plain Java) and a platform for testing JVM compiler optimizations, specifically targeting the Graal compiler.

The immediate audience is Java engineers who want to run Llama 3, 3.1, or 3.2 models locally. The project is the successor to llama2.java and is itself inspired by Andrej Karpathy's llama2.c. That lineage means the code priorities are clarity, portability, and JVM-native idioms over maximum raw throughput.

By constraining the implementation to a single file with no declared external dependencies, the project makes auditing and modification tractable. The entire inference path from GGUF parsing to token generation is readable in one file, which is a genuine advantage for anyone learning how transformer inference works at the implementation level.

## GGUF Format Parsing and Quantization Support

The project reads models in the GGUF format, the binary model container format used by llama.cpp and compatible tools. The tokenizer is based on the minbpe byte-pair encoding approach.

Supported weight formats are F32, F16, and BF16 at full precision, plus the Q4_0, Q4_1, Q4_K, Q5_K, Q6_K, and Q8_0 quantization types. The README flags a subtle point about quantization purity: in the wild, models labeled Q4_0 often use a higher-precision quantization such as Q6_K for the token embedding and output weight tensors rather than pure Q4_0 throughout. The README states that a pure quantization, where every tensor uses the same type, can be generated from a high-precision source using the llama-quantize utility from llama.cpp:

```bash
./llama-quantize --pure ./Meta-Llama-3-8B-Instruct-BF16.gguf ./Meta-Llama-3-8B-Instruct-Q4_0.gguf Q4_0
```

Pure quantizations matter here because llama3.java's Vector API acceleration paths are optimized for uniform tensor types. Mixed-quantization models will still load but may not hit the same throughput.

The project also implements Llama 3.1's ad-hoc RoPE scaling and Llama 3.2's tied word embeddings, so the GGUF compatibility extends across the current Llama 3 family.

## Building and Running llama3.java

Java 21 or higher is required. The 21 floor exists because the project uses the MemorySegment API for memory-mapped GGUF file loading, which reached stable status in Java 21.

The simplest way to run the project is via jbang, which handles compilation transparently:

```bash
jbang Llama3.java --help
```

Or make the file executable and run it directly:

```bash
chmod +x Llama3.java
./Llama3.java --help
```

For a more conventional deployment, the Makefile target `make jar` produces llama3.jar:

```bash
java --enable-preview --add-modules jdk.incubator.vector -jar llama3.jar --help
```

Models must be downloaded separately. The README points to Hugging Face repositories maintained by the author (mukel/Llama-3.2-1B-Instruct-GGUF, mukel/Llama-3.2-3B-Instruct-GGUF, mukel/Meta-Llama-3.1-8B-Instruct-GGUF, and mukel/Meta-Llama-3-8B-Instruct-GGUF) as well as unsloth-converted versions. The CLI exposes `--chat` for conversational use and `--instruct` for single-turn instruction following.

## GraalVM Native Image and AOT Model Preloading

Beyond the standard JVM path, llama3.java supports compilation to a native binary via GraalVM Native Image. Running `make native` produces a self-contained `llama3` executable:

```bash
./llama3 --model Llama-3.2-1B-Instruct-Q8_0.gguf --chat
```

The native image eliminates the JVM startup overhead and reduces the deployment footprint to a single binary with no JRE dependency on the target host.

The project also supports AOT model preloading, which embeds a specific GGUF model into the native binary at compile time. This approach trades binary size for zero parsing overhead at runtime:

```bash
PRELOAD_GGUF=/path/to/model.gguf make native
```

The resulting binary has effectively zero time-to-first-token for the preloaded model because the weights are already parsed and memory-mapped before execution begins. The README notes that a preloaded binary can still run other models, but they will incur the usual parsing overhead. This design is suitable for embedded or edge deployments where startup latency matters and the model is fixed.

## Performance, the Vector API, and GraalVM Tuning

llama3.java accelerates matrix-vector multiplications using Java's Vector API (JEP 469). The Vector API allows the JIT to emit SIMD instructions matching the host hardware's vector width. The README states that the preferred vector size is selected automatically but can be overridden with the JVM flag `-Dllama.VectorBitSize=0|128|256|512`, where 0 disables SIMD entirely.

The README recommends GraalVM 25 or later for best performance, citing partial but good Vector API support in the Graal compiler. The README includes benchmark commands comparing llama3.java against vanilla llama.cpp on a 2019 AMD Ryzen 3950X 16-core system. The benchmark runs the model pinned to one NUMA die using `taskset -c 0-15`, because the README notes that inference throughput is limited by memory bandwidth rather than compute, and crossing NUMA nodes introduces memory latency.

The project is one of several in the same author's family of Java LLM inference repositories, alongside gemma4.java, gptoss.java, and qwen35.java. Each handles a specific model family, and they share the same single-file, no-dependency design philosophy.

## Limitations and Cases Where llama3.java Is the Wrong Tool

The CLI provides only `--chat` and `--instruct` modes. There is no HTTP server, no OpenAI-compatible API endpoint, and no batching support. For production serving of inference requests from multiple clients, llama.cpp's server mode, Ollama, or vLLM are better-matched tools. llama.cpp in particular covers the same GGUF format and quantization types, but runs natively in C++ with a more complete feature set for serving.

The Vector API dependency is incubating as of Java 21; the `--add-modules jdk.incubator.vector` flag must be passed explicitly at runtime. If your organization restricts JVM module flags for security or compliance reasons, this may be a blocking issue.

The project targets x86-64 Linux in its benchmarks. ARM hardware, including Apple Silicon Macs and Raspberry Pi devices, is not mentioned in the README, so performance and correctness on those platforms are unconfirmed by the project's own documentation.

The last push to the repository was on 2026-04-24. The README does not describe a release cadence or a timeline for supporting future Llama model generations.

## Conclusion

llama3.java suits Java engineers who want to run Llama 3 models locally without leaving the JVM, researchers experimenting with compiler optimizations on the Graal backend, and developers who need a self-contained binary with minimal build overhead. It is the wrong choice for anyone targeting ARM mobile hardware, requiring batched inference, or expecting a feature-complete inference server. Java 21 is a hard prerequisite; older JDKs will not work because the code depends on the MemorySegment mmap API added in that release. Start with the Q4_0 quantization of the 1B or 3B model from the author's Hugging Face repositories and verify your JDK version before building.

## FAQ

### What Java version does llama3.java require?

The README states Java 21 or higher is required. The version floor exists because the project uses the MemorySegment mmap API that was stabilized in Java 21.

### Does llama3.java support GGUF models from llama.cpp?

Yes. The project includes a GGUF format parser and supports F16, BF16, F32, and several quantized types including Q4_0, Q4_K, Q5_K, Q6_K, and Q8_0.

### Can llama3.java compile to a native executable?

Yes. The README describes GraalVM Native Image support via the `make native` target, which produces a self-contained binary with no JRE dependency on the target host.

## Sources

- [Issues](https://github.com/mukel/llama3.java/issues)
- [License: MIT](https://github.com/mukel/llama3.java/blob/main/LICENSE)
- [mukel/llama3.java on GitHub](https://github.com/mukel/llama3.java)
- [README](https://github.com/mukel/llama3.java/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mukel-llama3-java
