Model or dataset
jamesob/local-llm avatar
jamesob/local-llm

jamesob/local-llm: A Hardware BOM and Docker Runners for Local SOTA Models

Everything I know about running LLMs locally

1,854 stars106 forksShellLicense varies

At a glance

What is it?
This repository is one person's build log for running near-frontier models at home, from a $2k two-GPU box to a $40k four-GPU machine. It is a hardware guide first and a software kit second, and the README says so.
Who is it for?
Adopt it if you are planning a multi-GPU local inference box and want a real parts list with prices, or if you already own four RTX PRO 6000 cards and want a starting vLLM compose file. Skip it if you want a one-command installer or a single-GPU laptop setup; the README points at jamesob/stt for the speech side but gives no equivalent path for the cheap end.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 82 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What jamesob/local-llm actually is, and who it is written for

This is not a framework and not a CLI. The repository holds a README, an images/ directory, a runners/ directory and a tools/ directory. The README opens by describing itself as "Everything I know about running LLMs locally," and it is written in the first person about one machine its author built. That framing matters, because the value here is the reasoning behind a parts list, not an abstraction you import into your own code.

The target reader has a budget and a decision to make. The README splits that decision into roughly $2k, $20k and $40k tiers, and it is honest about which ones it can help with. The $20k section is marked TODO and the author writes that he has no experience in that regime. The $2k tier is two RTX 3090s for 48GB of VRAM total, which the README says runs Qwen3.6-27B and SOTA speech-to-text through cohere-transcribe. The $40k tier is four RTX PRO 6000 Blackwell cards, 384GB of VRAM, and the README claims that gets you "pretty close to Claude Opus." That claim is the author's judgement, not a measured result.

So the audience is narrow and specific: people building a multi-GPU inference host who want to know where the money should go. The README's own answer is that it should go into VRAM, not into a current-generation host platform.

The PCIe switch argument: spend on VRAM, not on a Gen5 host

The central design decision in this build is to pair current GPUs with a last-generation host. The base system is an ASRock Rack ROMED8-2T with an EPYC Milan 7313P and 128GB of DDR4 ECC bought on eBay, totalling $5,687 including case, PSUs, storage and a PCIe switch. The four GPUs cost about $46,000. The README states that staying on Gen4 rather than Gen5 saves roughly $10,000 in host costs.

That only works if the GPUs can talk to each other without going through the CPU. Hence the c-payne Microchip Switchtec PM40100 Gen4 switch, roughly $1,330 with its host adapter and SlimSAS cables. The README's explanation is that tensor parallelism performs an allreduce step, and without a switch that traffic crosses the PCI root complex. With the switch, the cards communicate peer-to-peer inside the switch fabric at wire speed, which the README says reduces latency and removes the need for PCIe5 hardware.

The stated result is Gen4 line rate at 27.5/50.4 GB/s with sub-microsecond latency. Two caveats are worth stating plainly. First, this is the author's own measurement, and the README does not describe the methodology. Second, the repository ships tools/measure-gpu-speed.sh, a P2P bandwidth and latency benchmark, so you can produce your own number rather than trusting that one. The README also warns that spilling model weights into system RAM makes LLM performance unusably slow for agentic workloads, which is the reason the whole build optimises for VRAM capacity.

Getting the PCIe switches to work: bifurcation, link speed, ASPM and ACS

The hardware section is the part most likely to save you a weekend, because it documents the BIOS and kernel settings that a switch-based topology needs. The README lists BIOS bifurcation, link speed and ASPM as the settings that make the switch behave properly. It gives a kernel and GRUB parameter of iommu=off, with the note that otherwise NCCL hangs. It also calls ACS disable "critical for switch P2P," so that peer-to-peer traffic stays inside the switch fabric rather than being routed up to the root complex.

These are not optional tweaks. A machine that boots and enumerates four GPUs can still be functionally broken for tensor parallelism if ACS is left enabled or the IOMMU is on, and the failure mode is a hang rather than a clear error. If you are evaluating this repository before buying a switch, that single iommu=off line is the most transferable piece of information in it.

The build also has physical requirements the README does not soften. The GPU mount is a custom wooden enclosure that took about a day to fabricate. The PCI switch's built-in fan was loud and, in the author's words, seemingly useless, so he unplugged it from the board. Power limiting is covered as its own topic because the machine runs $46,000 of silicon on a 110V circuit. None of this is turnkey, and the README never presents it as such.

Installing the runners: from clone to a vLLM compose file

There is no install command for the repository itself. It is a collection of guides and configuration files, so the setup step is cloning it and reading the runner you intend to use. The README points at runners/ for ready-to-run serving configs and states that the GLM-5.2-594B runner is a vLLM docker-compose setup using DCP4+MTP5, which the README reports at roughly 80 t/s at 460k context.

Start by getting the repository onto the machine that will serve the model.

bash
git clone https://github.com/jamesob/local-llm
cd local-llm
ls runners/

The listing should show GLM-5.2-594B and stt. The stt runner is the cheaper entry point: the README says it assumes only about 11GB of VRAM on an Nvidia GPU and uses cohere-transcribe. The GLM runner is the opposite end, built for four RTX PRO 6000 cards.

Before trusting the interconnect on a new build, run the benchmark that ships with the repository.

bash
./tools/measure-gpu-speed.sh

The README describes this script as a P2P bandwidth and latency benchmark. On the author's machine it corresponds to the 27.5/50.4 GB/s Gen4 line-rate figure. If your numbers are far below that, the BIOS and kernel settings in the previous section are the first place to look, not the model configuration.

For the serving config itself, the README directs you to the runner directory rather than reproducing the compose file in the README. Read runners/GLM-5.2-594B before launching anything, because the DCP4 and MTP5 settings are tied to a four-GPU topology and will not behave the same on fewer cards.

Where this guide stops being useful

The clearest limitation is the one the README states itself: the $20k tier is a TODO, and the author writes that he is not of much help there. If your budget lands between the two extremes, the repository gives you a list of ideas (two RTX 6000 Pro cards, four networked DGX Sparks, Apple hardware) and no parts list, no prices and no configuration.

The second limitation is scope. This is a hardware and serving guide. It does not cover fine-tuning, quantisation workflows, evaluation, or how to point an editor or coding agent at the endpoint. Anyone arriving with a question about wiring a local model into an IDE will not find an answer here.

The third is that the numbers are single-source and unverified by the repository. The throughput figure, the context length and the bandwidth numbers all come from one machine described in prose. There is no test harness that reproduces the throughput claim, only the P2P benchmark. Treat 80 t/s at 460k context as a report from one build, not a specification.

Finally, the licence is not stated in the repository metadata. That matters if you intend to reuse the configuration files in anything commercial, and it is a question to resolve before you build on them rather than after.

How it differs from llama.cpp and Ollama-style setups

The obvious alternative for most readers is a llama.cpp or Ollama-style setup: one installer, a model pulled by name, a single GPU or none, and a chat interface within minutes. That approach optimises for time-to-first-token on modest hardware. This repository optimises for the opposite: maximum VRAM and interconnect bandwidth so that a very large model runs with tensor parallelism across four cards.

The difference in approach shows up in what each one assumes about your machine. A llama.cpp setup assumes one accelerator and lets the CPU or system RAM absorb the overflow. This README explicitly rejects that trade, stating that spilling into system RAM makes performance unusably slow for agentic workloads. That is why the build spends on a PCIe switch instead of accepting slower inter-GPU traffic.

The second difference is the serving stack. The runners here are vLLM docker-compose configurations. That is a server you operate, with a topology-specific parallel configuration, not an application you launch. If you want a desktop chat app, the alternative is the right choice and this repository is the wrong one. If you want to serve a 594B model across four GPUs and care about allreduce latency, the alternative does not address the problem this repository was written to solve.

Editorial conclusion

Adopt it if you are planning a multi-GPU local inference box and want a real parts list with prices, or if you already own four RTX PRO 6000 cards and want a starting vLLM compose file. Skip it if you want a one-command installer or a single-GPU laptop setup; the README points at jamesob/stt for the speech side but gives no equivalent path for the cheap end. Before spending anything, read runners/GLM-5.2-594B and tools/measure-gpu-speed.sh, and confirm your motherboard exposes the PCIe bifurcation and ACS settings the README describes, because those are the parts that decide whether peer-to-peer traffic stays inside the switch.

Frequently asked questions

What does local LLM mean?

It means running a large language model on hardware you own rather than calling a hosted API. This repository is about the hardware and serving configuration needed to do that at a high level of model quality.

Is it worth running a local LLM?

The README frames the question around budget: roughly $2k buys two RTX 3090s and 48GB of VRAM for Qwen3.6-27B plus speech-to-text, while roughly $40k buys four RTX PRO 6000 cards and 384GB of VRAM for something the author describes as close to Claude Opus. It also notes that spilling into system RAM makes performance unusably slow for agentic workloads, so the answer depends on whether you can afford enough VRAM.

What's the best LLM I can run locally with jamesob/local-llm?

The README's table lists GLM-5.2-Int8Mix-NVFP4-REAP-594B as the best model for four RTX 6000 Pro cards as of 2026-07-01, with a runner config in runners/GLM-5.2-594B. At the $2k tier it names Qwen3.6-27B and cohere-transcribe for speech-to-text.

How do I install jamesob/local-llm?

There is no installer. The repository is a README plus runners/ and tools/ directories, so you clone it and read the runner you intend to use. The GLM-5.2-594B runner is a vLLM docker-compose configuration, and tools/measure-gpu-speed.sh is a shell script you run directly.

Does jamesob/local-llm explain how to use a local LLM for coding?

No. The README covers hardware selection, speech-to-text, and serving configurations for large models, but it does not describe connecting a local endpoint to a coding tool or editor. That ground is not covered in the repository.

Official sources

  1. Issues
  2. jamesob/local-llm on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jamesob-local-llm.svg)](https://hysenlabs.com/projects/jamesob-local-llm)