Model or dataset
Nehanth/swarmllm avatar
Nehanth/swarmllm

swarmllm ships as pooled: a browser LLM pool split across the tabs in a room

Every device brings a slice. Together they run the whole model. Peer-to-peer LLM inference across browser tabs: a from-scratch WebGPU engine and a WebRTC runtime that split a 27B model over the devices in a room.

569 stars88 forksJavaScriptMIT

At a glance

What is it?
Nehanth/swarmllm is the repo name; the package, the site and the CLI all say pooled. Underneath sits a WebGPU inference engine that splits a 27B model into contiguous layer bands across browser tabs, plus an OpenAI and Anthropic compatible bridge that other agents can drive. The naming split, the version skew between tabs and host, and the room-wide memory floor are the parts worth reading before you open a room.
Who is it for?
Try Pooled when you have two or three devices with spare memory and a use for a private local model, especially if you want Codex CLI, Claude Code or opencode pointed at a machine that is not yours.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The repository is Nehanth/swarmllm but the package is pooled

The repository you cloned is Nehanth/swarmllm, and almost nothing inside it uses that name. package.json sets the package name to pooled, gives its homepage as https://pooled.run, and lists its own repository field as https://github.com/Nehanth/pooled.git. The local setup lines clone Nehanth/pooled. The wordmark at the top of the page links to pooled.run, and both demo videos, the September 27 recording and the earlier September 7 one, are downloaded from the pooled releases. The project listing gives swarmllm.ai as the homepage instead. The same package.json also carries private set to true, so pooled itself is never published to a registry; the piece meant to be installed is @pooled/cli, which the serve instructions invoke through npx. Two names, one code base, and the one a visitor searches for is the one that appears least.

Every token crosses the tab boundary as a hidden state

For each generated token the host embeds the previous token into a hidden state and passes it onward. The diagram in the readme splits a 27B into contiguous bands: the laptop holds layers 0 to 21, the desktop 22 to 42, the phone 43 to 63, and each device runs its band on its own GPU using the project's own WGSL kernels rather than a library. The payload is small and fixed size, 10 KB of hidden state on the 27B and 4 KB on the MoE. When the band comes back, the host does the final norm, the LM head and the sample, and at the same time drafts the tokens that follow. Speculative decoding, batched prefill and each kernel trick are held against golden tests, so the speculative stream is the same as plain decoding. The transport underneath is direct WebRTC, which uses UDP, and work networks often block UDP, so a site with a relay configured falls back to TCP or TLS on port 443. Links still go direct when they can, and model weights never go through the relay.

The room has to hold 17 or 22.5 GB of weights at once

Pooling does not make memory disappear, it divides it, and the division has to land. Qwen 3.8 27B at Q4_0 needs roughly 17 GB across the whole room. Qwen 3.6 35B MoE at Q4_0 needs roughly 22.5 GB, and it carries 256 experts with 8 active per token, which is why the project says it decodes several times faster than the 27B despite being larger on disk. Qwen3 1.7B at Q8_0 needs about 5.6 GB across the room for 16K context, or about 4 GB for 8K, and it is the row aimed at rooms of phones and light laptops. Context ceilings scale the same way: 16K for the 27B and up to 64K, 32K for the MoE and up to 128K. The one number that moves on its own is the 1.7B window, which drops to 8K when the room is short of memory, so the context you get depends on which devices are still connected rather than on a setting chosen up front.

Layers come from Hugging Face or from the device that already has them

A joining device downloads only its own band, either from Hugging Face or from another device in the room that already holds those layers, and keeps the result cached for next time. That second source is the part that changes the cold start: the second laptop to arrive is not re-fetching what the first laptop already has in its cache. Departures are handled by announcement and one click. When a device leaves, the room says so, and the host deals its layers out again with a single click, which is a manual step rather than an automatic rebalance. So the sequence around a dropout is: the room announces the departure, the host assigns the orphaned layers somewhere, and the participants wait. If a phone drops mid answer, the cost is not a corrupted stream but a host who has to act before decoding continues.

The bridge holds no layers and the room answers one request at a time

A room can serve its model to anything speaking the OpenAI Chat Completions, OpenAI Responses or Anthropic Messages API, tool calling included, which covers Codex CLI, Claude Code, opencode, Continue, Open WebUI, LiteLLM, the openai and anthropic SDKs and curl. Inside the room, the Serve API button opens the page carrying the command for that room:

bash
npx @pooled/cli serve "https://pooled.run/r/4TKG9P#k=…"   # the room's invite link, or its code
# pooled serve · room 4TKG9P · Qwen3.6 35B MoE · Q4 · 32768 tokens of context
#   OpenAI     http://127.0.0.1:8080/v1         (OPENAI_BASE_URL, any API key: chat/completions, responses)
#   Anthropic  http://127.0.0.1:8080            (ANTHROPIC_BASE_URL: messages)

curl http://127.0.0.1:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model": "pooled", "messages": [{"role": "user", "content": "Hi"}], "stream": true}'

That machine needs Node 22 or newer and no GPU of its own, because the bridge joins the room as an API client with no layers. Requests run on the room's GPUs one at a time in the room's queue, which is the ceiling an agent-driven workload feels first. The listener binds to 127.0.0.1 only, and POOLED_TOKEN makes it require a key. Two endpoints are deliberately missing: images come through as a note, since the models read text only, and the legacy /v1/completions path is not served at all.

An older host answers every tool call with a 400

A room is a federation of independently loaded tabs, so version skew is structural rather than incidental. The bridge side is thorough: tool calls stream, run in parallel and cover every tool_choice form, tool results work, JSON mode and JSON schemas work, reasoning works, and the Responses previous_response_id works, on all three APIs. Every call and every structured answer is grammar-constrained to its schema, so calls parse even when they come from the 1.7B. The constraint is on the host: it has to be a Pooled build that has tool calling, and older hosts answer a tool request with a 400 that tells the client to reload. The message lands on the caller, so a room can hold up to date guests and a host that still answers tools that way. Prompts travel the same social graph: they reach the host and, unless the host limits who sees answers, everyone in the room.

Code mode waits for the host before touching a folder on disk

Switching the room to Code mode starts a small agent loop: the room's model writes files, serves them on a virtual localhost on port 5173 inside a sandboxed preview, reads its own console errors and fixes them. Everyone in the room sees the files and can run the same preview on their own screen. The storage choice is where the boundary sits. Files live either in the host's browser or in a folder on disk that the host picks, and edits to a real folder wait for the host's approval. Nothing about the approval is described in the room itself: it is the host's action, taken after the model has already proposed the write, which means the host is reviewing generated code rather than approving a plan. What the agent may and may not reach is defined separately in the security document rather than in the room UI.

npm test runs Deno, and the check script lists files by hand

The first command most people will try is npm test, and its body is deno test --allow-read tests/unit. The lockfiles agree with that: deno.lock sits beside package-lock.json in the repository root, and the GPU and benchmark paths are Deno programs started with the --unstable-webgpu flag. The heavier script is check, which runs deno check over engine/*.js and then loops node --check across a hand maintained list of paths, covering api, room, harness, site/js, engine/wgsl, cli/bin, cli/lib, cli/test, several files under tests/e2e, and every package directory with a glob. That enumeration is the brittle part, since a new file added under any of those trees gets no syntax check until somebody edits the script. GPU runs are delegated to tests/run.sh with quick, q38, q38once, all and selftest as the modes, and serving the project locally needs two rewrites kept in step:

bash
git clone https://github.com/Nehanth/pooled && cd pooled
npx -y serve -l 8080 .        # then open http://localhost:8080/room

serve reads serve.json for the /room and /r/:code rewrites, and production uses the same two rewrites in vercel.json, so the two files drift apart silently when only one is edited.

Editorial conclusion

Try Pooled when you have two or three devices with spare memory and a use for a private local model, especially if you want Codex CLI, Claude Code or opencode pointed at a machine that is not yours. Check three things first: that the summed weight footprint of the model fits across the devices actually present, that every tab is on a Pooled build with tool calling if you plan to use the bridge for agents, and that the repo you cloned under the name swarmllm is the one whose clone instructions say pooled, because that mismatch is still unresolved at v1.0.0. Do not expect it to replace a server-hosted model on latency: one request at a time through a room queue is the design, not a bug.

Frequently asked questions

How much total memory does a Pooled room need to run Qwen 3.8 27B?

About 17 GB across the whole room at Q4_0, spread over however many devices are connected. The 35B MoE is the heavier option at about 22.5 GB, while Qwen3 1.7B at Q8_0 asks for roughly 5.6 GB for 16K context and roughly 4 GB for 8K.

Where do my prompts go when I drive a Pooled room with an external agent?

Prompts go to the room's host and, unless the host limits who sees answers, to everyone in the room. The serve bridge listens on 127.0.0.1 only, and setting POOLED_TOKEN makes it require a key.

Can I point Codex CLI or Claude Code at a Pooled room?

Yes, through npx @pooled/cli serve on Node 22 or newer, which exposes OpenAI Chat Completions and Responses on 127.0.0.1:8080/v1 and Anthropic Messages on 127.0.0.1:8080. The room's host must be running a Pooled build with tool calling, because older hosts answer tool requests with a 400 telling the client to reload.

Will a Pooled room work on my office network?

Rooms use direct UDP connections, which work networks often block. When the site has a relay set up, rooms fall back to it on their own over TCP or TLS on port 443, and model weights never go through the relay.

Official sources

  1. License: MIT
  2. Nehanth/swarmllm on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nehanth-swarmllm.svg)](https://hysenlabs.com/projects/nehanth-swarmllm)