# sparkDash reads a DGX Spark from inside a privileged container

> The MiaAI-Lab/sparkDash repository is one privileged container that watches NVIDIA DGX Spark (GB10) machines, plus any Linux host with an NVIDIA GPU, and probes the local LLM servers running on them. It carries version 1.8.9 with no matching release tag, and its history lives in a changelog file rather than in GitHub releases.

**MiaAI-Lab/sparkDash** — sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard

- Repository: https://github.com/MiaAI-Lab/sparkDash
- Website: https://x.com/MiaAI_lab
- Stars: 511 · Forks: 102
- Language: JavaScript
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/miaai-lab-sparkdash

## The container takes the host network, the host PID namespace and root

The image is built for one platform only, linux/arm64, because the target is the GB10, and its base is pulled from public.ecr.aws instead of Docker Hub: Hub often resolves to IPv6, and a Spark with no IPv6 route fails auth.docker.io with a network is unreachable error. The build also retries npm ci once and then clears the npm cache, because a single long install has been seen to fail with a message that the exit handler was never called. Three runtime settings decide what sparkDash can see, and each is a compromise made on purpose. `network_mode: host` exists so local LLM probes can reach servers bound to loopback: ds4-server and start.sh default to `--host 127.0.0.1`, bridge NAT cannot reach the host's own 127.0.0.1, and a LAN-IP probe misses a loopback bind entirely. `pid: host` exists so the compute-apps list of nvidia-smi can see GPU processes, and without it local VRAM falls back to a minute-old `config/gpu-memory.json`. `privileged: true` covers the remaining host metrics. The read-only mounts follow from those three: `/proc`, `/sys` and the whole root filesystem at `/host/root`, plus `nvidia-smi` and `libnvidia-ml.so.1` mapped in as a fallback for when the namespace route fails. The compose file also runs the watcher in production, with `command: ["node", "--watch", "server/index.js"]` and `./server` bind-mounted, so an edit to the server code reloads the process without an image rebuild. What you end up running is a container that sees every process on the machine, shares its network stack and reads its filesystem. Put it on a box that does inference and nothing else.

## Loopback by default, and one switch that removes the token

The environment template is where the network posture is decided, and it says so in its own comments. Startup on a non-loopback address fails closed, because this release does not authenticate direct LAN clients, and the comment points at `docs/REMOTE-ACCESS.md` for the supported way in. The four values that matter at first start are these. `SPARKDASH_ALLOW_OPEN_REMOTE` defaults to 1, which permits a tokenless remote bind; set it to 0 and the server refuses to come up without `SPARKDASH_TOKEN`. So the shipped default is permissive in exactly the case the comment above it warns about, and anyone putting this on a shared network has to change it or put a tunnel in front. `LLM_PORT=8888` is where the first local LLM server is expected, and since each additional port gets its own panel with independent backend detection, a second server on another port is a second panel rather than a reconfiguration. The dashboard's own listener is `PORT=5555`.

## Seven LLM backends, one label field, and a rate that can read zero

The probe auto-detects llama.cpp, vLLM, sglang, ds4-server, EXL3, TensorFold and q27, reports live decode and prefill tok/s, and separates cached from uncached prefill on ds4, llama.cpp, SGLang and q27, with a daily peak kept on the LLM card. TensorFold arrived in 1.8.9 and is identified from `/v1/models` when the response carries `owned_by: tensorfold`, then labeled on the LLM card and in the Overview. On four of those seven, ds4, llama.cpp, SGLang and q27, the prefill figure is reported twice, once against a prefix the server already holds and once against a cold one, which is the difference between a prompt that can be reused and one that cannot. A daily peak is kept on the LLM card, so the live number is not the only number on screen. The TensorFold rate is a different matter: it is computed from cumulative token totals on `/health`, and stock TensorFold does not publish those yet, so the card reads 0 tok/s until the server starts publishing them. The label and the rate come from two different endpoints, which means a zero there can mean not reporting rather than not serving, and the benches and the prompt demo still work against that server like any OpenAI-compatible one. The inference-health tiles are narrower still: KV cache share, run and wait queue, TTFT, E2E and ITL p95, preemptions, prefix cache and MTP accept all come from Prometheus `/metrics`, and only vLLM and q27 are wired up. q27 FIFO-queues, so its Requests tile reads "N run" with no wait gauge behind it.

## The benchmarks change the prompt with the concurrency and the context size

The decode bench streams at several concurrencies at once and offers a type picker of Structured, Prose, Code and JSON, where the Code type hands each stream a different Python task so the streams do not share one prompt and quietly measure cache hits instead of generation. Everything runs on what the README calls the lab protocol, temp 0 with thinking off, and the last run is persisted so a number stays on screen after a reload. The prefill bench sweeps context size from 1k to 300k, reporting prefill tok/s and TTFT, and gives every size its own unique prefix: a prefix cache warmed by the previous size cannot flatter the next one. Remote units are reached over LAN HTTP or through an SSH tunnel to loopback, and a Remote button takes an on-demand HTTPS host and port instead, chosen for the run rather than configured as a unit, which is what lets you point a decode bench at an endpoint that is not in the unit list at all. 1.8.9 added remote-Spark benches over an SSH tunnel, a custom prefill size, that on-demand bench host, and a share-as-image card for bench results, and it fixed the decode-bench request quota and its 24x and 32x budget, long prefills dying at about 5 minutes, and SGLang prefill latching. Read a decode figure as valid within one run, and treat a cross-machine comparison as a different measurement.

## Two clocks, one in the environment file and one in the frontend bundle

How often a number is sampled and how far back the charts can look are configured in two different places, and only one of them is a runtime setting. The environment template gives per-subsystem poll intervals in milliseconds, and the storage panel is deliberately the slowest while bandwidth is the fastest. Metrics arrive over a WebSocket and land in a central history store, which is what lets a sparkline survive a tab switch instead of resetting when you click another unit. Retention is the other half: the compose file passes `VITE_HISTORY_HOURS` into the build with a default of 8, and both its comment and the Dockerfile say the same thing, that changing it means rebuilding, with the Dockerfile naming `docker compose build --build-arg VITE_HISTORY_HOURS=4` as the override and pointing at `src/hooks/metricsStore.ts` as the file that consumes it. So the sampling rate is something you tune in an env file and restart, while the window the frontend can display is compiled into the bundle. A practical consequence: the storage panel is not a live view, and raising the GPU interval to smooth a busy chart silently drops detail rather than lengthening the history.

## Units are added in the browser, and the SSH passwords sit in an encrypted file

Units are managed as configuration, not as code: they can be added, edited, reordered and removed from the UI with no process restart, and each one carries a role. Head, Worker and Standalone each behave differently, a worker is labelled and linked to its head, a standalone unit can switch its own LLM monitoring off, and workers can be hidden from the Overview and the tab strip when you only want to look at the machine serving traffic. Secrets are handled on the same principle of keeping them out of the data path: SSH passwords are encrypted with AES-256-GCM and are never written into `sparks.json` or returned in API responses, and the config directory is bind-mounted at `./config` so they survive a container restart. Remote units authenticate with an SSH key or a password, and two variables in the environment template govern the transport. `SSH_CONTROL_PERSIST_SECONDS=60` reuses authenticated SSH transports for remote polling and can be set to 0 to turn that off. `SSH_IDENTITY_FILE` points at a key inside the container when the name is not a default OpenSSH one, with a comment requiring mode 600 on the host mount. 1.8.9 also fixed remote SSH session churn and Tailscale address classification. Since units come and go from the browser, that config directory is the real source of truth, and neither the compose file nor the environment template says where the encryption key itself comes from, so treat a key change as a question to ask before you rotate one.

## Three opt-in probes report the work on top, not the hardware underneath

A separate family of probes looks at what a unit is doing rather than how fast it is, and all three are opt-in. ComfyUI monitoring is switched on per unit with `comfyMonitoring`, which defaults to off, alongside `comfyPort` at a default of 8188; head, worker and standalone units can all run it. It reads the queue and the jobs, shows progress, offers cancel and an Open link, and adds an inventory and an overview chip, while deliberately not drawing a second copy of the GPU and RAM bars that already sit on those panels. The ComfyUI probe borrows the same request path as the LLM probe, one instance per unit, and reports the job list rather than a second copy of the hardware counters. Hermes Agent is also per unit, with a background update check every 10 minutes, status badges, and a one-click or batch `hermes update` across units, so it is the only one of the three that can change the machine rather than describe it. The tailnet probe is the smallest of the three and the one that catches a failure nothing else reports: it flags a unit that is healthy on the LAN but off its tailnet. None of the three is on in a fresh install, so a dashboard that shows plenty of hardware and inference numbers may still be telling you nothing about the jobs running on the machine.

## Version 1.8.9 exists in the manifest but has no tag

The package manifest declares version 1.8.9, and the changelog section at the top of the README covers that release: the TensorFold backend, the q27 backend, a custom prefill size, remote-Spark benches over an SSH tunnel, an on-demand Remote bench host, an option to hide worker nodes, and a share-as-image card for bench results, plus the fixes for the decode-bench quota, long prefills, SGLang prefill latching, `SPARKDASH_TOKEN` in compose, Tailscale address classification and remote SSH session churn. Full history is kept in `CHANGELOG.md`, linked from the table of contents as the full changelog. The only GitHub release is a media asset named Media: LLM Prompt Showcase, published 2026-07-23, carrying the demo video that also sits at `assets/llm-showcase.mp4`. No tag corresponds to 1.8.9, so a deployment either tracks the main branch or copies the changelog itself. The repository is not archived, its last push was 2026-10-01, and it has 511 stars, 102 forks and 17 open issues. Tests are split into two suites, node's own runner over the collectors, sparks and llmtokens test directories, and vitest for the frontend, with `tsc --noEmit` for types. The working tree splits into `server/` for the Node collectors, `src/` for the React client, `src/shared/` for the prompt catalog that the decode bench and the prompt demo import, `config/` for the encrypted secrets, plus `scripts/`, `deploy/`, a `deploy.sh` at the root and a `docs/` folder that the environment template points at. Two compose files sit next to two Dockerfiles, and the package scripts pair the watch-mode server with vite under `npm run dev`, which is the same pair the production compose runs. One oddity sits in the root: a file named SPARKDASH-REMEDIATION-MERGE-SUMMARY-2026-09-07.md, a merge summary that ships to everyone who clones.

## Conclusion

Treat sparkDash as a tool for a machine whose only job is inference. It gives you a defensible read on a fleet you control: per-unit roles, per-port backend detection, and two benchmarks designed so their numbers are comparable inside a run and not across machines. Before deploying, settle three things. Can you accept a container that shares the host network and PID namespace, runs privileged and mounts the root filesystem read only? Do your LLM backends publish what the probe reads, since TensorFold does not publish token totals and its tok/s card reads 0 until it does? And do you need a pinned version, because 1.8.9 has no tag. If you need LAN authentication for the dashboard itself, do not expose this build: the environment template states that direct LAN clients are not authenticated, and the tokenless remote bind defaults to on.

## FAQ

### What is sparkDash?

It is a real-time web dashboard for one or more NVIDIA DGX Spark (GB10) machines in a single browser window, streaming GPU, CPU, unified memory, storage, network and local LLM metrics. Any Linux machine with an NVIDIA GPU can join as a dedicated GPU host over SSH, with its system RAM and discrete VRAM reported separately.

### What operating system does sparkDash expect on a DGX Spark?

The repository does not state the operating system of the Spark itself. Its own image is a Debian Bookworm base, node:22-bookworm-slim, built for linux/arm64, and any non-Spark unit it monitors has to be a Linux machine with an NVIDIA GPU.

### How much does a DGX Spark cost according to sparkDash?

Nothing in the repository records a price, and the dashboard is free under the MIT licence. The hardware facts it does state are the GB10 platform and a 128 GB LPDDR5X unified memory pool of about 273 GB/s, split between GPU and CPU.

### Is the DGX Spark worth it, by sparkDash's own measures?

The repository takes no position on value, but it does give you the numbers to judge with: a prefill sweep from 1k to 300k context, multi-concurrency decode tok/s, TTFT, KV cache share and per-process VRAM, all measured at temp 0 with thinking off.

### What use cases does sparkDash assume for a DGX Spark?

Serving local LLMs is the assumption: the probe detects llama.cpp, vLLM, sglang, ds4-server, EXL3, TensorFold and q27, and the default LLM port is 8888. On top of that it can watch ComfyUI jobs and Hermes Agent status, so image generation and agent workloads are the other two the panels are built around.

## Sources

- [License: MIT](https://github.com/MiaAI-Lab/sparkDash/blob/main/LICENSE)
- [MiaAI-Lab/sparkDash on GitHub](https://github.com/MiaAI-Lab/sparkDash)
- [Project website](https://x.com/MiaAI_lab)
- [README](https://github.com/MiaAI-Lab/sparkDash/blob/main/README.md)
- [Releases](https://github.com/MiaAI-Lab/sparkDash/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/miaai-lab-sparkdash
