gpu-hot publishes port 1312 without a login, and the compose file takes host process access by default
🔥 Real-time NVIDIA GPU dashboard
At a glance
- What is it?
- gpu-hot is a small FastAPI service in a container that polls NVML and serves an NVIDIA GPU dashboard on port 1312, plus a hub mode that aggregates several machines over plain HTTP. Its compose file, its Dockerfile and its README answer the exposure question three different ways.
- Who is it for?
- gpu-hot fits a trusted lab machine or a network segment you already treat as trusted, because its value is a browser dashboard with no account system, no per-user access and no TLS in the documented setup. Check four things before you publish it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 39 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Port 1312 is published on every interface, and nothing in front of it asks who you are
The same port number appears in five places: the single machine command publishes `1312:1312`, the compose service maps the same pair, the Dockerfile declares `EXPOSE 1312`, the hub example points at `http://server1:1312` style URLs, and the fix offered when a hub cannot reach its nodes is `sudo ufw allow 1312/tcp`. Through none of that does a password, a token or a required header appear. There is no environment variable for one in the configuration block either, so the dashboard answers to anyone who can reach the socket, and the hub, which needs no GPU of its own, reads each node over plain HTTP:
# On each GPU server
docker run -d --gpus all -p 1312:1312 -e NODE_NAME=$(hostname) ghcr.io/psalias2006/gpu-hot:latest
# On a hub machine (no GPU required)
docker run -d -p 1312:1312 -e GPU_HOT_MODE=hub -e NODE_URLS=http://server1:1312,http://server2:1312,http://server3:1312 ghcr.io/psalias2006/gpu-hot:latestThe only diagnostic offered for reachability is to fetch `/api/gpu-data` on a node directly, which is the same unauthenticated snapshot the dashboard itself uses.
The compose file takes host process visibility that the one-liner asks you to opt into
Process monitoring is documented as something you add. The instruction is to add `--init --pid=host` to see process names, followed by a warning that this allows the container to access host process information. `docker-compose.yml` sets `pid: "host"` and `init: true` in the service with nothing conditional around them, so the compose route arrives with host PID namespace access already granted while the quick start treats it as a decision. Both routes publish the same port and expose the same metrics, which means the two documented installs are not the same security posture.
The compose file also takes GPU selection away from the command line. Devices are reserved through `count: all` under the nvidia driver, and pinning specific cards means editing YAML to replace `count` with `device_ids` and list the UUIDs shown under each card in the UI. Two example UUIDs sit there commented out. The `docker run` form is the opposite, taking the whole GPU set from `--gpus all` and narrowing it later through `NVIDIA_VISIBLE_DEVICES`.
The fallback path polls four times slower, and you enable it by symptom
Two timers drive the sampling. The NVML path uses `UPDATE_INTERVAL`, documented as 0.5 seconds by default, and the fallback that shells out to nvidia-smi uses `NVIDIA_SMI_INTERVAL`, documented as 2.0 seconds. On the same dashboard that is a four times coarser picture for older cards, and nothing in the configuration warns about the difference when you switch.
The switch itself is `NVIDIA_SMI=true`, and the documented condition for setting it is that metrics do not appear. That makes the decision symptom based: you switch after an empty dashboard, not from a capability the container could have detected. The layout puts `core/monitor.py` and `core/nvidia_smi_fallback.py` side by side, so both paths ship in the image, and the NVML path depends on the pinned `nvidia-ml-py` package. The seven documented settings, with their defaults, are short enough to read in one screen:
NVIDIA_VISIBLE_DEVICES=0,1 # Specific GPUs (default: all)
NVIDIA_SMI=true # Force nvidia-smi mode for older GPUs
GPU_HOT_MODE=hub # Set to 'hub' for multi-node aggregation (default: single node)
NODE_NAME=gpu-server-1 # Node display name (default: hostname)
NODE_URLS=http://host:1312... # Comma-separated node URLs (required for hub mode)
UPDATE_INTERVAL=0.5 # Optional. NVML polling interval in seconds (default: 0.5)
NVIDIA_SMI_INTERVAL=2.0 # Optional. nvidia-smi fallback polling interval (default: 2.0)Polling pauses when nobody is connected, and the health check keeps knocking
Idle cost is handled explicitly: polling is paused automatically when no clients are connected, so idle CPU usage stays near zero. Both the image and the compose file then declare a health check that curls `http://localhost:1312/api/gpu-data` on a timer, so a container nobody is watching still receives a request every 30 seconds. Whether an HTTP snapshot request counts as a connected client is not something the configuration block or the API section answers, so the two claims about idle behaviour cannot both be read as a guarantee about a container that is only being monitored.
The metrics themselves are broader than the summary suggests. The feature list names utilization, temperature, memory, power draw, fan speed, clock speeds, PCIe info, P-State, throttle status and encoder or decoder sessions, and it advertises sub-second updates, which matches the half second NVML interval. Historical charts cover utilization, temperature, power and clocks, and system metrics are listed as CPU and RAM.
The same health check is declared twice, with a 5 second start period in one file and 40 in the other
The Dockerfile health check runs every 30 seconds with a 10 second timeout, 3 retries and a start period of 5 seconds. The compose service repeats the same curl against the same endpoint with the same interval, timeout and retry count, and raises the start period to 40 seconds. Which one is in force depends on whether you started the image or the compose file, and the difference matters on a cold node where NVML takes time to answer.
The diagnostic offered when no GPUs are detected points at a third image again:
nvidia-smi # Verify drivers work
docker run --rm --gpus all nvidia/cuda:12.1.0-base-ubuntu22.04 nvidia-smi # Test Docker GPU accessThat is a `base` CUDA image at 12.1.0, while the project builds from `nvidia/cuda:12.2.2-runtime-ubuntu22.04`. A green result there proves the driver and the container runtime can see a card, not that the dashboard image will, since the two differ in tag, in variant and in what is installed.
Eight Python dependencies are frozen at exact versions and nothing upgrades them
Every line of `requirements.txt` uses `==`:
fastapi==0.104.1
uvicorn[standard]==0.24.0
websockets==12.0
psutil==5.9.6
nvidia-ml-py==13.580.82
requests==2.31.0
websocket-client==1.6.3
aiohttp==3.9.1The Dockerfile installs them with `--no-cache-dir` and no constraint relaxation anywhere in the build, so a fix to any of them arrives only when somebody edits the file. The image installs `python3`, `python3-pip` and `curl` from apt, copies the tree, creates `templates` if it is missing and starts `python3 app.py`. The environment also differs between the two documented starts: the compose service sets `NVIDIA_VISIBLE_DEVICES` and `NVIDIA_DRIVER_CAPABILITIES` explicitly, while the single machine command sets neither and relies on `--gpus all`, which is also why the compose file pins `restart: unless-stopped` and the one-liner does not mention a restart policy at all.
The browser sample opens a raw WebSocket against a socket.io path
The documented client is a few lines of browser JavaScript, and the path it opens is the interesting part:
const ws = new WebSocket('ws://localhost:1312/socket.io/');
ws.onmessage = (event) => {
const data = JSON.parse(event.data);
};That path is the one socket.io uses, and the sample connects with the plain WebSocket constructor rather than a socket.io client. On the server side the pinned libraries are `websockets` and `websocket-client`, and no socket.io server package appears in the dependency list. The payload carries more than the feature list claims: the summary says system metrics are CPU and RAM, while the sample describes `data.system` as host CPU, RAM, swap, disk and network, with `data.gpus` and `data.processes` alongside it. Two HTTP endpoints sit next to the socket, `/` for the dashboard and `/api/version` for version and update information.
The documented layout stops before the tests, the docs and the compose file
The Project Structure block names `app.py`, `version.py`, and a `core/` directory holding `config.py`, `monitor.py`, `handlers.py`, `hub.py`, `hub_handlers.py`, `nvidia_smi_fallback.py` and a `metrics/` pair of `collector.py` and `utils.py`, then stops at `static`. The root actually holds eighteen entries. `Dockerfile`, `docker-compose.yml`, `requirements.txt`, `run_tests.sh`, `tests/`, `docs/`, `templates/`, `.dockerignore`, `.editorconfig` and `gpu-hot.png` are all missing from that block, so a test runner exists in the repository with no documented command to invoke it and no documented statement of what it asserts.
Two smaller mismatches sit alongside it. The repository records JavaScript as its primary language while the documented server is Python, with the single JavaScript in the README being the WebSocket sample. And the feature list advertises a scale from 1 to 100+ GPUs, a figure that arrives with no measurement, no configuration note and no test behind it. Last push is 2026-08-25 and the newest tag is v1.9.2 from 2026-07-18, which is also the same day as v1.9.1.
Editorial conclusion
gpu-hot fits a trusted lab machine or a network segment you already treat as trusted, because its value is a browser dashboard with no account system, no per-user access and no TLS in the documented setup. Check four things before you publish it. Put an authenticating proxy or a tunnel in front of port 1312, since nothing in the configuration adds a login and the troubleshooting steps tell you to open the port in the firewall. Decide whether host PID access is acceptable, remembering that docker-compose.yml requests it by default while the documented one-liner makes you ask for it. Pin the ports and node URLs you intend to expose, since the hub reads `http://` node endpoints. And read the pinned dependency list as a maintenance obligation rather than a convenience, since every package is frozen at an exact version and the fallback polling path is four times coarser than the default one.
Frequently asked questions
How do I run the gpu-hot dashboard on one machine?
With `docker run -d --gpus all -p 1312:1312 ghcr.io/psalias2006/gpu-hot:latest`, then open `http://localhost:1312`. Docker and the NVIDIA Container Toolkit are the stated requirements. From source it is `git clone`, `cd gpu-hot` and `docker-compose up --build`.
Does the gpu-hot dashboard need a login or an API key?
Nothing in the documented configuration sets a password, a token or a header. The port is published as 1312:1312, the hub reads nodes over plain `http://` URLs, and the firewall step offered for unreachable nodes is `sudo ufw allow 1312/tcp`. Putting an authenticating proxy in front of it is left to you.
Why does gpu-hot show no metrics on an older GPU?
Add `NVIDIA_SMI=true` to force nvidia-smi mode, which is the documented response when metrics do not appear. That path is coarser: `NVIDIA_SMI_INTERVAL` defaults to 2.0 seconds against the 0.5 second NVML interval, and `UPDATE_INTERVAL` can be raised if sampling itself is the problem.
How does gpu-hot aggregate several GPU machines?
Set `GPU_HOT_MODE=hub` on a machine with no GPU and give it `NODE_URLS` as a comma-separated list of node URLs, which is required in hub mode. Each GPU server runs the same image with `NODE_NAME` set, defaulting to the hostname. `list_dumps`-style discovery is not involved; the hub reads each node's `/api/gpu-data`.
How do I see which processes are using the GPU in gpu-hot?
Add `--init --pid=host`, and note that this allows the container to access host process information. The compose service already sets `pid: "host"` and `init: true`, so process names appear there without any extra flag.
Does gpu-hot read the GPU through NVML or through nvidia-smi?
NVML by default, through `core/monitor.py` and the pinned `nvidia-ml-py` package, with `core/nvidia_smi_fallback.py` as the path for older cards. Metrics cover utilization, temperature, memory, power draw, fan speed, clock speeds, PCIe info, P-State, throttle status and encoder or decoder sessions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/psalias2006-gpu-hot)