kaggle-tpu-lab: Run Large Open Models on a Free Kaggle TPU
Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.
At a glance
- What is it?
- kaggle-tpu-lab is a set of Kaggle notebooks and scripts that run Qwen3.8-27B and GLM-5.3-Flash on Kaggle's free TPU v5e-8 and expose a public OpenAI/Anthropic-compatible API endpoint. It targets engineers who want a capable inference server for a coding agent without paying for cloud compute.
- Who is it for?
- Engineers who want a cost-free inference endpoint for a coding agent and can accept session restarts every nine hours will find kaggle-tpu-lab a direct fit. Teams needing continuous uptime, production traffic volumes, or stable URLs should use dedicated GPU cloud instances instead.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Free Kaggle TPU Quota as the Compute Source
kaggle-tpu-lab solves one specific problem: running a large language model for a coding agent costs money when done on rented GPU cloud instances. Kaggle gives phone-verified accounts around 20 TPU v5e-8 hours per week at no charge. The project turns that quota into a usable inference server. The intended user is a developer who wants to point Claude Code, Codex CLI, opencode, or a similar agent at a capable model and has no budget for cloud compute.
The TPU v5e-8 is the eight-chip variant of Google's fifth-generation edge TPU. The project uses it to run two models that would require multiple high-end GPUs in a typical cloud setup. The free-tier constraint shapes every design choice: sessions are capped at just under nine hours, the URL changes on each restart, and the weekly quota limits how often you can run back-to-back sessions.
Qwen3.8-27B on vllm-tpu and GLM-5.3-Flash on a Custom JAX Engine
The repository contains two model configurations. Qwen3.8-27B loads in bfloat16 with no quantization and runs on vllm-tpu with one patch applied by the project. The README reports around 130 tokens per second on a single stream, around 540 tokens per second across eight concurrent streams, and a prefill rate of about 10,300 tokens per second.
GLM-5.3-Flash is a 320-billion-parameter mixture-of-experts model. The repository loads it with 3-bit expert weights and int8 for the remaining components. It runs on a JAX inference engine written by the project, which the README describes as, to their knowledge, the first engine to run GLM-5.3-Flash on a TPU. Single-stream throughput is about 64 tokens per second, three-stream throughput is about 90 tokens per second, and prefill runs at around 1,600 tokens per second.
Both models expose a 262k token context window with prefix caching enabled. The numbers in the README are measured on the shipped configuration; the per-model folder READMEs describe how they were obtained.
Starting the Server from the Kaggle Notebook or the Terminal
The notebook route requires no local installation. Open the Kaggle notebook for the model you want, press Run All, and the kernel attaches public datasets that hold the weights and, where it helps, a pre-built compile cache. It then builds the engine across the eight TPU chips, opens a Cloudflare tunnel, and prints a URL and an API key. Leave the last cell running. A keepalive holds the session for just under nine hours.
For terminal users who prefer driving it from a local shell, the README gives these steps:
git clone https://github.com/ARahim3/kaggle-tpu-lab
cd kaggle-tpu-lab
python launch.py serveThe `launch.py` script pushes the kernel using the Kaggle CLI and follows its output. The `status` and `stop` subcommands check or end a running session. The `--model` flag will select a different recipe once additional recipes are added. Running from the terminal requires Python 3.9 or later and the Kaggle CLI.
Connecting Claude Code, Codex CLI, and opencode to the Endpoint
Once the notebook prints its URL and API key, wire any OpenAI or Anthropic-compatible client to it by pointing it at the tunnel address. The README gives the one-line form for Claude Code:
ANTHROPIC_BASE_URL=<url> ANTHROPIC_AUTH_TOKEN=<key> ANTHROPIC_MODEL=<model> claudeSet ANTHROPIC_BASE_URL to the Cloudflare tunnel URL the notebook printed, ANTHROPIC_AUTH_TOKEN to the API key, and ANTHROPIC_MODEL to the model identifier. Each per-model folder README provides the exact values and the equivalent commands for Codex CLI and opencode. Because the endpoint speaks both the OpenAI and the Anthropic APIs, any client that supports one of those wire formats can be pointed at it without code changes.
Repository Layout and Adding a New Model
Each model lives in its own top-level folder: `qwen38-27b/` and `glm53-flash/` are the two that ship. Each folder contains a README with throughput numbers and measurement notes, a `kernel/` directory with the serving script, and a `notebook/` directory with the run-all notebook generated from that script. An important design decision is that the launcher and the notebook share the same configuration block: changing a setting in one place changes it in both.
The CONTRIBUTING.md file describes the expected folder structure for new model recipes. Adding a model means creating a top-level folder with that layout, adding whatever the recipe needs (a patch file, a custom engine), and following the shared configuration convention. At the time of writing, only the notebook route is fully automated for Qwen3.8-27B; the `--model` flag in `launch.py` is wired for future recipes.
Where kaggle-tpu-lab Fails and Who Should Not Use It
The nine-hour session cap is the main operational constraint. When the session ends, you must run the notebook again and the tunnel URL changes. Any tool or script that hard-codes the endpoint URL will break on each restart. The free quota of about 20 TPU hours per week is exhausted by roughly two full sessions; heavy or continuous use is not feasible on the free tier.
The project has no GPU path. It is TPU-only, which means the quantization formats and the engines it uses are specific to TPU execution. PyTorch users looking for a drop-in inference backend will find no support here. The model weights are not covered by the MIT license that governs the project code; each model folder documents the applicable weight license, and some commercial restrictions may apply to the models themselves.
Phone verification on the Kaggle account is required before TPU access is granted. This is a one-time step but is a blocker in environments where account verification is restricted.
Comparison with Renting a GPU Cloud Instance
The most direct alternative is renting a multi-GPU instance on a provider such as Lambda Labs, Vast.ai, or a major cloud vendor, and running vllm or llama.cpp on it. The core difference is cost versus control. A rented instance runs as long as needed, keeps a stable endpoint URL, supports GPU-optimized tooling, and can be scaled or reconfigured without depending on a platform's free quota. kaggle-tpu-lab provides none of those guarantees but also charges nothing up to the weekly TPU limit.
For a developer who wants to test whether a 27B or 320B model is useful for their coding workflow before committing to a paid setup, kaggle-tpu-lab provides that test at no cost. For anyone running a team workflow, an integration, or a long-running agent that cannot tolerate URL rotation, a rented instance is the appropriate choice.
Editorial conclusion
Engineers who want a cost-free inference endpoint for a coding agent and can accept session restarts every nine hours will find kaggle-tpu-lab a direct fit. Teams needing continuous uptime, production traffic volumes, or stable URLs should use dedicated GPU cloud instances instead. Before starting, confirm that your Kaggle account has phone verification complete, check the free weekly TPU quota, and read the model weight license in the folder for the model you plan to use.
Frequently asked questions
What is TPU on Kaggle?
Kaggle provides phone-verified accounts with access to a TPU v5e-8 accelerator, available at no charge for approximately 20 hours per week. kaggle-tpu-lab uses this hardware to serve Qwen3.8-27B and GLM-5.3-Flash with a public API endpoint that speaks the OpenAI and Anthropic wire formats.
Is Kaggle TPU free?
Kaggle's TPU access is free for accounts that have completed phone verification, with a quota of approximately 20 TPU hours per week. kaggle-tpu-lab is built around this quota: each session runs for just under nine hours, and the README describes re-running the notebook when a session ends.
How long does a kaggle-tpu-lab session stay up?
The README states that a keepalive holds the session for just under nine hours, after which a new run is needed. Each new session opens a fresh Cloudflare tunnel and prints a new URL and API key, so clients pointing at the old URL must be updated.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arahim3-kaggle-tpu-lab)