production-ocr-course: a six-week syllabus whose numbers outrun its evidence
Build, deploy, and scale a production-grade OCR pipeline using Rust, vLLM, Redis, KEDA, and Kubernetes.
At a glance
- What is it?
- An Apache-2.0 repository that packages a six-week cohort for running a generative OCR pipeline on AKS or GKE, with a Rust ingest gateway, a Python batch collector and vLLM behind Qwen 3.5 4B. The architecture choices are documented in outline only, the headline throughput figure ships without a benchmark, and four of the six weeks have walkthrough directories while two do not.
- Who is it for?
- Treat this as a syllabus and an architecture sketch, not a runnable system. It earns a look from engineers who already run GPU clusters and want the layout-first, two-stage OCR framing plus the scaling vocabulary, and it should be skipped by anyone expecting a repo that deploys on its own, because the pages carrying the security configuration, the document workflow and the scaling math are all below the point where the available copy of the page breaks off.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The 1.86 pages/second claim arrives with no benchmark attached
The overview leads with a throughput number: generative OCR at 1.86 pages/second, credited to vLLM continuous batching, PagedAttention and multi-token prediction. Nothing in the retrieved page backs it. There is no input document set, no page complexity breakdown, no latency percentile, no hardware configuration beyond the mention of T4 and A100 node pools, and no test harness among the top-level entries, which are client_rt_consumer, client_rt_producer, docs, images, infra-gke, k8s, server and the four week deployment walkthroughs. The load-testing tooling that Friday office hours promise is therefore not something a reader can clone and rerun. Read 1.86 as a course figure rather than a number to plan capacity against, because the distance between a synthetic benchmark and real traffic is exactly where autoscaling budgets get decided.
Weeks 1 and 2 have no walkthrough directory, weeks 3 through 6 do
The course is organised as a six-week cohort, and the top-level layout encodes that unevenly. Four directories carry standalone deployment walkthroughs: week3_vllm_deployment, week4_rust_gateway_deployment, week5_async_architecture_deployment and week6_apim_mcp_deployment. No equivalent directory exists for the Kubernetes for AI Systems week or the SOTA OCR Approaches week, and the week-by-week table attaches a standalone link only from week 3 onward. The first two weeks ship as articles plus live sessions: pods, services, node pools, resource scheduling, GPU drivers and taints in week 1, then evaluating the GLM-OCR SDK layout detection in week 2. That split is defensible, since infrastructure has less to script than a model server, but a reader who wants to move at their own pace cannot get the AKS and GKE node pool walkthrough as self-serve documentation, and the docs directory only carries onboarding, quota and deployment guides.
The Rust gateway caps ingest at 10MB while a Python worker does the batching
Week 4 builds the ingest side as an Axum service named client_rt_producer, tuned for heavy payloads with a 10MB limit and writing into Redis through atomic HSET calls instead of read-modify-write. Week 5 then passes the work to a Python worker named client_rt_consumer. The repository description summarises the stack as Rust, vLLM, Redis, KEDA and Kubernetes, which holds for the gateway and the control plane but leaves the Python consumer out entirely. That omission matters when sizing a team, because the component the description implies you own in Rust is only the front door, while the collector that groups work into batches is not Rust at all. The 100ms collection window in week 5 is the knob deciding how many pages reach the inference engine together, and a window that short only pays off if the producer can keep it full, which is where the 10MB ceiling and the write atomicity start to shape burst behaviour.
A RAM-disk handoff is paired with two node pools that scale independently
Documents move from the layout encoder to the inference engine through a zero-copy handoff on /dev/shm, presented as a RAM-disk transfer rather than a network round trip. In parallel, KEDA scales the T4 layout pool and the A100 inference pool as separate deployments from zero to bursting load, and week 5 stacks scale-to-zero on top of that. The retrieved page never says how a document written to /dev/shm on one node reaches a model server scheduled onto another. With two independently scaled pools and a transfer mechanism that relies on local shared memory, that boundary is the load-bearing question of the architecture, and the only place it could be answered is the section on the exact document workflow, which sits below the point where the available copy breaks off. Hold the two claims as unverified against each other until you can read that section or inspect the manifests.
Zero public exposure is asserted in the overview and configured somewhere else
The overview states that the pipeline sits behind an internal load balancer plus an enterprise API gateway with zero public exposure, and week 6 assigns JWT verification, rate limiting and policy work to Azure APIM or the GCP API Gateway. Nothing in the retrieved page shows the policy, the manifest or the config file that enforces any of it. The top level includes an infra-gke directory and a k8s directory, which is where such configuration would live, but their contents are out of view. The internal load balancer framing also implies private cluster settings and DNS inside the provider account, so following along means working inside a cloud VPC rather than a laptop. Reviewing this design for security means reading the week 6 walkthrough and the k8s resource definitions directly, because the summary bullets stay aspirations until those definitions are checked.
The retrieved copy ends mid-word and every section after it is unreachable
The table of contents lists ten sections past Getting Started: the SLM advantage, pipeline architecture and throughput, the exact document workflow, a pre-layout encoder fidelity argument, a document handoff technical report, a formal architecture assessment, a scaling philosophy, the tech stack, contributors and the license. The copy that could be retrieved stops partway into the first of them, mid-sentence, on a line about OCR having evolved from simple character recognition. Everything below that point is absent, Contributors included, so the people behind the course cannot be identified from what a reader can see. The visible page is also padded with two empty centred paragraph blocks and a newsletter signup table sitting between the overview and the audience section. What a reader receives is the marketing arc and the syllabus, and none of the architecture.
No releases to pin against, and a main branch that moved on 24 September
Repository metadata records 421 stars, 123 forks and no open issues, with the last push dated 24 September 2026 on the default branch main, and no GitHub releases at all. A course whose weeks are meant to be followed in order has nothing versioned to pin against: there is no tag to check out for week 3, so a reader who falls behind cannot return to the state of the vLLM deployment walkthrough as it was written. The MAX_NUM_BATCHED_TOKENS tuning, the 100ms collector window and the 10MB gateway ceiling are all values that drift as vLLM and the model move. Apache-2.0 covers the code, and the LICENSE file sits at the top level beside the four week directories, so reuse and modification are unencumbered even though the material is delivered as a live cohort rather than as self-paced documentation.
Editorial conclusion
Treat this as a syllabus and an architecture sketch, not a runnable system. It earns a look from engineers who already run GPU clusters and want the layout-first, two-stage OCR framing plus the scaling vocabulary, and it should be skipped by anyone expecting a repo that deploys on its own, because the pages carrying the security configuration, the document workflow and the scaling math are all below the point where the available copy of the page breaks off. Before committing cluster spend, check three things yourself: whether 1.86 pages per second reproduces on your document mix, how a /dev/shm handoff crosses between separately scaled T4 and A100 nodes, and whether an unversioned main branch still matches the walkthrough you read.
Frequently asked questions
What does production-ocr-course actually teach?
It teaches how to run a generative OCR pipeline in production on Kubernetes, aimed at ML and platform engineers who already know how to call an OCR API. The syllabus covers GPU node pools, autoscaling, network security and the systems tradeoffs behind serving a generative model at throughput.
Which models and SDK does production-ocr-course serve?
The inference model is Qwen 3.5 4B, deployed on vLLM with continuous batching, PagedAttention and multi-token prediction. Layout detection comes from the GLM-OCR SDK, which the course evaluates against single-stage end-to-end models in week 2.
Does production-ocr-course ship runnable install commands?
No shell commands appear on the retrieved page. Entry points are documents instead: docs/azure_onboarding.md and docs/gcp_onboarding.md for accounts, docs/azure_gpu_prereqs.md and docs/gcp_gpu_prereqs.md for GPU quota, then docs/aks_deployment.md, docs/gke_deployment.md and docs/cloud_comparison.md.
Can production-ocr-course run on free or trial cloud credits?
The page says free and trial accounts cannot run GPUs, and names T4 plus A100 as the quota to request first. Azure is the primary cloud and GCP is the optional path, with a provider comparison document listing where the two diverge.
Which programming languages appear in the production-ocr-course codebase?
Two: Rust for the Axum ingest gateway client_rt_producer, and Python for the batching worker client_rt_consumer. The repository description names Rust, vLLM, Redis, KEDA and Kubernetes, and does not mention the Python worker.
Is the 1.86 pages per second figure reproducible from production-ocr-course?
Nothing in the retrieved page supports reproducing it. There is no benchmark harness among the top-level directories, no document set and no latency percentile, only the figure itself alongside the vLLM techniques credited for it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/neural-maze-production-ocr-course)