# 470 notebooks on Vertex AI MLOps, and the counts do not add up

> statmike/vertex-ai-mlops is a collection of Jupyter notebooks covering Google Cloud machine learning and AI work end to end, from AutoML and BigQuery ML to feature stores, serving and agents. It is genuinely useful as a map of the platform, and its own index has arithmetic gaps that tell you which folders to trust before you plan a reading path around it.

**statmike/vertex-ai-mlops** — Google Cloud Platform Vertex AI end-to-end workflows for machine learning operations

- Repository: https://github.com/statmike/vertex-ai-mlops
- Stars: 715 · Forks: 314
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/statmike-vertex-ai-mlops

## The platform was renamed twice more after Vertex AI, and the index follows

The first thing this README does is explain that the service it teaches has changed name several times: Cloud ML Engine, then AI Platform, then Vertex AI, and as of April 2026 the Gemini Enterprise Agent Platform. Each rename is presented as an expansion of scope, from custom model training to a unified platform with AutoML and managed notebooks, and now to model building, agent development, orchestration and governance together. Two more renames appear further down, Agent Engine for what was Vertex AI Agent Builder, and Managed Airflow for Cloud Composer.

That framing matters for how you read the repository. It is not a frozen tutorial for one API surface; the index describes itself as tracking the platform, and the most recent material is agents built with ADK, agent deployment on Agent Engine, and agentic workflows. Older folders are named for services that still exist, so the naming churn shows up as much in the directory list as in the prose.

What the collection deliberately does not do is pick one service per job. Spanner, AlloyDB, Cloud SQL, Memorystore, Firestore and Bigtable all appear as feature stores, vector search backends and SQL inference endpoints, which is the collection's actual argument: these services compose, and the notebooks are arranged to show the composition.

## The headline says 470 notebooks, the three section tables account for 167

The README announces a collection of 470+ interactive notebooks covering custom ML, generative AI and agent development. The three section tables that are spelled out add up to something much smaller. MLOps is 74 notebooks, data+ai is 40, and Applied GenAI is 53, which is 167 in total, leaving more than 300 notebooks unaccounted for by the index.

Those notebooks are not missing so much as uncounted, and the top level of the repository shows where they live: a numbered path from `00 - Setup` through `01 - Data Sources`, `02 - Vertex AI AutoML`, `03 - BigQuery ML (BQML)`, `04 - scikit-learn`, `05 - TensorFlow`, `06 - XGBoost`, `07 - PyTorch` and `08 - R`, ending at `99 - Cleanup`, followed by a long alphabetical tail of Applied and Learn folders.

So the 470 figure is a sum nobody recomputed, and the index describes three curated sections rather than the whole tree. Anyone planning a reading path has to walk the folder names instead, which is the practical reason this repository works better as a directory listing than as a curriculum.

## MLOps claims 74 notebooks and its six sub-sections add up to 72

The MLOps section is the biggest of the three, and its arithmetic is short by two. Serving is 32 notebooks, Feature Store is 21, Pipelines is 13, Model Evaluation is 3, Model Monitoring is 2 and Experiment Tracking is 1. Those six numbers total 72 against a section header that says 74.

The shape of the section is worth more than the total. Two folders carry 53 of the 72: Serving covers online endpoints in dedicated, shared and private with PSC flavours, batch inference through Vertex AI, Dataflow, Dataproc and Airflow, SQL-based inference through BigQuery ML, AlloyDB and Spanner, deployment to Cloud Run, GKE and Cloud Functions, Triton Inference Server and vLLM for large language model serving. Feature Store pairs the Vertex AI managed feature store with a 15 notebook deep dive into a self managed Bigtable one, touching serialization, sync patterns, history, schema evolution, vector search, replication and a recommendation engine capstone.

The remaining three folders are thin by design. Model Monitoring is two notebooks on feature skew and drift detection, Experiment Tracking is one on logging parameters, metrics and artifacts, and Model Evaluation is three. Read the section as two substantial paths and four spot checks rather than as 74 equal entries.

## data+ai says 40 and the counted folders already reach 40

The arithmetic runs the other way here. The data+ai header says 40 notebooks, and the folders that carry a number add up to exactly 40: BigQuery AI Functions at 32, Dataflow at 3, Dataproc at 4 and Tabular Data at 1. Two entries in the same list carry no number at all, Managed Airflow for orchestrating batch inference across Dataproc, Dataflow, KFP and Vertex AI, and AlloyDB with Spanner for ML.PREDICT() against Vertex AI Endpoints.

The same pattern repeats one level down. BigQuery AI Functions is described as 32 notebooks made of 21 individual function guides covering all 20 built-in AI functions such as AI.GENERATE, AI.EMBED, AI.FORECAST and AI.CLASSIFY, plus 9 end to end workflows for RAG, semantic search, document intelligence, content moderation and time series. Twenty one guides plus 9 workflows is 30, so that header is two high as well.

Both errors point the same way. The counts are maintained per folder by hand and are correct more often than not, but the section totals above them are written separately and drift. It is a small thing, and it is also the cheapest way to tell whether a number in this README describes files or intent.

## Retrieval compares 11 databases, names nine, and publishes no numbers

The Applied GenAI section is where the comparisons live, and where the README is most careful to promise without evidencing. Retrieval is 11 notebooks on vector search across what the README calls 11 Google Cloud databases: BigQuery, Vertex AI Vector Search, Feature Store, Spanner, AlloyDB, Cloud SQL, Memorystore, Firestore, Bigtable, and more. Nine named, two implied, with cost and latency comparisons promised for the set.

Similar promises sit elsewhere. Dataflow is 3 notebooks on streaming and batch inference with the Beam RunInference API, including model hot swap patterns and a GPU inference benchmarking study. Tabular Data is a single notebook on BigQuery read patterns with benchmarks and cost analysis. The BigQuery AI Functions set covers document intelligence and content moderation.

None of those comparisons carry a figure on the index page: no latency number, no cost per query, no ranking of the nine named databases. The value is real but it is locked inside the notebooks, so the cost of choosing a vector backend is one notebook read per candidate, not one page read.

## The top level mixes a lesson path with a research pile, in folders with spaces

The directory naming tells you how the repository grew. The numbered spine reads like a course for someone starting from an empty project: setup, data sources, AutoML, BigQuery ML, scikit-learn, TensorFlow, XGBoost, PyTorch, R, then cleanup. Alongside it sit folders that look like a working notebook author's own bench, among them `Dev/`, `IDE/`, `Explorations/`, `Framework Workflows/`, `scripting/`, `core/` and `architectures/`, and a set of thematic Applied folders that run from Applied Autoencoders to Applied Reinforcement Learning.

The names carry spaces and punctuation on purpose, `00 - Setup`, `03 - BigQuery ML (BQML)`, so a shell command or a link has to quote them, and the numbered prefix only sorts correctly inside a viewer that sorts by name rather than alphabetically. The two index files sit side by side as `readme.md` and `readme-legacy.md`, which is the repository admitting that its own front page was reorganised and kept the old one.

Three more top level entries are not course material at all: `.agents/`, `.claude/` and `agent-skills/`, which is agent configuration sitting next to 470 teaching notebooks.

## One interpreter pin, no dependency file, no workflow directory

The top level listing is short enough to audit. It holds `.python-version`, `.gitignore`, `LICENSE`, the two readme files, and the content directories. What it does not hold is a `requirements.txt`, a `pyproject.toml` or a `Dockerfile`, and there is no `.github/` directory in the listing at all, so no workflow file and no automation is visible at the top of the tree.

That leaves dependency resolution to the notebooks themselves, with the interpreter version pinned once for the whole repository. For a collection meant to be read by people at different stages, that is a reasonable trade, since a student on a fresh project does not want a shared lockfile from someone else's account. For a team that wants reproducible runs in CI, it is the wrong shape, and the fix is per notebook rather than per repository.

Versioning is equally informal. The repository has no GitHub releases, and the last push was 2026-09-16. A notebook collection this size will keep moving, so a documented reference to a specific set of notebooks has no version string to point at.

## The README opens with a share table and a right-click Save As link

The first thing on the page is an HTML table, not documentation. It offers share links for LinkedIn, Reddit, Bluesky and Twitter, a row of the author's own profiles, a link to view the file on GitHub, and a raw download link to readme.md on raw.githubusercontent.com with the instruction to right click and Save As.

That framing belongs to a site rather than a repository. Anyone who clones the work does not need to save the readme, and anyone who only wants the index cannot get the notebooks from a download link. It also means the entry point carries no statement of what a notebook needs before it runs, which is the first question a new reader has.

Licensing is the one thing the page resolves cleanly. The license field reads Apache-2.0 and a LICENSE file sits at the root, so the notebooks can be reused, adapted and vendored into internal material with attribution rather than only read in place.

## Conclusion

Treat this repository as a map, not a course. It answers one question well, which Google Cloud service belongs in which stage of a machine learning system, and it answers it with notebooks you can open and run against your own project IDs. It is the wrong thing to adopt as a versioned dependency: there are no releases, no tags, the last push was 2026-09-16, and the only environment pin at the top level is a .python-version file. Before you build a curriculum or an internal doc on it, open the folder you care about and count the notebooks yourself, because the index sections disagree with their own sub-sections, and pin your own snapshot of the commit rather than tracking main.

## FAQ

### How many notebooks does statmike/vertex-ai-mlops contain?

The README headline says 470+ interactive notebooks. The three section tables spelled out account for 167 of them, MLOps at 74, data+ai at 40 and Applied GenAI at 53, with the rest living in the numbered topic folders such as 02 - Vertex AI AutoML and 05 - TensorFlow.

### Is Vertex AI being renamed, and does that affect this notebook collection?

The README says the platform was renamed as of April 2026 to the Gemini Enterprise Agent Platform, after Cloud ML Engine and AI Platform, and it also notes Agent Engine for what was Vertex AI Agent Builder. The collection tracks the change rather than pinning one name, so its newest material is about agents.

### Does statmike/vertex-ai-mlops ship a pinned environment or CI?

Only an interpreter pin. The top level holds a .python-version and no requirements file, no pyproject.toml and no .github directory, so dependency installation is left to the individual notebooks and no workflow is visible in the tree.

### Which Google Cloud databases does the Retrieval section of these notebooks compare?

The README says vector search across 11 Google Cloud databases and names nine of them: BigQuery, Vertex AI Vector Search, Feature Store, Spanner, AlloyDB, Cloud SQL, Memorystore, Firestore and Bigtable. The cost and latency comparisons are inside the 11 Retrieval notebooks.

### What license are the notebooks in statmike/vertex-ai-mlops under?

Apache-2.0, with a LICENSE file at the repository root and the same identifier in the repository's license field, so the notebooks can be reused and adapted with attribution.

### Are there releases or tags for the notebooks in this repository?

No. The repository has no GitHub releases, and the last push was on 2026-09-16, so there is no version string to reference when you pin a specific set of notebooks.

## Sources

- [Issues](https://github.com/statmike/vertex-ai-mlops/issues)
- [License: Apache-2.0](https://github.com/statmike/vertex-ai-mlops/blob/main/LICENSE)
- [README](https://github.com/statmike/vertex-ai-mlops/blob/main/README.md)
- [statmike/vertex-ai-mlops on GitHub](https://github.com/statmike/vertex-ai-mlops)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/statmike-vertex-ai-mlops
