# GenMedia Creative Studio: a demo front end for Vertex generative media

> Google's reference application for Gemini, Veo, Lyria and Chirp, built on Mesop and FastAPI, and explicit that it is a demonstration rather than a product.

**GoogleCloudPlatform/vertex-ai-creative-studio** — GenMedia Creative Studio is a generative media user experience highlighting the use of Gemini, Gemini Omni, Veo, Gemini Image 🍌, Gemini TTS, Chirp 3, Lyria and other generative media APIs on Google Cloud.

- Repository: https://github.com/GoogleCloudPlatform/vertex-ai-creative-studio
- Website: https://googlecloudplatform.github.io/vertex-ai-creative-studio/
- Stars: 1,208 · Forks: 372
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/googlecloudplatform-vertex-ai-creative-studio

## Running the app locally with uv and Cloud credentials

The quick start is four commands and one browser tab. The application is a Python app, Mesop on FastAPI, that requires Python 3.14 or newer and uses uv for dependency management. One detail worth catching early: the UI calls Vertex for Gemini, Veo, Imagen and Lyria regardless of where the web server runs, so a local process still needs a Google Cloud project with Vertex AI access and credentials to reach it.

```bash
# 1. Install uv (macOS/Linux) if you don't have it:
curl -LsSf https://astral.sh/uv/install.sh | sh

# 2. Authenticate to Google Cloud (provides Application Default Credentials):
gcloud auth application-default login
export PROJECT_ID=$(gcloud config get-value project)

# 3. Sync dependencies (uv installs Python 3.14 automatically) and run:
uv sync
uv run main.py
```

After that you open localhost port 8080, where the root path redirects to /home. The README sends you to the Documentation Hub for `.env` configuration and the list of environment variables, which is the right place to look because the repository carries both an environment_variables.md file and a dotenv.template at its root.

There is also a Cloud Shell path for people who would rather not install uv locally, and a tutorial.md at the repository root that the Cloud Shell button wires up. The README asks for Google Chrome specifically, warning that some features may not work as expected on Safari or Firefox, which is a fair thing to know before you file a browser bug.

## The model catalogue the UI actually exposes

The feature list is grouped by medium rather than by API, which tells you how the application is organised. Image work covers Gemini 3.1 Flash-Lite Image, also called Nano Banana 2 Lite, Gemini Flash Image Generation, Gemini 3 Pro Image, and Virtual Try-On. Video lists Gemini Omni Flash, Veo 3.1 and Veo 3. Music is Lyria 3 and Lyria 2. Speech is Chirp 3 HD plus Gemini Text to Speech.

Underneath that there is a Workflows group with four named pieces: Character Consistency, Shop the Look, Starter Pack Moodboard, and Interior Designer. These are the parts of the repository most worth reading if you are building something similar, because they are about composition and prompt structure rather than about calling one endpoint. An Asset Library sits alongside them, which suggests generated output is treated as a managed collection rather than as a download.

The repository topics confirm the range: chirp, gemini, gemini-omni, gemini-tts, google-cloud, lyria, nano-banana, veo and vertex-ai. There are 1208 stars, 372 forks and 29 open issues, and the last push was 2026-09-20.

The honest caveat is the one in the description itself. This is a demonstration project, and the README opens with a blockquote saying it is not an officially supported Google product, is not eligible for the Google Open Source Software Vulnerability Rewards Program, and is intended for demonstration purposes only rather than production use. Anyone evaluating it for a real deployment has to take that at face value.

## A boot check you can run before opening a pull request

scripts/smoke_test.sh is a small piece of engineering that more repositories should copy. It runs uv sync, boots the app under gunicorn and uvicorn in APP_ENV=local mode, and verifies that the UI serves by checking that GET / redirects to /home and returns 200, and that GET /__login returns 200.

```bash
./scripts/smoke_test.sh        # boot + UI check
./scripts/smoke_test.sh -l     # also run the live Vertex generation leg
```

The -l flag adds a single live generation call against gemini-2.5-flash, gated on a PROJECT_ID and Application Default Credentials being present. So the script has two tiers: a hermetic boot check anyone can run, and an optional leg that proves the credentials and the model call both work.

The README is careful about what this is not. It says to run it before merging any pull request that touches core-app runtime code, naming pages/, models/, state/, config/ and main.py as the directories that matter, and it says the script is a fast pre-merge sanity check rather than a substitute for the test suite. It is also not a deployment, IAP or Load Balancer validation, since that lives on the Terraform deployment path.

The `__login` endpoint in that check is a hint about how the app handles authentication when it is not behind Identity Aware Proxy, which is one of the deployment options the README describes.

## Deploying to Cloud Run with one worker and eight threads

Deployment is Terraform plus Cloud Run, and the README gives you a genuine choice at the top: a custom domain fronted by Identity Aware Proxy and a Load Balancer, or the autogenerated Cloud Run domain. The step-by-step instructions live in the Deployment Guide on the Documentation Hub.

The Dockerfile is short enough to read in full and it says something about the intended traffic shape. It starts from python:3.14-slim, installs libgl1 and libglib2.0-0 because OpenCV needs them, copies the project in, installs uv, runs uv sync, and exposes port 8080. The command it defines comes from the Procfile at the repository root:

```bash
CMD ["/app/.venv/bin/gunicorn", "--bind", ":8080", "--workers", "1", "--threads", "8", "--timeout", "0", "--forwarded-allow-ips", "*", "-k", "uvicorn.workers.UvicornWorker", "main:app"]
```

A single worker with eight threads and no timeout is a deliberate configuration for an app whose requests are long. Generative media calls do not return in a second, so scaling out processes would mostly buy memory. The repository root also carries cloudbuild.yaml, a build.sh, a terraform.tfvars.example and a deploy/ directory, so the deployment story is version controlled rather than described in a wiki.

The tree is large and tells you the shape of the codebase: pages/, routers/, components/, models/, services/, state/, config/, workflows/, prompts/, experiments/, tools/, test/, docs/ and docs-site/. Two of those deserve a note. prompts/ is a directory rather than strings inside Python files, which is the right call for a project built around prompt iteration. state/ and app_factory.py suggest the Mesop app is assembled through a factory, which matters if you are trying to reuse any of this.

## What the dependency list reveals about the media pipeline

pyproject.toml is the most information-dense file in the repository for anyone deciding whether this codebase is useful to them. It declares version 1.13.1 for the project named vertex-ai-genmedia-creative-studio, pins Python to the 3.14 line with an upper bound, and then lists dependencies that map almost one to one onto the feature list.

The generative core is google-genai, google-cloud-aiplatform and google-cloud-texttospeech. The web layer is mesop on fastapi with uvicorn and gunicorn behind it. firebase-admin is present, which is how a demo app gets user identity without standing up a full account system. google-cloud-tasks is there too, which is the usual way to make a long running media job survive an HTTP request.

The media side is where it gets specific: mediapy for rendering, moviepy for video work, opencv-python for image processing, librosa and praat-parselmouth for audio analysis, numpy, pandas and scipy for the numeric work, and pillow for images. There is also c2pa-python, which is the Content Credentials library, and that choice is a statement about provenance rather than about pixels.

Two of those are held in place deliberately. A uv override block forces pillow at or above 12.3.0 because MoviePy 2.2.1 caps Pillow below 12, and the comment says the override exists because a local audit requires the vulnerability-fixed 12.2.0 line, to be removed once MoviePy publishes compatible metadata. requirements.txt is generated by `uv export` rather than maintained by hand, so the pin resolution lives in one place.

## Experiments, MCP servers, and a scrubbed git history

The experiments/ directory is where this repository stops being an app and starts being a shelf of prototypes: stand-alone applications, new workflows, and Model Context Protocol servers covering combined video generation workflows, advanced prompting techniques, image recontextualization and audio exploration. The Documentation Hub has an Experiments section with architecture diagrams and per-tool installation instructions. The release notes suggest the MCP side is the most actively developed part of the project, since v3.20.0 and mcp-v3.20.0 both shipped on 2026-09-12 and the changelog entries are almost entirely MCP work, including a standalone Gemini 3.5 Transcribe MCP server and a speech-to-text tool in the Gemini MCP server.

That same changelog exposes the project's internal vocabulary. Several entries describe tiered agentic producers, with a Tier 3 preview agent that supports delegation and interrupt, wrapping the generative media MCP tools for sub-agents. So the experiments include not just demos of model capability but scaffolding for letting an agent drive those capabilities, which is a more interesting thing to read than another prompt template.

One operational note deserves attention from anyone with an existing clone. The README carries a notice that the git history on main was scrubbed in August 2026 to remove legacy compiled binaries, which cut clone size by about 60 percent, and it points at reset instructions on the changelog page for synchronizing an existing checkout. A rewritten main branch is the kind of change that quietly breaks CI caches and shallow clones, so it is worth reading that page before you try to merge anything.

The repository also carries AGENTS.md, GEMINI.md, FAQ.md, a developers_guide.md, renovate.json and release-please configuration, which together describe a project with real process around it even if the product itself is a demo.

## Conclusion

GenMedia Creative Studio is best read as a worked example rather than a foundation. It shows how a team wires several generative media models into one interface with Mesop, how it keeps them behind a FastAPI app, and how it ships that app to Cloud Run with Terraform, all in one repository you can read top to bottom. The parts that will save a reader the most time are the workflow modules rather than the model calls, since character consistency and shop the look show how multi-step prompting is organised. What the project cannot give you is a support commitment: it says plainly that it is a demonstration, not for production, and outside the vulnerability rewards program. Start by running it locally with uv sync and uv run main.py against your own project, then use scripts/smoke_test.sh to confirm a change still boots.

## FAQ

### What is Vertex AI Studio used for?

GenMedia Creative Studio is a web application that showcases Google Cloud's generative media models, covering Gemini image generation, Veo for video, Lyria for music, and Chirp 3 HD and Gemini TTS for speech. It exists so you can try those models side by side and study how the workflows that combine them are built, rather than as a product you would run in production.

### What do I need before I can run it locally?

Python 3.14 or newer, uv, and a Google Cloud project with Vertex AI access. The UI calls Vertex for Gemini, Veo, Imagen and Lyria no matter where the web server runs, so `gcloud auth application-default login` and an exported PROJECT_ID are needed even for a local process. After `uv sync` and `uv run main.py`, the app serves on port 8080.

### Is it safe to use this in production?

The README says no. It opens by stating the project is not an officially supported Google product, is not eligible for the Google Open Source Software Vulnerability Rewards Program, and is intended for demonstration purposes only rather than production use. It is Apache licensed open source, so nothing prevents you reading it or adapting it, but you inherit the maintenance burden yourself.

### What are the MCP servers in the experiments folder?

The experiments/ directory holds stand-alone applications, new workflows and Model Context Protocol servers for generative media, covering combined video generation, advanced prompting, image recontextualization and audio. Recent releases added a standalone Gemini 3.5 Transcribe MCP server and a gemini_transcribe speech-to-text tool, and the Documentation Hub carries architecture diagrams and installation steps per tool.

## Sources

- [GoogleCloudPlatform/vertex-ai-creative-studio on GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-creative-studio)
- [License: Apache-2.0](https://github.com/GoogleCloudPlatform/vertex-ai-creative-studio/blob/main/LICENSE)
- [Project website](https://googlecloudplatform.github.io/vertex-ai-creative-studio/)
- [README](https://github.com/GoogleCloudPlatform/vertex-ai-creative-studio/blob/main/README.md)
- [Releases](https://github.com/GoogleCloudPlatform/vertex-ai-creative-studio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/googlecloudplatform-vertex-ai-creative-studio
