# screenshot-to-code needs Replicate in a file the UI cannot edit, and ships no version

> screenshot-to-code turns screenshots, mockups, Figma designs, and screen recordings into code across six stacks, backed by several model providers. The code is MIT licensed, but every conversion costs provider tokens, one key can only be set outside the settings dialog, and one of the tools degrades silently when a browser is missing.

**abi/screenshot-to-code** — Drop in a screenshot and convert it to clean code (HTML/Tailwind/React/Vue).

- Repository: https://github.com/abi/screenshot-to-code
- Website: https://screenshottocode.com
- Stars: 79,719 · Forks: 9,743
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/abi-screenshot-to-code

## The free part is the code, not the inference

The licensing is MIT and the repository is readable and forkable, but nothing about the tool is free to run. You need at least one model provider key, and the three named are OpenAI, Anthropic, or Gemini, with Replicate added on top. Gemini and Replicate are both strongly recommended, and the reason given is not stylistic: Gemini powers asset extraction, which is how the tool reuses the real logos and images from your screenshot rather than inventing them, and Replicate powers image generation, background removal, and image editing. The local-model escape hatch is explicitly discouraged, with Ollama named as not recommended due to poor-quality results. What the project cannot do is give you a conversion without a paid provider behind it. The readme frames the hosted site as the easiest way to try it, which is the tell: self-hosting moves the cost rather than removing it.

## Replicate is the one key the settings dialog cannot set

There are two ways to configure keys, and they do not cover the same set. You can set up OpenAI, Anthropic, and Gemini keys in the frontend settings dialog, reached by clicking the gear icon after the app loads. Replicate is the exception: it must be configured in backend/.env as REPLICATE_API_KEY, and there is no UI path for it. What that cannot do is fail loudly. The documented consequence of the missing key is that edit_images and remove_backgrounds are unavailable, so an operator who configures everything through the settings dialog gets an install that looks complete and quietly has two image features switched off, with nothing in the interface telling them why those features do nothing. The settings dialog does report one thing, whether screenshot preview is available, which makes the silence on Replicate all the more confusing. Add that key to the file before you judge the tool.

## Without Chromium the agent stops checking its own work

The screenshot preview tool is the part of the pipeline that sounds most valuable, because it lets the agent render its own generated page in a headless browser and visually check the result. It is optional, and it is enabled automatically once Chromium is installed, either through the playwright install chromium step in the local setup or automatically in the Docker image. The failure behavior is what matters. If Chromium is missing, the app just skips the tool, and the settings dialog is what tells you whether it is available. What this cannot do is report that your output got worse. There is no error, no warning in the log, and no degraded banner, because a missing browser is indistinguishable from a successful run that happened to be correct. So a reader evaluating output quality who skipped the Chromium step is measuring the tool without its self-check and has no signal that they did. Check the settings dialog first.

## Video mode and asset extraction both hard-require a Gemini key

Beyond generating code, the tool can take a screen recording of a website in action and turn it into a functional prototype, and it can lift the real assets out of a screenshot rather than approximating them. Both of those capabilities are attached to one provider in the key table. The Gemini row is marked as one of the three optional keys but also as strongly recommended, and it is the row that carries two hard dependencies: it extracts real assets from the screenshot, and it is required for video mode. What that cannot do is be substituted. A team standardized on OpenAI can generate code fine, and the OpenAI row only unlocks the GPT code-gen variants of GPT-5.5 and GPT-5.4 Mini, but the recording workflow and the asset reuse are simply unavailable without adding a Google key. Model choice and feature availability are entangled here in a way the key table makes easy to miss.

## Docker gives you a running app you cannot develop

The container path is two commands from the root directory:
```bash
echo "OPENAI_API_KEY=sk-your-key" > .env
docker-compose up -d --build
```
and it leaves the app at http://localhost:5173. The compose file defines a backend built from ./backend with a Dockerfile, reading configuration from an env_file of .env, publishing the port through a defaultable BACKEND_PORT, and starting the server with poetry run uvicorn main:app bound to 0.0.0.0. A frontend service builds from ./frontend and publishes 5173 on 5173. The project states the limitation plainly: you cannot develop the application with this setup, as file changes will not trigger a rebuild. What the container path cannot do is shorten your edit loop. It is for evaluation, and note the asymmetry in the two services: the backend port is parameterized, the frontend port is not, so a taken 5173 has no override, and the backend binding to 0.0.0.0 publishes it on every interface.

## The frontend reaches the backend over two env var families, with no auth described

The local setup runs the backend on port 7001 and the frontend on 5173, and the two are coupled through environment variables in frontend/.env.local. If you want the backend on a different port you update VITE_WS_BACKEND_URL, and to change the host the frontend connects to you configure both VITE_HTTP_BACKEND_URL and VITE_WS_BACKEND_URL, with the example given being a VITE_HTTP_BACKEND_URL pointing at a specific address on port 7001. The compose file carries the same warning in a comment: if you change the port, make sure to also change the VITE_WS_BACKEND_URL at frontend/.env.local. What this cannot do is stay quiet about the consequence. There are two channels here, plain HTTP and a websocket, and the documentation does not describe any authentication on either. Combined with the backend command binding to 0.0.0.0, a port left at its default on a shared machine exposes an API that spends your provider keys.

## The two model lists disagree, and there is no version to pin

The project lists its default models in one place and its per-key unlocks in another, and the two do not line up. The default list names Gemini 3 Flash Preview and Gemini 3.1 Pro Preview, GPT-5.5 and GPT-5.4 Mini, and Claude Opus 4.6 and Claude Opus 4.8. The Anthropic row in the key table instead lists Opus 5, Opus 4.8, Fable 5, and Sonnet 4.6. Neither list is marked as authoritative, so a reader trying to work out which Claude variants actually exist in the app has to guess. The versioning story is thinner still. The repository has no GitHub releases, the manifest is marked private with a version of 0.0.0, and there is no changelog among the top level files. What this cannot do is let you hold a working configuration still. Choosing an output stack is a decision the tool makes easy, since it supports HTML with Tailwind, HTML with CSS, React with Tailwind, Vue with Tailwind, Bootstrap, and Ionic with Tailwind, but choosing a version of the tool is not a decision it offers at all.

## Conclusion

screenshot-to-code fits a team that already pays for model access and wants a screenshot turned into a starting point in a stack of its choice, and that is willing to wire a FastAPI backend to a Vite frontend. It does not fit anyone expecting a local, offline, or zero-cost tool, or anyone who needs a version they can pin. Before you commit, set the Replicate key in backend/.env rather than in the UI, confirm the screenshot preview is reported as available, and make sure the backend port is not reachable from a network you do not control.

## FAQ

### what is screenshot to code

screenshot-to-code is an MIT licensed open-source tool that converts screenshots, mockups, Figma designs, and screen recordings into clean functional code. Supported output stacks are HTML with Tailwind, HTML with CSS, React with Tailwind, Vue with Tailwind, Bootstrap, and Ionic with Tailwind, and it also turns a screen recording of a site in action into a functional prototype.

### is screenshot to code free

The code is free, since the project is MIT licensed and the repository has no paid tier. Running it is not free, because you need at least one model provider key from OpenAI, Anthropic, or Gemini, and the project strongly recommends adding Replicate for image generation, background removal, and image editing.

### How do I turn a picture into code with screenshot-to-code?

Run it locally or use the hosted app, drop in the image, and pick one of the six supported stacks. Locally the setup is a FastAPI backend on port 7001 and a React and Vite frontend on port 5173, with the app available at http://localhost:5173 once both are started.

## Sources

- [Official documentation](https://screenshottocode.com)
- [Official README](https://github.com/abi/screenshot-to-code#readme)
- [Project repository](https://github.com/abi/screenshot-to-code)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/abi-screenshot-to-code
