Self-hosted service
abi/screenshot-to-code avatar
abi/screenshot-to-code

abi/screenshot-to-code: turning mockups into HTML, Tailwind, React or Vue

Drop in a screenshot and convert it to clean code (HTML/Tailwind/React/Vue).

79,719 stars9,743 forksPythonMIT

At a glance

What is it?
A self-hostable FastAPI and React tool that sends a screenshot or screen recording to a model provider and returns frontend code. The setup cost is real, the model keys are mandatory, and the output quality tracks whichever provider you pay for.
Who is it for?
Adopt abi/screenshot-to-code if you already hold an OpenAI, Anthropic or Gemini key and want a local FastAPI plus Vite app you can modify, or if you need the screen recording to prototype path that most converters do not attempt. Do not adopt it if you want a one-click converter with no keys, no Poetry, no pnpm and no Chromium install, because the README makes all four part of the local path.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

Who abi/screenshot-to-code is actually for

The README frames the project around one job: take a screenshot, a mockup, a Figma design or a screen recording and produce frontend code in one of six stacks (HTML with Tailwind, HTML with CSS, React with Tailwind, Vue with Tailwind, Bootstrap, or Ionic with Tailwind). That list matters more than the pitch. This is not a general image-to-code engine; it is a converter with a fixed set of output targets, and if your project uses Svelte or plain Angular, nothing in the README suggests a path.

The intended user is someone who can run a Python backend and a Node frontend locally. The README splits its Getting Started section into two explicit paths, run locally or use the hosted app at screenshottocode.com, and describes the local path as best for customizing, self-hosting, or contributing. That is a fair description of the audience: engineers who want to modify the prompts, swap models, or keep screenshots off a third-party service. Designers who only want the conversion result are pointed at the hosted product instead.

There is a second audience the README does not name directly: people who want a screen recording converted into a working prototype. The project supports recording a website in action and turning it into a functional prototype, which is a different input class from a static mockup and is the main thing that separates it from the crowded field of screenshot converters.

How the FastAPI backend and Vite frontend fit together

The repository is a two-service application. The README states the app has a React/Vite frontend and a FastAPI backend, and the top-level layout confirms it: separate backend/ and frontend/ directories, plus a docker-compose.yml that builds both. The root package.json declares pnpm workspaces over frontend and backend and a test script that runs frontend tests through pnpm and backend tests through poetry run pytest, so the two halves are tested with different toolchains.

The data flow is: the browser sends an image (or a recording) to the FastAPI backend, the backend calls a model provider, and the generated code comes back to the UI. The backend runs on port 7001 by default, and the frontend dev server serves on 5173. The two are wired by environment variables rather than by a reverse proxy: VITE_HTTP_BACKEND_URL and VITE_WS_BACKEND_URL in frontend/.env.local tell the browser where the backend lives. The WebSocket variable is the interesting one, because it implies the generation streams back rather than arriving as one response, which is what you would expect from a code-generation loop.

Two optional capabilities sit on top of that core. Gemini keys power asset extraction, described as reusing the real logos and images from your screenshot. A separate optional tool, screenshot preview, lets the agent render its own generated page in a headless browser and visually check its work. That second one is enabled by installing Chromium through Playwright, and the README is explicit that if Chromium is missing the app just skips the tool.

Installing it locally and generating your first component

The README requires at least one model provider key from OpenAI, Anthropic or Gemini. It strongly recommends adding Gemini and Replicate on top, because Gemini powers asset extraction and Replicate powers image generation, background removal and image editing. Without REPLICATE_API_KEY, the README says edit_images and remove_backgrounds are unavailable.

The backend uses Poetry. From the backend directory you write a .env file with your keys, install dependencies, install the Chromium browser Playwright needs, activate the environment, and start uvicorn on port 7001.

bash
cd backend
echo "OPENAI_API_KEY=sk-your-key" > .env
echo "GEMINI_API_KEY=your-key" >> .env
poetry install
poetry run playwright install chromium
poetry env activate
poetry run uvicorn main:app --reload --port 7001

On Linux the README points at poetry run playwright install --with-deps chromium instead, noting that it needs sudo or apt to pull the system libraries. The env activate step prints a command you run yourself, typically a source line pointing at the virtualenv.

The frontend is a pnpm project. Install and start it in a second terminal:

bash
cd frontend
pnpm install
pnpm dev

The README says to open http://localhost:5173 to use the app. If you moved the backend off 7001, set VITE_WS_BACKEND_URL in frontend/.env.local to match. Keys can also be entered through the settings dialog behind the gear icon, except Replicate, which the README says must be configured in backend/.env as REPLICATE_API_KEY.

If you would rather not manage two processes, the README gives a Docker path from the repository root. It writes a .env with your OpenAI key and brings up both services, with the app available at http://localhost:5173. The README warns that you cannot develop the application this way, because file changes do not trigger a rebuild.

bash
echo "OPENAI_API_KEY=sk-your-key" > .env
docker-compose up -d --build

The compose file maps the backend port through ${BACKEND_PORT:-7001} and the frontend on 5173, and its comments note that changing the backend port means also changing VITE_WS_BACKEND_URL.

The model key is a hard dependency, not a configuration detail

The most consequential constraint is that this project does not generate code on its own. Every conversion goes to a hosted model provider, and the README's key table makes the requirement unambiguous: at least one of OPENAI_API_KEY, ANTHROPIC_API_KEY or GEMINI_API_KEY is mandatory. That means the true cost of running it is your provider bill plus the setup, not zero. The hosted app exists precisely because that barrier is real for a lot of people.

The README also ties output quality to which keys you hold. With more keys, the app automatically picks a stronger mix of models per variant; with a single key it uses only that provider's models. So a single-key install is not just a narrower feature set, it is a different quality tier, and the README says so without hedging. It calls Gemini and Replicate strongly recommended for screenshot-to-code accuracy, which is an admission that the OpenAI-only and Anthropic-only paths are weaker for this task.

The Ollama path deserves a direct read. The README links to a GitHub comment for running with Ollama open-source models and labels it not recommended due to poor-quality results. That is unusually blunt for a README, and it should be taken at face value: if your reason for self-hosting is to avoid paying a provider, the project's own documentation tells you the free local-model route produces worse output.

Screenshot preview has a quieter failure mode. It is optional, it is enabled automatically once Chromium is installed, and if Chromium is missing the app skips the tool without erroring. You can run the whole stack, get plausible-looking code, and never know that the self-checking step was silently absent. The Settings dialog shows whether screenshot preview is available on your backend, which is the only place the README says you can confirm it.

Where the self-hosted route stops being the right choice

The README's own split is the honest answer to this question. If you want to try the converter, the README calls the hosted app the fastest way with no local setup. If you want to customize, self-host or contribute, run it locally. Choosing the local path for anything other than those three reasons buys you complexity without benefit.

The local path is not light. It needs Poetry, pnpm, a Python environment, a Node environment, and for full functionality a Chromium install that on Linux requires system packages through apt. The Docker route removes the toolchain problem but the README states you cannot develop with it, so it is a deployment convenience rather than a development environment.

There are also gaps in what the README documents. It does not describe a rollback or version-history mechanism for generated output, so if a generation goes wrong you are re-running it rather than reverting. It gives no cost estimate for a generation, which is the number most people would want before pointing it at a large mockup. And it does not document offline operation, because there is no offline mode: without a reachable provider API, the core function does not run. The README does address one related case, noting that if you cannot reach the OpenAI API directly you can set OPENAI_BASE_URL in backend/.env or in the settings dialog, with v1 in the path.

How it differs from v0, Bolt and other prompt-to-UI tools

The closest comparison is Vercel's v0 or StackBlitz's Bolt, and the difference is the input and the deployment model rather than the output. Those tools take a text prompt and generate a UI inside a hosted environment. abi/screenshot-to-code takes an image or a recording and generates code you run yourself. You cannot describe a layout in words and get a result here; the input is visual. That constraint is the point, since a mockup or a Figma frame is already a specification and re-describing it in prose loses information.

The second difference is provider neutrality. v0 is tied to a single vendor's models. This project exposes Gemini, OpenAI and Anthropic variants side by side in one UI, and the README notes that adding all four keys lets you compare multiple models per generation. If comparing model output on the same screenshot is part of your evaluation process, that is a capability the hosted single-vendor tools do not offer.

The third difference is where the code lives. With a hosted generator, your mockup leaves your network. Running this locally means the screenshot goes to whichever provider you configured, but the orchestration, the prompts and the generated files stay on your machine. For teams with unreleased designs, that distinction is often the deciding factor. The counterweight is that you now own the upgrade path, the dependency versions and the provider keys.

Editorial conclusion

Adopt abi/screenshot-to-code if you already hold an OpenAI, Anthropic or Gemini key and want a local FastAPI plus Vite app you can modify, or if you need the screen recording to prototype path that most converters do not attempt. Do not adopt it if you want a one-click converter with no keys, no Poetry, no pnpm and no Chromium install, because the README makes all four part of the local path. Before committing, verify that your chosen provider key works against the exact model names the app offers, and check whether your deployment can install Chromium, since the screenshot preview tool is skipped silently when it is missing.

Frequently asked questions

What is abi/screenshot-to-code?

It is a tool that converts screenshots, mockups, Figma designs and screen recordings into frontend code using AI models. It supports HTML with Tailwind, HTML with CSS, React with Tailwind, Vue with Tailwind, Bootstrap, and Ionic with Tailwind, and runs as a FastAPI backend with a React/Vite frontend.

Is abi/screenshot-to-code free?

The source is MIT licensed and the hosted app at screenshottocode.com is offered as the fastest way to try it, but running it locally requires at least one paid model provider key from OpenAI, Anthropic or Gemini. The project itself does not generate code without a provider.

How do I turn a picture into code with abi/screenshot-to-code?

Start the FastAPI backend on port 7001 and the frontend with pnpm dev, open http://localhost:5173, and upload the image. The backend sends it to the model provider you configured and returns code in the stack you selected.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/abi-screenshot-to-code.svg)](https://hysenlabs.com/projects/abi-screenshot-to-code)
Community notes

Community notes