ScreenCoder: a multi-agent pipeline that turns UI screenshots into editable HTML/CSS
ScreenCoder — Turn any UI screenshot into clean, editable HTML/CSS with full control. Fast, accurate, and easy to customize.
At a glance
- What is it?
- ScreenCoder is a research system from CUHK MMLab that converts a UI screenshot into HTML/CSS through four separate Python stages. Here is what each stage does, how to run it, and why the multi-step workflow is both the point and the main cost.
- Who is it for?
- ScreenCoder is worth adopting if you already have an API key for Doubao, Qwen, GPT or Gemini and you want to study or extend a staged visual-to-code pipeline, including the released SFT and RL post-training code and the ScreenBench benchmark.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap ScreenCoder targets: screenshots that need to become editable markup
A screenshot is a flat bitmap. A front-end codebase is a tree of elements with classes, nesting and styling rules. Going from one to the other is the problem ScreenCoder addresses, and the README frames the audience explicitly: developers and designers who want to prototype quickly or build pixel-perfect interfaces. The intended output is not a static image export but "clean, production-ready HTML/CSS code" that a person can then edit. That distinction matters, because a system that emits a single flattened div with a background image would technically satisfy "image to HTML" while being useless to anyone who wants to change a button colour. ScreenCoder's answer is a modular multi-agent design: visual understanding, layout planning and code synthesis are separate stages, and the repository layout reflects that split with distinct scripts for detection, generation, mapping and replacement. The project is a research artifact rather than a product. It was accepted to the EMNLP 2026 Main Conference, and the repository ships a paper link, a Hugging Face demo, post-training code for SFT and RL, and ScreenBench, a benchmark of 1000 real-world web screenshots paired with their HTML source. If you want a hosted demo, the README points to the Hugging Face Space. If you want to understand or modify the pipeline, you clone the repo.
How the four-stage pipeline works, and why placeholders sit in the middle
The architecture is visible in the top-level files. block_parsor.py performs block detection, then html_generator.py produces an HTML layout in which image regions are represented by gray placeholder blocks rather than real assets. That intermediate state is the key design decision. The generator is not asked to reproduce pixels; it is asked to lay out structure, and the visual content is deferred. The second half of the pipeline resolves those placeholders. image_box_detection.py finds the placeholder boxes in the generated page, UIED/run_single.py runs the UIED element-detection engine over the original screenshot, and mapping.py aligns the two sets of boxes. Finally, image_replacer.py swaps each placeholder div for the cropped image taken from the source screenshot. The README also lists mapping.py as the component that "maps the detected UIED components to logical page regions", which is where the two coordinate systems meet. The appeal of this decomposition is debuggability: if the final page looks wrong, you can inspect the placeholder layout before any image is inserted, which isolates whether the fault is in generation or in alignment. The cost is that every stage is a separate program with its own inputs and outputs, and the README does not document an intermediate file format, a resume mechanism, or an exit-code convention. Treat the stages as a chain you drive by hand, not as a job runner.
Installing ScreenCoder and running your first screenshot
The README gives a standard clone, virtualenv and pip sequence. Note that the clone command and the directory it changes into do not match in the README: the repository is ScreenCoder, while the cd line reads screencoder. Use the directory that actually exists after cloning.
git clone https://github.com/leigest519/ScreenCoder.git
cd ScreenCoder
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtThe dependency list is heavy. Beyond Pillow, beautifulsoup4, requests, numpy, scipy, scikit-learn, pandas, tqdm and opencv-python, requirements.txt pulls in paddlepaddle and paddleocr for the detection side, tensorflow and keras, playwright, and four model SDKs: volcengine-python-sdk[ark] for Doubao, openai, google-generativeai and the Qwen path. Expect a long install and expect the Paddle and TensorFlow wheels to be the parts most likely to fight your Python version.
Before running anything, choose a model. The README states that you set the desired model in block_parsor.py and html_generator.py, with Doubao as the default and Qwen, GPT and Gemini supported. The key goes in a plain-text file in the project root named after the model.
# create the file matching your chosen model, e.g. for Doubao
doubao_api.txtThe README notes that doubao_api.txt should be kept private and is listed in .gitignore. The same naming pattern applies to qwen_api.txt, gpt_api.txt and gemini_api.txt.
For a first run, the README offers a shortcut that executes the whole chain:
python main.pymain.py is described as the main script that generates final HTML code for a single screenshot. If you want to see the intermediate state instead, run the stages in order: block_parsor.py, html_generator.py, image_box_detection.py, UIED/run_single.py, mapping.py, image_replacer.py. The gray placeholder blocks after html_generator.py are what you should look for before the replacement stage runs. The README does not document where the input screenshot is placed or how it is passed to main.py, so check the script arguments in the repository before assuming a path convention.
The dependency and API-key surface is the real adoption cost
requirements.txt is the clearest limitation in the repository. It pins minimum versions rather than exact ones, and it spans two large ML stacks (Paddle and TensorFlow) plus a browser automation library, all in a single flat list. For a research release that is normal; for a tool you want in CI it is a liability, because a transitive bump in paddleocr or tensorflow can change detection output without any change to ScreenCoder itself. The API-key design has a similar character. Keys live in plain-text files at the project root, one per provider, and the model selection is made by editing block_parsor.py and html_generator.py rather than by an environment variable or a config file. That means switching from Doubao to Gemini is a source edit, and the README does not describe a way to override the choice at runtime. There is also no released version: the README's News section announces the EMNLP acceptance and the release of post-training code and ScreenBench, but no tagged releases were retrieved, so the practical way to consume ScreenCoder is to track the main branch. The last push to the repository was on 2026-09-15. If you need reproducible builds, pin the commit yourself, because the project does not do it for you.
Where the staged design breaks down
The placeholder round trip is the failure surface. Stage one decides where an image block belongs; stage two tries to find that block in the rendered page, detect the corresponding element in the original screenshot, and match the two. Any mismatch leaves either a gray box in the final HTML or a cropped image inserted in the wrong place. The README does not describe a fallback for unmatched placeholders, a confidence threshold for the mapping, or a report of how many boxes were resolved. That silence is the thing to test before trusting the output on a non-trivial page. There is a second, more basic constraint: the pipeline is built for web UI screenshots. The example material covers a YouTube page, an Instagram page and a design draft, and ScreenBench consists of web screenshots with corresponding HTML source. Nothing in the README suggests native mobile layouts, canvas-rendered applications, or interfaces whose meaning depends on interaction states, so a screenshot of a native iOS settings screen is outside what the documented workflow addresses. Finally, because generation goes through a hosted model, the quality of the HTML depends on the provider you select, and the README does not compare the four supported models or state which one the published results use. If visual fidelity is the acceptance criterion, that comparison is on you.
ScreenCoder against Design2Code and DCGen
The README lists Design2Code and DCGen among the projects ScreenCoder builds upon, and both appear in the related searches. The difference is in the shape of the system. Design2Code is the reference point for the visual-to-code evaluation task itself, and ScreenCoder's contribution is a modular pipeline with an explicit layout-planning stage and a placeholder-then-replace asset step, plus the ScreenBench benchmark of 1000 real-world screenshots. DCGen comes from the WebPAI group, which the README also credits as a source of research resources and datasets for webpage generation, and ScreenCoder's acknowledgement places DCGen in the same lineage of decomposition-based approaches. The practical distinction for a reader choosing between them is what you get out of the box: ScreenCoder ships runnable scripts for detection, generation, mapping and replacement, a Hugging Face demo, and post-training code for SFT and RL. If your goal is to evaluate or reproduce a visual-to-code method, ScreenBench and the released training code are the concrete artifacts to look at. If your goal is a maintained library with versioned releases and a documented API, none of these research repositories is that, and the correct alternative is a commercial or hosted conversion service rather than a sibling paper.
Licence and what Apache-2.0 does not cover
ScreenCoder is released under Apache-2.0, which permits commercial use, modification and redistribution provided the licence and notices are preserved. Two things sit outside that grant and are worth checking before you ship anything. First, the pipeline calls external model providers, and your use of Doubao, Qwen, GPT or Gemini is governed by those providers' terms, not by Apache-2.0; the API keys you place in doubao_api.txt and its siblings are your responsibility to keep private, as the README itself notes. Second, the README credits UIED, DCGen and Design2Code as prior work, and the UIED directory is a vendored copy of a separate project. The repository's LICENSE file covers ScreenCoder's own code; whether the vendored UIED code carries its own terms is something the README does not state, so read the files under UIED/ directly. This is a description of what the licence text and README say, not legal advice.
Editorial conclusion
ScreenCoder is worth adopting if you already have an API key for Doubao, Qwen, GPT or Gemini and you want to study or extend a staged visual-to-code pipeline, including the released SFT and RL post-training code and the ScreenBench benchmark. It is the wrong tool if you need a one-command converter with a documented error contract: the README does not specify what happens when UIED finds no components, when a placeholder fails to match, or how to resume after a single stage fails. Before committing, verify three things on your own hardware: that paddlepaddle and paddleocr install cleanly on your Python version, that block_parsor.py and html_generator.py point at the model you actually have a key for, and that your screenshots survive the placeholder round trip through image_box_detection.py, UIED/run_single.py and mapping.py without orphaned divs.
Frequently asked questions
What is ScreenCoder and what does it produce?
ScreenCoder is a UI-to-code generation system that transforms a screenshot or design mockup into HTML/CSS. It uses a modular multi-agent architecture combining visual understanding, layout planning and code synthesis, and the README describes the output as clean, production-ready and editable.
How do I install ScreenCoder?
Clone the repository, create a virtual environment, activate it and run pip install -r requirements.txt. The README then asks you to set the model in block_parsor.py and html_generator.py and to place an API key in a plain-text file such as doubao_api.txt at the project root.
Which models and API keys does ScreenCoder support?
The README lists Doubao as the default and also supports Qwen, GPT and Gemini. You choose the model by editing block_parsor.py and html_generator.py, and you create the matching key file (doubao_api.txt, qwen_api.txt, gpt_api.txt or gemini_api.txt) in the project root.
Can I try ScreenCoder without installing it?
Yes. The README links a Hugging Face demo space, and it also notes that you can run the demo locally by downloading it from the Hugging Face space and executing python app.py.
What is ScreenBench?
ScreenBench is a benchmark released alongside ScreenCoder for visual-to-code and web UI generation. The README describes it as 1000 up-to-date real-world sampled web screenshots with corresponding HTML source code across diverse topics, hosted on Hugging Face.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/leigest519-screencoder)