Model or dataset
boyang-hu/website-rebuild-skill avatar
boyang-hu/website-rebuild-skill

website-rebuild-skill: an Agent Skill that rebuilds a site from its own minified code

复刻网站的 Agent Skill:抓只读镜像、从压缩代码逐行还原、自动比对验收。An agent skill that mirrors a website, rebuilds it from the minified code, and verifies the result with automated diffs.

1,342 stars247 forksJavaScriptMIT

At a glance

What is it?
The skill mirrors a site as read-only evidence, ports its minified bundles line by line, then proves the result with console, network, DOM, geometry and pixel diffs. It is built for engineers who need a rebuild that can be verified, not a page that merely looks close.
Who is it for?
Adopt this skill when you need a rebuild you can defend with evidence: an archived site, a WebGL or Canvas scene, a static builder output, or a platform-layer storefront where the source is the specification. Do not adopt it when you want a quick visual approximation, a redesign, or a tool that makes the copyright call for you, because the skill only gathers evidence and hands the decision to you.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 25 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: a rebuild you cannot prove correct

Most attempts to reproduce a website stop at resemblance. Someone opens the page, looks at the two versions side by side, and declares it close enough. That works until the site has a WebGL scene driven by a physics simulation, a scroll timeline with nine checkpoints, or a lazy-loading placeholder that never resolves. At that point resemblance is not a standard, and every later change is a guess.

website-rebuild-skill treats the source site as the specification. The README states the goal directly: it is not "making a page that looks very similar", it takes the source implementation as the spec, captures the whole site as read-only evidence, restores logic line by line from minified code, and then proves with a set of automated acceptance gates that the rebuild and the source are doing the same thing, including pixel-by-pixel comparison.

The intended user is an engineer with a supported agent runtime who needs a faithful port: recovering a site that is about to disappear, studying how a production WebGL engine is assembled, or separating a Shopify store into platform, app, upstream theme and site-specific layers. It is a poor fit for anyone who wants a new design.

Mirror-first forensics, then a line-numbered port

The pipeline is a fixed chain of stages, and the ordering matters more than any single step.

Grading comes first. The agent probes what class of site it is looking at and whether the job is feasible. If it is not, the README says it states the reason and does not force a run that produces garbage. A site that has disappeared but has an archive is routed into a Wayback rescue path: pick a coherent moment by anchor plus time window, land the original bytes as a standard mirror, and register permanent holes honestly rather than filling them in.

Then the mirror. The whole site is captured as a read-only snapshot with a per-file sha256 ledger, reference-closure checks and authenticity checks. This mirror is the only evidence baseline for the entire project, and the first of six disciplines says it is never modified.

Then the reverse pass. Minified code is expanded so that every line in the rebuild can point back to a specific line in the source bundle. The README is explicit that bugs and strange idioms are copied, not fixed, because every oddity in minified code may be the behaviour itself.

Then the port, then acceptance, then closure. Acceptance compares both sides across five layers: console, network, DOM, geometry and pixels. Differences are either fixed or registered with an explanation of what the source does, what the rebuild does and why. An unregistered difference counts as a bug.

The final stage, source-ification, is deliberately last. The README's reasoning is that refactoring is dangerous when you cannot tell whether you broke something; here a pixel-exact referee already exists, so every split and rename can be falsified. Its stated summary is blunt: refactoring without a referee is blind editing.

Installing the skill with npx and running a first rebuild

Prerequisites are listed in a table: Node.js 22 or newer (the built-in fetch and WebSocket connect directly to CDP), a locally installed Chrome or Chromium for headless comparison, and npx for the few stages that spawn version-pinned external tools without importing them.

The installer copies files only. It does not run the skill, open a browser or touch the network, and it has no dependencies of its own. After copying it verifies every file in both directions by sha256, and only a full match counts as success. If the target already exists it prints the installed version and refuses, requiring --force to replace. --dry-run shows what it intends to write.

bash
npx website-rebuild-skill
npx website-rebuild-skill --project
npx website-rebuild-skill --dir <skills dir>

The three forms place the skill in the user-level skills directory, the project-level skills directory, or another runtime's skills directory, where the directory name website-rebuild is appended automatically.

If you have already cloned the repository, the manual route is to copy the skill directory into your agent's skills directory. The README gives the Claude Code example.

bash
cp -R skills/website-rebuild ~/.claude/skills/website-rebuild

Other runtimes that follow the Agent Skills specification use the same directory under their own convention. The README claims cross-runtime operation is measured rather than asserted: the verification checklist includes targets completed by Claude Code and targets completed end to end by Codex, one of them with 166 of 166 responses byte-identical.

Using it is a sentence, not a command. Give your agent a URL and say "rebuild this site" or "1:1 rebuild this website". The agent will come back with questions: the scope (whole site or specific pages), the level to reach (L1 mirror archive, L2 engineered rebuild, L3 source-ification, where the ladder is monotonic so choosing low costs nothing and you can resume upward later), how to handle external dependencies, and every judgement about whether something may be published. The README frames these as your decisions, not the agent's.

On timing, the README says schedules have converged across versions: small and medium sites, within a few dozen routes, now run unattended from grading through source-ification in a single session, while heavy WebGL or custom-engine sites take roughly one to three days. The agent gives a difficulty rating and estimate before starting.

What the acceptance gates actually compare

The five comparison layers are console, network, DOM, geometry and pixels. Pixel comparison uses a mean absolute difference figure, and the README reports 0.00 for several showcase pairs, meaning the two frames matched exactly at the same viewport, scroll position and animation moment.

The showcase examples are worth reading as boundaries rather than trophies. For lusion.co, a 1.25 MB custom WebGL engine with 156 shaders, the README notes that the pose of the 3D physics stack comes from the simulation itself, so the simulation's random state had to be pinned before both sides matched pixel for pixel. That is a real constraint: determinism has to be established before comparison means anything.

Hubtown is reported differently. It is a Theatre.js full-screen WebGL long take on Nuxt 3 and three.js, rendered live on both sides, and the pixel difference lands within a same-side noise band at meanAbsDiff 0.96. Samsy is a WebGPU scene with 238 MB of heavy assets, rendered live on both sides at the same moment, with film grain that belongs to the scene. The README does not claim 0.00 there, and that honesty is more useful than a uniform success claim would be.

The Raycast keyboard example carries the clearest lesson about discipline. The README points to an unloaded lazy-loading placeholder bar above the keyboard and notes that both sides are identical, because bugs and odd states are copied rather than fixed. If you want a rebuild that silently repairs what the original got wrong, this tool is pointed the other way on purpose.

Where the approach breaks down

The most concrete limitation is server-rendered component sources. For React Server Components and the Next.js App Router, the server component source is not shipped to the client. The skill's answer is to treat the complete output, the flight stream inlined in each page's HTML, as the specification and reconstruct a buildable Next project from it, closing with a flight-semantics gate. The README reports one Next 16 / Turbopack blog site at 18 of 18 routes semantically consistent and a 144-route heavy site passing 144 of 144, plus a blind reverse-engineering exercise against public source scoring roughly 95 percent on structure and 98 percent on behaviour. Those are reconstructions from output, not recoveries of the original source tree, and the difference matters if you were hoping to obtain the author's actual files.

Dead sites are another hard boundary. The README describes five Wayback rescues: four revived, one of which reached L3, and one where the visual layer was confirmed entirely lost, with the failure shape recorded. A site with no coherent archive window cannot be rescued at all.

Discipline three, that everything the source has must be present and nothing it lacks may be invented, has a stated boundary at the source-ification stage. Renaming, splitting modules and adding comments in the human-readable source do not count as invention, because that copy is an explicitly registered derivative rather than a claim about the source. Two lines stay fixed: no opportunistic refactoring (merging duplicates, extracting shared functions or changing algorithms all make equivalence undecidable), and speculation in comments must be labelled as speculation.

Finally, the skill does not make legal calls. The README states that the legal decision belongs to the user, that the skill only gathers evidence and presents it, and that output defaults to private, noindex and undeployed.

How it differs from a site scraper

A conventional site scraper or mirroring tool, such as wget's recursive mode or a headless-browser screenshot pipeline, produces an artifact and stops. You get HTML, assets and maybe a rendered image. Nothing in that workflow tells you whether the copy behaves like the original, and nothing records the differences you chose to accept.

This skill inverts the emphasis. The mirror is not the deliverable; it is the evidence baseline that the deliverable is measured against. Every deliberate difference must be registered with three fields: what the source does, what the rebuild does, and why. An unregistered difference is treated as a bug by definition. A scraper has no concept of an unregistered difference.

The second difference is the source-ification stage, which a scraper does not have at all. Verbatim-ported output is rewritten into a readable project: modules split, variables named from evidence, provenance comments added, assets copied. When the artifact has no module container, the README describes a concatenative decomposition: cut into semantically named parts and reassemble in order so the result is byte-equal to the original. The stated outcome is a source project that can be copied anywhere and run offline, with a per-file byte self-check before every start.

The third difference is scope of comparison. A scraper verifies that files downloaded. This verifies console output, network behaviour, DOM structure, geometry and pixels. The README's framing of the goal is that it watches whether the two sides are correct, not whether they look alike.

Maintenance, licence and what a fork inherits

The repository is not archived, and the last push was on 2026-09-07, which is recent enough that the codebase is moving. There are no retrieved releases, so version tracking happens through the package manifest, which lists version 0.3.23, and through CHANGELOG.md in the repository root. Upgrading means re-running the installer, which will refuse to overwrite an existing installation without --force and will verify the new copy by sha256 in both directions. That refusal is a feature: it prevents a silent half-upgrade of a skill whose behaviour depends on many small scripts.

The licence is MIT, declared in package.json and accompanied by a LICENSE file and an MIT badge in the README. MIT covers the skill's own code. It does not cover the sites you point it at, and the README separates the two questions explicitly, noting that whether something can be done and whether it should be published are different matters. Nothing here is legal advice, and the skill's own default is private, noindex output that is not deployed.

For anyone who forks the output, the README points to a handoff document, skills/website-rebuild/references/beyond-the-rebuild.md, and describes the one thing the skill deliberately leaves behind: the byte manifest and reassembly gate travel with the deliverable, so after forking you can tell precisely where each step departed from the source. The skill does not name your project, write its story or replace its content, because those decisions are yours.

Editorial conclusion

Adopt this skill when you need a rebuild you can defend with evidence: an archived site, a WebGL or Canvas scene, a static builder output, or a platform-layer storefront where the source is the specification. Do not adopt it when you want a quick visual approximation, a redesign, or a tool that makes the copyright call for you, because the skill only gathers evidence and hands the decision to you. Before starting, verify that Node.js is 22 or newer, that Chrome or Chromium is installed locally for headless comparison, that the target is still reachable or has a coherent Wayback capture, and that the output stays private and noindex until you decide otherwise. The one thing to check first is the difficulty grade the agent returns before any work begins: a site it refuses to grade is a site you should not force.

Frequently asked questions

What runtime does website-rebuild-skill need?

It follows the Agent Skills open specification and is designed for any agent that supports skills. The README states that the same skill directory runs the full pipeline in either Claude Code or Codex, and the installer's --dir flag places it in another runtime's skills directory.

Does website-rebuild-skill install npm dependencies?

No. The README states the toolchain has zero dependencies and that the entire pipeline installs no npm packages before the source-ification stage. The 75 Node scripts are split into 51 stage and acceptance scripts, 14 shared libraries and 10 source-ification or reverse-reconstruction tools.

Can website-rebuild-skill rebuild a site that no longer exists?

It has a Wayback rescue path that selects a coherent moment by anchor plus time window, lands the original bytes as a standard mirror, and registers permanent holes honestly. The README records five such rescues: four revived, one of which reached L3, and one where the visual layer was confirmed entirely lost.

Does website-rebuild-skill decide whether I am allowed to publish the rebuild?

No. The README states the legal decision belongs to the user and that the skill only gathers evidence and presents it. Output defaults to private, noindex and undeployed.

Official sources

  1. boyang-hu/website-rebuild-skill on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/boyang-hu-website-rebuild-skill.svg)](https://hysenlabs.com/projects/boyang-hu-website-rebuild-skill)