The AI-native hiring guide is one HTML file with a five-minute role sort
AI-Native 工程师招聘面试官手册 · Builder / Reviewer 双岗评分 · AI 驾驶能力实操考核 · 易哈佛医疗
At a glance
- What is it?
- A single-file interview manual from a healthcare company's internal hiring process, built around a Builder and Reviewer split, 14 questions in three scoring tiers, and six knockout rules. The scoring arithmetic is precise and inspectable, while the knockout criteria and the 14 questions themselves live inside the HTML where you cannot audit them from the repository.
- Who is it for?
- This manual is worth reading as a template if you are redesigning technical interviews, and worth copying only after you rewrite the parts the repository does not show you. The schedule, the role split and the pass thresholds are visible and arguable, but the 14 questions, the three scoring tiers and the six knockout criteria are all inside the HTML, so audit those against your own bias risks before you run them on a real candidate.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three files, one of which is the entire manual
The top-level tree is a README, an index.html and a preview image. The manual, the 14 interview questions, the scoring rubric, the interviewer support material and the printable score sheet all live inside that single HTML file, described as zero dependency and openable directly in a browser. Local use is two options:
# 直接打开
open index.html
# 或者用任意 HTTP 服务器
npx serve .There is no build step, no package manifest and no test. The published copy is served from a GitHub Pages address, and the source of the online version and the local file are the same document.
That structure is a real virtue for a document like this. A hiring manual that can be emailed, archived and opened offline has no supply chain, no toolchain to keep current and nothing to break when a dependency is deprecated. The cost is equally concrete: there is no diff-friendly source, so a question can only be versioned by rewriting the file, and two interviewers comparing versions are comparing opaque bytes rather than changed sentences.
The pass line is 35 out of 50 and the bands leave no room
Each candidate is assessed across five modules worth ten points each, for a maximum of 50. Three bands follow: PASS at 35 or above, which is stated as 70 percent, with no veto item triggered; HOLD at 30 to 34, recorded as a backup; and FAIL below 30, or any veto item triggered regardless of score.
The arithmetic is tight by design. A candidate scoring 34 is a hold and a candidate scoring 35 passes, so a single point on a single module decides the outcome. A candidate at 30 is still a hold rather than a fail, which means the floor of the pass band sits 5 points above the floor of the hold band.
The veto path is the more interesting mechanism, because it makes the score irrelevant in one direction. Six hard knockout criteria are described as a guard against subjective bias, and triggering any one of them sends a candidate to FAIL no matter how high the total is. That is a strong structural protection, since it denies a favorable impression any purchase over a disqualifying signal.
The knockout criteria that guard against bias are not enumerated
The bias protection is stated as a feature and counted as a feature: six hard elimination criteria exist to prevent subjective judgement from deciding outcomes. They are not listed in the README, and they are not in the repository as a separate file either, which means they live inside the index.html alongside the questions.
That has two consequences. An organization adopting the rubric cannot audit the guardrails before using them, because the only way to read them is to open the HTML and find them. And a candidate-facing process has six ways to be eliminated that are not documented in the handbook itself, which is the opposite of what an anti-bias mechanism should look like when the handbook is the artifact that gets shared.
The same pattern applies to the scoring tiers. Each question carries criteria in three tiers, 5, 3 and 1, so the scale is a coarse one where the middle is a single point of 3 and the bottom is a single point of 1. Whether a reviewer reliably distinguishes a 5 from a 3 on an issue-writing task is a calibration question, and no calibration procedure, rubric training or inter-rater agreement check appears anywhere in the visible documentation.
Role typing happens in the first five minutes and then fixes the test
The 60-minute schedule starts with five minutes of icebreaking and typing, whose stated purpose is to decide whether the candidate is a Builder, a Reviewer or the dual type. Everything after that is shaped by the verdict: the block from minute 30 to 45 is a specialized module, and it is either the Builder set or the Reviewer set.
So a five-minute impression sets the agenda for the next forty. There is no block in the schedule for revisiting the typing, and no stated procedure for what to do when a candidate answers as a Builder in the opening and then reviews well in the specialized block, or the reverse. The dual path is offered as a third outcome, but the schedule still contains only one specialized block, so a dual candidate is assessed against one of the two sets rather than both.
The role definitions are also asymmetric in a way that matters for scoring. A Builder drives output and a Reviewer guards quality, and the whole thesis rests on the idea that these are different scarce abilities rather than senior and junior versions of one job. Under that framing, typing someone into the wrong role is not a formatting error, it is measuring the wrong thing for forty minutes.
Fourteen questions against roughly thirty minutes of question time
The advertised inventory is 14 structured questions, each with a junior and a senior version and each carrying 5/3/1 scoring criteria. The schedule allocates 15 minutes to the general A and B segment covering issue writing and solution review, 10 minutes to the general G segment for the hands-on AI driving exercise, and 15 minutes to the specialized module.
Set against the schedule, 14 questions have to fit inside the segments where questioning happens, and the final five minutes are reserved for scoring on the spot. That leaves roughly thirty minutes of question time across three segments, or a little over two minutes per question, and that figure is before setup time for the hands-on exercise and before any candidate questions. The manual also ships interviewer backup material, described as ready-made solution summaries and code diffs usable by non-technical interviewers, which implies those artifacts are read or presented during the session rather than skimmed beforehand.
The hands-on segment is the one to watch, since its stated purpose is live observation of how a candidate drives an AI agent. Observation of behavior is hard to compress, and a two-minute average hides that the practical segment plausibly eats its own budget and pushes the question count down.
The thesis is that writing code stopped being the scarce skill
The framing is explicit about the year and the analogy: programmers are described as mecha pilots, and the stated position is that hiring should target people who can drive AI agents rather than people who can write code. The core claim follows from it, that using AI to write code is easy while knowing what to build, how to build it and whether the result is right is the scarce ability.
Five things are assessed rather than technical knowledge: issue quality, meaning whether the candidate can write instructions an AI agent can execute directly; review judgment for spotting risk in AI output; AI driving ability; product intuition and systems thinking, expressed as judging whether to build something at all rather than only how; and decision-making under pressure, described as giving clear direction without avoiding it.
Issue quality is the most operational of the five, because it is the one that can be scored from text before the candidate touches a tool. Review judgment is the one with no stated method. The manual measures whether a human can spot risk in machine output, but says nothing about how that human's own judgment is checked, calibrated against a known answer set, or audited after the fact. For a manual whose premise is that machine output needs review, that is the largest methodological gap.
The score sheet is filled in a browser and leaves as a PDF
Scoring happens on the spot in the last five minutes of the session, on an online record form that is filled in within the browser and supports printing or PDF export. Non-technical interviewers are given prepared solution summaries and code diffs so they can run the same process without reading code themselves.
There is no backend in that design. The repository has three files, none of them a service, so scores exist in the browser tab and then in an exported file. There is no candidate identifier field described, no record retention, and no way for two interviewers to score the same candidate independently and compare, which is the mechanism that would catch the rater disagreement that three-tier rubrics tend to produce.
The origin is stated plainly. This is described as an internal interview tool from a healthcare company, open sourced for the reference of peers, MIT licensed in the README footer, with no homepage beyond the GitHub Pages address. The last push to the repository is dated 2026-03-12 and there are no tagged releases, so the published page and the repository copy may differ without a version marker to tell you which is which.
Editorial conclusion
This manual is worth reading as a template if you are redesigning technical interviews, and worth copying only after you rewrite the parts the repository does not show you. The schedule, the role split and the pass thresholds are visible and arguable, but the 14 questions, the three scoring tiers and the six knockout criteria are all inside the HTML, so audit those against your own bias risks before you run them on a real candidate.
Frequently asked questions
What does the ai-native-hiring-guide manual assess?
It assesses five things instead of technical knowledge: issue quality, review judgment for AI output, AI driving ability, product intuition and systems thinking about whether to build something, and decision-making under pressure. Candidates are typed as Builder, Reviewer or dual in the first five minutes.
How is the ai-native-hiring-guide scored?
Five modules at ten points each for a maximum of 50. PASS requires 35 or above with no veto item triggered, HOLD is 30 to 34 and recorded as a backup, and FAIL is below 30 or any veto item. Six hard knockout criteria can trigger a fail regardless of the total.
How do I open the ai-native-hiring-guide locally?
It is a single HTML file with no dependencies, so you can open index.html directly or serve the directory with any HTTP server, for example npx serve . An online copy is published at a GitHub Pages address.
How long is the ai-native-hiring-guide interview?
Sixty minutes, in six segments: five minutes of icebreaking and typing, fifteen for issue writing and solution review, ten for the hands-on AI driving exercise, fifteen for the role-specific module, ten for candidate questions, and five for scoring on the spot.
Can non-technical interviewers use the ai-native-hiring-guide?
Yes, that is the stated purpose of the interviewer support material, which provides ready-made solution summaries and code diffs. Scoring is done in the browser on a form that supports printing and PDF export, and no backend is required.
Who maintains the ai-native-hiring-guide?
It is described as an internal interview tool from a healthcare company, open sourced for the reference of peers, with an MIT license stated in the README footer. The last push to the repository is dated 2026-03-12 and there are no tagged releases.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vorojar-ai-native-hiring-guide)