Model or dataset
SeanJ1ang/design-judge-skills avatar
SeanJ1ang/design-judge-skills

The 1,612 entries were scored three days before the winners were known

Evidence-driven Agent Skills for design award research, evaluation, award matching, entry writing, and submission readiness.

712 stars102 forksPythonApache-2.0

At a glance

What is it?
A set of agent skills that covers the design award submission workflow from researching past winners to checking a submission package. What makes it unusual is the evaluation section: the whole of one award's entry pool was scored before the results existed, and the README is careful about what the numbers mean.
Who is it for?
design-judge-skills is unusual in a way worth rewarding: it published the scoring date, the entry count and the fact that the winners were unknown at the time, which is more than most prompt collections offer, and its principles are written to stop a reader treating a score as a verdict. Two things to weigh before you rely on it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 42 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The scoring was frozen before the results existed

The section that answers whether the skills work is placed before the project description, and it starts with a methodology rather than a claim.

Every entry in the 2026 K-Design competition was collected, all 1,612 of them, and all were scored and saved on 2026-07-03, before the winning works were announced. The README states plainly that the winners list was not available when the scoring happened, so the scores cannot be reverse engineered or adjusted afterwards from the final results. That is the single most important line in the file, because a scoring benchmark produced after the winners are known measures nothing.

The reported metric is a top-115 hit count and hit rate for six configurations:

| Model | Top-115 hits | Rate | |---|---:|---:| | GPT-5.5 | 32 | 27.83% | | GPT-5.4 | 33 | 28.70% | | Claude Sonnet 4.6 | 19 | 16.52% | | Design Judge with GPT-5.5 | 41 | 35.65% | | Design Judge with GPT-5.4 | 47 | 40.87% | | Design Judge with Claude Sonnet 4.6 | 44 | 38.26% |

The rates confirm the denominator: every figure is a count out of 115, which is the number of winners. So the question being asked is how many actual winners a configuration placed inside its own top 115, out of the full 1,612.

The skills layer moves the number, the model swap moves it too

Read the table by row pair and two effects separate.

Adding the skills to the same underlying model lifts the hit count in all three cases: 32 to 41 for one GPT-5.5 configuration, 33 to 47 for the other, and 19 to 44 for Claude Sonnet 4.6. The last of those is the largest single jump in the table, and it is also the weakest baseline.

Swapping the model underneath the skills moves the result too, from 41 to 47 between the two GPT configurations, so the skills are not the only variable. The three skills-augmented rows all land between 38 and 41 percent, which is a narrower band than the baselines, which run from 16.5 to 28.7 percent.

What the table does not say is what a top-115 ranking is, whether the skills output a ranked list at all, or how the ranking was produced from a score. Those are the questions a reader should carry away, and a document link on coverage, privacy and limits is given for exactly that kind of detail. The honest summary is that the skills roughly halve the distance between a bare model and a good ranking on this one competition, and that a single competition is one competition.

Five of six user-facing skills are Beta

The skill index carries a status column, and the distribution is worth reading before you plan around any of them.

One skill is Stable: the search skill, which retrieves and verifies same-category winning works from official sources. Five are Beta, meaning they have passed the examples and the automated tests but may still have edge case problems: the pipeline skill that picks a route, the evaluation skill, the award matching skill, the entry text preparation skill and the submission check. The seventh entry is the shared support package, marked as a support item that is not triggered on its own.

The definitions are given in the file rather than assumed, which is better than most status vocabularies, and there is a note that the badge count covers six user-facing workflows and excludes the shared package. So the honest headline is one stable skill and five beta ones.

The three READMEs of the repository itself, in Chinese, English and Japanese, and a per-skill pair of a Chinese and an English README, are the same idea applied to documentation.

Two skills cannot be installed without a third one

Each skills directory in the repository is an installable unit, and one of them is a dependency rather than a tool. The shared package provides the category taxonomy and the official source rules for the search and matching skills, and the README is explicit: a full install includes it automatically, but installing the search skill or the matching skill on its own means installing the shared package too.

The documented route is a skills command line, which needs Node.js 18 or newer and no global install of its own:

bash
npx skills add SeanJ1ang/design-judge-skills --list

Then a full global install to a named agent, with a wildcard skill selector, an agent name, a confirmation flag and a copy flag:

bash
npx skills add SeanJ1ang/design-judge-skills --global --agent codex --skill '*' --yes --copy

The same command installs a single skill into the current project instead of globally, and repeating the skill flag installs the search skill together with the shared package. Passing the all flag covers every agent the command line supports. Listing and updating are separate commands, with an update command that takes a global flag and a confirmation.

For agents that only understand a skill file, the fallback is spelled out as four steps: clone the repository to a stable path, copy or link the whole skill directory, keep the skill file along with its agents, references, scripts, examples and tests directories, and keep the shared package whenever the search or matching skill is in use.

The pipeline skill chooses a route, it does not force one

The workflow diagram has one entry point and a chain, and the accompanying sentence matters more than the diagram.

Design materials go to the pipeline skill when the route is uncertain or the full process is wanted. From there the diagram branches to searching for comparable winners, evaluating the design and its presentation, and choosing an award, track and category, and the last of those leads to preparing the entry information and then to the submission check.

The note under it says the modules can be used independently, and that a complete submission usually starts at evaluation or matching, moves to information preparation once the target award is settled, and ends with the submission check. The quick start then adds the part that makes the design coherent: if you do not know which skill you need, use the pipeline one, because it selects only the necessary modules and does not force the full pipeline.

Six ready made prompts are given, one per skill, each naming the skill with a dollar sign prefix and stating the task. The matching example, for instance, asks the agent to compare five awards against a project, and the submission check example asks for a go, conditional go or no-go conclusion against that year's official rules.

Scores are for deciding, not for estimating a win

Three of the five design principles exist to stop the skills being read as an oracle.

First, first party sources win: award rules and cases come from official pages, and search summaries and third party pages are only for finding leads. Second, facts are separated from inferences and from items still needing user confirmation. Third, and the one that matters most for a reader with an entry to submit, scoring stays transparent: fit and evaluation scores support a decision and do not represent a probability of winning.

The same limit appears in the data section. The evaluation module carries 22,125 aggregated observation records of winning or shortlisted works from four awards, iF, iF Student, Red Dot and IDEA, and the file says these records provide descriptive background only, do not change the core scoring, and cannot be used to estimate a winning probability.

A fourth principle is about honesty in a different direction: the skills align to published judging criteria but do not simulate undisclosed jury preferences or internal processes. Combined with the requirement that deadlines, fees, eligibility, categories and submission specifications be re-verified against official pages at run time, that is a coherent position for a project whose entire pitch is evidence.

Eleven awards are in scope, and dates are re-checked at run time

The configured coverage is a list of eleven award programmes: iF DESIGN AWARD, iF DESIGN STUDENT AWARD, Red Dot Product Design, Red Dot Design Concept, IDEA, DIA, K-Design, GOOD DESIGN AWARD Japan, Core77, James Dyson and EPDA.

The list is about routing, not about content. The rules are explicit that whenever deadlines, fees, eligibility, categories or submission specifications are involved, the skill requires the official page to be checked again at run time. A skills repository cannot keep eleven programmes' terms current in prose, so the design is to fetch rather than to remember, and to say so.

The structure supports that. Every skill directory holds its own references directory, and the shared package holds a category taxonomy and a source registry, which is where a mapping from the user's project to the right award track would live.

One detail in the layout is easy to miss: every user-facing skill carries a file named after OpenAI in its agents directory, alongside its own references, scripts, examples and tests. There is an evals directory at the repository root as well, which is the natural place for the scoring described at the top of the file, and the development conventions section ends mid sentence.

Editorial conclusion

design-judge-skills is unusual in a way worth rewarding: it published the scoring date, the entry count and the fact that the winners were unknown at the time, which is more than most prompt collections offer, and its principles are written to stop a reader treating a score as a verdict. Two things to weigh before you rely on it. The evaluation is one competition with 1,612 entries and 115 winners, so the effect is real but the sample is one, and the reading is a top-115 hit rate rather than a claim about your own work. And five of the six user-facing skills are Beta by the project's own definition, so treat the search skill as the settled one and the rest as drafts. On the substance, the rules that matter most are the ones about first party sources and re-verifying dates and fees at run time, and both are reasons to keep a human in the loop before you submit anything.

Frequently asked questions

What is design-judge-skills?

A set of agent skills covering the design award submission workflow: retrieving and verifying comparable winning entries from official sources, evaluating a design and its presentation with an evidence rubric, matching a project to awards, tracks and categories, preparing entry text with source constraints, and checking a submission package against that year's official rules.

How were the Design Judge skills evaluated?

All 1,612 entries in the 2026 K-Design competition were collected and scored, with the scores saved on 2026-07-03, before the winning works were announced. The README notes the winners list was unavailable at that point, so the scores could not be revised from the results. The table reports how many winners each configuration placed in its own top 115.

How do I install design-judge-skills?

With the skills command line and Node.js 18 or newer, with no global install of the command line itself. A listing command shows what the repository offers, and an add command with a global flag, an agent name, a wildcard skill selector, a confirmation flag and a copy flag installs everything for one agent. An all flag covers every supported agent, and separate list and update commands manage what is installed.

Which of the design-judge skills are stable?

One, the search skill, which retrieves and verifies comparable winning entries from official sources. The pipeline, evaluation, award matching, entry preparation and submission check skills are all marked Beta, meaning they passed the examples and automated tests but may still have edge case issues. The shared package is a support item and is not counted among the six user-facing workflows.

Do I need the design-judge-shared package?

Yes, if you install the search skill or the award matching skill on their own. It provides the shared category taxonomy and official source rules, it is included automatically in a full install, and the install instructions show adding it in the same command with a second skill flag.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. SeanJ1ang/design-judge-skills on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/seanj1ang-design-judge-skills.svg)](https://hysenlabs.com/projects/seanj1ang-design-judge-skills)