The missing layer was judgment
I already had the production line: give a design skill and a website prompt to a model, then save the result. It produced an archive of attractive HTML files and one stubborn question: which part was actually good?
A polished screenshot could hide weak prompt coverage. A faithful design system could still produce a bad website. A beautiful desktop layout could quietly become a horizontal scrolling incident at 375 pixels.
So I added the evaluation layer.
What went through the machine
This is a real baseline, not a balanced tournament bracket. The uneven coverage matters later.
- Outputs
- 32
- Models
- 6
- Skills
- 3
- Viewports
- 6
Every single one of the 32 scored outputs has a dedicated record page with full-resolution desktop and mobile review frames, sub-score metrics, review notes, and live HTML source.
Six dimensions, one hundred points
The rubric separates instruction-following from design judgment and browser behavior.
Visual language, layout grammar, typography, components, prohibited patterns, and distinctiveness.
Explicit requirement coverage tied back to the prompt stored in the repository.
Hierarchy, composition, spacing, readability, rhythm, polish, coherence, and intentionality.
Semantics, accessibility basics, asset health, runtime failures, and code weight.
Rendered behavior at 375, 430, 768, 1024, 1440, and 1920 pixels.
The artifact-versus-template test: did this particular website feel deliberately designed?
The first scoreboard
These averages describe this dataset. They do not prove that one model is generally superior.
| Model | Outputs | Overall | Responsive | Design |
|---|---|---|---|---|
| Grok 4.6 | 4 | 93.4 | 7.5 / 10 | 17.9 / 20 |
| GPT-5.6 Sol | 11 | 93.0 | 7.9 / 10 | 17.3 / 20 |
| Gemini 3.7 Flash | 10 | 90.4 | 6.0 / 10 | 17.2 / 20 |
| MiMo V2.5 | 2 | 84.6 | 5.7 / 10 | 14.5 / 20 |
| Sonnet 4.6 Thinking | 3 | 84.0 | 3.2 / 10 | 15.3 / 20 |
| Kimi K2.6 | 2 | 82.4 | 2.6 / 10 | 15.0 / 20 |
Pretty desktop work was common.
The meaningful separation appeared when prompt coverage, negative constraints, and mobile translation were scored separately.
Responsive behavior moved the rankings.
Several convincing split-flap designs remained fixed-width. One Kimi cinema output overflowed until the viewport reached 1920 pixels.
Specific structure beat decoration.
The strongest work turned the subject matter into the interface instead of applying the design skill as a surface treatment.
The design language survived. Sometimes.
Three representative pairs show how each system translated the same artifact from desktop into a narrow viewport.
Want to inspect any specific site's desktop and mobile frames at 1:1 scale? Browse the full visual gallery in the benchmark archive.
Inspect all 32 grades
Filter the baseline, sort by any dimension, and open any record to see its score breakdown, violations, review note, and direct links to its evidence page and live source output.
Consistency is its own result
Stability uses distinct prompt means, so duplicate variants of one prompt do not dominate. Lower deviation is steadier.
| Model | Skill | Prompts | Mean | Deviation |
|---|---|---|---|---|
| GPT-5.6 Sol | Blueprint | 2 | 91.9 | 0.1 |
| Gemini 3.7 Flash | Blueprint | 2 | 90.5 | 0.5 |
| GPT-5.6 Sol | Field Manual | 2 | 95.8 | 0.8 |
| Grok 4.6 | Field Manual | 2 | 96.2 | 0.9 |
| Gemini 3.7 Flash | Field Manual | 4 | 90.1 | 1.0 |
| GPT-5.6 Sol | Solari | 4 | 92.5 | 1.3 |
| Gemini 3.7 Flash | Solari | 2 | 91.9 | 1.6 |
| MiMo V2.5 | Solari | 2 | 84.6 | 2.8 |
| Grok 4.6 | Solari | 2 | 90.7 | 2.9 |
| Kimi K2.6 | Solari | 2 | 82.4 | 3.4 |
What this baseline cannot prove
- The model field is unbalanced. GPT-5.6 Sol has 11 outputs. Kimi and MiMo have two each. The averages are descriptive, not a controlled model ranking.
- One artifact is provisional. Ridgeline Trail Conservancy has no matching source prompt in the repository, so its prompt score is omitted and its total is normalized across available dimensions.
- Skill differentiation is not measured yet. No identical brief appears under more than one skill. The system cannot honestly say how far Blueprint and Field Manual diverge on the same subject.
- This is not a Lighthouse study. Technical quality covers source and browser checks, not Lighthouse Performance, Accessibility, Best Practices, or SEO scores.
- The visual review is a baseline judgment. The next useful measurement is the difference between this grade and the owner's independent grade.