Skip to content

[ OPERATOR'S LOG / EXPERIMENT 001 ]

I Built a Test Bench for Website Design Skills

32 generated websites. Six models. Three design systems. One attempt to stop saying “this looks pretty good.”

BASELINE 001 RECORDED 31 COMPLETE / 1 PROVISIONAL

The missing layer was judgment

I already had the production line: give a design skill and a website prompt to a model, then save the result. It produced an archive of attractive HTML files and one stubborn question: which part was actually good?

A polished screenshot could hide weak prompt coverage. A faithful design system could still produce a bad website. A beautiful desktop layout could quietly become a horizontal scrolling incident at 375 pixels.

So I added the evaluation layer.

01 / BASELINE

What went through the machine

This is a real baseline, not a balanced tournament bracket. The uneven coverage matters later.

Outputs
32
Models
6
Skills
3
Viewports
6
Design Skill Benchmark dashboard showing model averages and the beginning of the ranked artifact table
The local report keeps every dimension visible. A total is useful; the reason behind it is the actual record.
Live Evidence Archive

Every single one of the 32 scored outputs has a dedicated record page with full-resolution desktop and mobile review frames, sub-score metrics, review notes, and live HTML source.

02 / RUBRIC

Six dimensions, one hundred points

The rubric separates instruction-following from design judgment and browser behavior.

01Skill fidelity25

Visual language, layout grammar, typography, components, prohibited patterns, and distinctiveness.

02Prompt fidelity25

Explicit requirement coverage tied back to the prompt stored in the repository.

03Design quality20

Hierarchy, composition, spacing, readability, rhythm, polish, coherence, and intentionality.

04Technical quality15

Semantics, accessibility basics, asset health, runtime failures, and code weight.

05Responsive quality10

Rendered behavior at 375, 430, 768, 1024, 1440, and 1920 pixels.

06Originality5

The artifact-versus-template test: did this particular website feel deliberately designed?

03 / RESULTS

The first scoreboard

These averages describe this dataset. They do not prove that one model is generally superior.

ModelOutputsOverallResponsiveDesign
Grok 4.6493.47.5 / 1017.9 / 20
GPT-5.6 Sol1193.07.9 / 1017.3 / 20
Gemini 3.7 Flash1090.46.0 / 1017.2 / 20
MiMo V2.5284.65.7 / 1014.5 / 20
Sonnet 4.6 Thinking384.03.2 / 1015.3 / 20
Kimi K2.6282.42.6 / 1015.0 / 20
BLUEPRINT91.35 outputs
FIELD MANUAL90.912 outputs
SOLARI89.215 outputs
FINDING A

Pretty desktop work was common.

The meaningful separation appeared when prompt coverage, negative constraints, and mobile translation were scored separately.

FINDING B

Responsive behavior moved the rankings.

Several convincing split-flap designs remained fixed-width. One Kimi cinema output overflowed until the viewport reached 1920 pixels.

FINDING C

Specific structure beat decoration.

The strongest work turned the subject matter into the interface instead of applying the design skill as a surface treatment.

04 / EVIDENCE

The design language survived. Sometimes.

Three representative pairs show how each system translated the same artifact from desktop into a narrow viewport.

Desktop ATLAS systems architecture website using Blueprint technical drawing language Mobile ATLAS systems architecture website preserving Blueprint hierarchy
Blueprint / ATLAS / GPT-5.6 Sol — the architecture diagram becomes the page structure rather than decorative engineering wallpaper.
Desktop RELAY AI deployment handbook using Field Manual procedural document language Mobile RELAY AI deployment handbook preserving recorded state and procedural hierarchy
Field Manual / RELAY / Grok 4.6 — recorded state, procedural hierarchy, and a strong mobile document produced the highest score in the baseline: 97.
Desktop Municipal League Baseball website using Solari split-flap scoreboard language Mobile Municipal League Baseball website preserving the Solari scoreboard mechanism
Solari / Municipal League Baseball / Gemini 3.7 Flash — a civic scoreboard identity survives the narrow translation without turning into an airport parody.
Contact sheet containing desktop review frames for all 32 benchmark outputs
All 32 desktop review frames. Open the image for the full sheet.
Contact sheet containing mobile review frames for all 32 benchmark outputs
All 32 mobile review frames. This sheet caught what the desktop sheet politely omitted.
Inspect Individual Frames

Want to inspect any specific site's desktop and mobile frames at 1:1 scale? Browse the full visual gallery in the benchmark archive.

05 / FULL RECORD

Inspect all 32 grades

Filter the baseline, sort by any dimension, and open any record to see its score breakdown, violations, review note, and direct links to its evidence page and live source output.

Loading the baseline… Higher scores first
06 / STABILITY

Consistency is its own result

Stability uses distinct prompt means, so duplicate variants of one prompt do not dominate. Lower deviation is steadier.

ModelSkillPromptsMeanDeviation
GPT-5.6 SolBlueprint291.90.1
Gemini 3.7 FlashBlueprint290.50.5
GPT-5.6 SolField Manual295.80.8
Grok 4.6Field Manual296.20.9
Gemini 3.7 FlashField Manual490.11.0
GPT-5.6 SolSolari492.51.3
Gemini 3.7 FlashSolari291.91.6
MiMo V2.5Solari284.62.8
Grok 4.6Solari290.72.9
Kimi K2.6Solari282.43.4
07 / LIMITS

What this baseline cannot prove

  • The model field is unbalanced. GPT-5.6 Sol has 11 outputs. Kimi and MiMo have two each. The averages are descriptive, not a controlled model ranking.
  • One artifact is provisional. Ridgeline Trail Conservancy has no matching source prompt in the repository, so its prompt score is omitted and its total is normalized across available dimensions.
  • Skill differentiation is not measured yet. No identical brief appears under more than one skill. The system cannot honestly say how far Blueprint and Field Manual diverge on the same subject.
  • This is not a Lighthouse study. Technical quality covers source and browser checks, not Lighthouse Performance, Accessibility, Best Practices, or SEO scores.
  • The visual review is a baseline judgment. The next useful measurement is the difference between this grade and the owner's independent grade.
08 / NEXT RUN

Make the experiment harder

  1. 01
    Shared-brief differentiation

    Give one municipal water authority brief to every skill, including a no-skill control.

  2. 02
    Balanced model matrix

    Run the same number of prompts per model and skill before comparing averages.

  3. 03
    Repeated trials

    Generate each cell more than once and report variance instead of treating one output as the model.

  4. 04
    Transferability

    Stress each design language across civic, industrial, cultural, scientific, editorial, consumer, and commerce briefs.

  5. 05
    Human comparison

    Record the owner's grade beside this baseline and inspect where the rubric or reviewer disagrees.