← Lemonteed Home DESIGN SKILL BENCHMARK
← Read Case Study Post BASELINE 001 · FROZEN EVIDENCE

OPERATOR'S LOG / EXPERIMENT 001

Every generation, on the record.

This is the inspectable evidence archive for Design Skill Benchmark Baseline 001: 32 generated websites, six models, three design systems, and the grades behind the case study.

Records
32
Models
6
Skills
3
Frames
64

What this is: a static snapshot of every scored generation, its desktop and mobile review frames, score breakdown, review note, violations, and a sanitized source archive. What it is not: a live evaluator or a model leaderboard beyond this dataset.

BASELINE 001 / DATA VIEWER

The record, without the CSV goggles.

DIMENSION 01 / DESIGN SKILLS

Three design systems, strict rules.

Each model was tested against specific design skill definitions to evaluate style fidelity, layout grammar, and whether negative constraints were respected.

BLUEPRINT 91.3 AVG · 5 OUTPUTS

Blueprint System

Technical and architectural diagrammatic language: high-density grids, technical monospace typography, dimension markings, schematic drawing frames, and blueprint cyan accents.

Key requirements: Spec-sheet hierarchy, title block/schedule, engineering notes, dimension markers.
Prohibitions: No marketing gradients, no rounded cartoon pills, no generic hero banners.
FIELD MANUAL 90.9 AVG · 12 OUTPUTS

Field Manual System

Government, military, and institutional operational handbooks: numbered procedure lists, tactical callouts, classification stamps, warning blocks, and utilitarian olive/gold accents.

Key requirements: Strict section numbering, procedure tables, classification banners, emergency callouts.
Prohibitions: No SaaS glassmorphism, no marketing fluff, no non-functional interactive gimmicks.
SOLARI 89.2 AVG · 15 OUTPUTS

Solari System

Mechanical split-flap displays and transit timetable boards: dark airport/station departure aesthetics, high-contrast amber/yellow glow, and tabular multi-column schedule matrices.

Key requirements: Split-flap typography/character cells, board header status, status codes (ON TIME, DELAYED).
Prohibitions: No light mode, no fluid unconstrained cards that break character grid alignment.

DIMENSION 02 / PROMPT CATALOG

Eleven real-world scenario prompts.

Prompts combined dense domain requirements, data structures, and responsive layouts to test whether models followed instructions or took shortcuts.

Prompt Scenario Design Skill Domain & Complexity Outputs
RELAY — AI Agent Deployment & Operations Field Manual Multi-agent operations handbook, orchestrator routing, telemetry tables 4 runs
Municipal Trail Operations Field Manual Park trail maintenance, incident status report, ranger log 3 runs
Northstar Observatory Field Manual Telescope optics schedule, observation logs, aperture coordinates 3 runs
Municipal League Baseball Solari Civic baseball split-flap scoreboard, standings board, pitch count matrix 6 runs
Terminal 17 — Regional Freight Exchange Solari Intermodal rail/truck timetable board, bay assignments, cargo status 4 runs
Night Shift 91.3 FM Solari Late-night radio program board, broadcast queue, audio track schedule 3 runs
ATLAS — Distributed Intelligence Architecture Blueprint High-density multi-node diagram, memory architecture, execution topology 3 runs
Halloway Instruments HX-4 Field Manual Environmental sensor field manual, calibration chart, sensor readings 2 runs
Aeroline R7 — Endurance Cycle Platform Blueprint Endurance engineering spec sheet, frame stress diagrams, torque tolerances 1 run
Northline Waterworks Blueprint Regional water authority engineering sheet, pipeline pressure monitors 1 run
Analog Cinema Club Solari Repertory theater 35mm schedule, screening room split-flap board 2 runs

DIMENSION 03 / EVALUATED MODELS

Six leading LLMs on the test bench.

Scoreboard of model performance across 32 runs, measured on 6 independent grading dimensions.

Model Outputs Overall Score Responsive (10) Design (20) Key Observation
Grok 4.6 4 93.4 7.5 / 10 17.9 / 20 Highest consistency; tightest adherence to negative constraints and typography.
GPT-5.6 Sol 11 93.0 7.9 / 10 17.3 / 20 Strongest responsive behavior; excellent split-flap and table translations.
Gemini 3.7 Flash 10 90.4 6.0 / 10 17.2 / 20 High design polish and prompt fidelity; occasional mobile viewport overflow.
MiMo V2.5 2 84.6 5.7 / 10 14.5 / 20 Solid visual styling; missed a subset of explicit table requirements.
Sonnet 4.6 Thinking 3 84.0 3.2 / 10 15.3 / 20 Strong desktop layouts, but severe mobile responsive failure (fixed widths).
Kimi K2.6 2 82.4 2.6 / 10 15.0 / 20 Creative layouts, but widespread horizontal scrolling on screens under 1024px.

DIMENSION 04 / EVIDENCE & RUBRIC

100 points, 6 dimensions, 64 review frames.

Every score is backed by full-resolution desktop (1440px) and mobile (375px) WebP screenshots, rubric deductions, and live source code.

01Skill Fidelity25 PTS

Visual language, layout grammar, typography, components, prohibited patterns, and distinctiveness.

02Prompt Fidelity25 PTS

Explicit requirement coverage tied back to the test prompt and domain data points.

03Design Quality20 PTS

Hierarchy, composition, spacing, readability, rhythm, polish, coherence, and intentionality.

04Technical Quality15 PTS

Semantics, accessibility basics, asset health, runtime failures, and code weight.

05Responsive Quality10 PTS

Rendered behavior at 375, 430, 768, 1024, 1440, and 1920 pixels without horizontal overflow.

06Originality5 PTS

The artifact-versus-template test: did this website feel deliberately designed rather than generic?

DIMENSION 05 / OUTPUT RECORDS

All 32 Benchmark Output Records

Click any record card to inspect its full evidence page, desktop & mobile review frames, sub-score breakdown, and archived HTML source.