Back to Blog
Column

Which model do I trust with eight hundred websites?

FastMetal

Hawaii Tech Week · 3 September 2026

I own eight hundred dropped domains and about four hours a week. This is the benchmark I built to decide which language model gets to write on all of them, and the two runs it took to stop measuring my own instrument.

Eight hundred domains, four hours a week

Where the problem comes from, and why my credentials are not an answer to it.

§1the portfolio

Over eight hundred domains

No websites on any of them.

A hand-picked sample of the portfolio, roughly a quarter of it Japanese IDN. These are dropped domains with real search value and nothing on them.

  • waikiki.jp
  • paris.jp
  • berlin.jp
  • amsterdam.jp
  • athens.jp
  • copenhagen.jp
  • brussels.jp
  • budapest.jp
  • geneva.jp
  • montreal.jp
  • moscow.jp
  • denver.jp
  • detroit.jp
  • houston.jp
  • sandiego.jp
  • sanfrancisco.jp
  • manhattan.jp
  • monterey.jp
  • napa.jp
  • cancun.jp
  • caribbean.jp
  • kualalumpur.jp
  • kathmandu.jp
  • johannesburg.jp
  • beijing.jp
  • hong-kong.jp
  • liverpool.jp
  • beverlyhills.jp
  • neworleans.jp
  • washingtondc.jp
  • netherlands.jp
  • deutschland.jp
  • poland.jp
  • sweden.jp
  • switzerland.jp
  • austria.jp
  • greece.jp
  • bulgaria.jp
  • cyprus.jp
  • serbia.jp
  • ukraine.jp
  • egypt.jp
  • colombia.jp
  • uruguay.jp
  • saudiarabia.jp
  • bahamas.jp
  • seychelles.jp
  • tanzania.jp
  • new-zealand.jp
  • sri-lanka.jp
  • alabama.jp
  • alaska.jp
  • colorado.jp
  • illinois.jp
  • kentucky.jp
  • ohio.jp
  • oregon.jp
  • virginia.jp
  • cheese.jp
  • cigars.jp
  • cola.jp
  • dinner.jp
  • buffet.jp
  • bowling.jp
  • surfing.jp
  • skate.jp
  • jetski.jp
  • golfclub.jp
  • winery.jp
  • tapas.jp
  • teppanyaki.jp
  • restaurants.jp
  • supermarket.jp
  • ice-cream.jp
  • cookbook.jp
  • tuna.jp
  • beaujolais.jp
  • olympics.jp
  • presentation.jp
  • online.jp
  • wifi.jp
  • usb.jp
  • electronics.jp
  • electricity.jp
  • investing.jp
  • investment.jp
  • securities.jp
  • hedgefund.jp
  • venturecapital.jp
  • wallstreet.jp
  • dollar.jp
  • funding.jp
  • lease.jp
  • advertising.jp
  • iptv.jp
  • ultrabook.jp
  • allergy.jp
  • diabetes.jp
  • dermatology.jp
  • cleanenergy.jp
  • solarpanel.jp
  • ecoenergy.jp
  • iconiq.jp
  • iphone-repair.jp
  • カフェ.jp
  • デザイン.jp
  • システム.jp
  • ビーチ.jp
  • ツーリズム.jp
  • ナビゲーション.jp
  • スマートフォン.jp
  • クレジットカード.jp
  • ソーラー.jp
  • ウィスキー.jp
  • イタリアワイン.jp
  • カリフォルニアワイン.jp
  • アロハ.jp
  • オアフ.jp
  • 中華料理.jp
  • 日本料理.jp
  • 野菜.jp
  • 九州.jp
  • 博多.jp
  • 富士.jp
  • 丸の内.jp
  • 香港.jp
  • 交通.jp
  • 眼科.jp
  • 元気.jp
  • 代替エネルギー.jp
  • 新エネルギー.jp
  • カーボン市場.jp

§2hawaii tech week · 3 september 2026

Which model do I trust with eight hundred websites?

The whole talk is one question, and it is a purchasing question rather than a technology one.

§3qualifications

You want to know my references?

  • Pioneered the Japanese domain market from 1999
  • Board, Japan Domain Name Business Association
  • Vice-board, Japan Internet Providers Association
  • ICANN member
  • Millions of dollars in Japanese-language paid media
  • Companies and offices in Honolulu, Tokyo and Hanoi
  • EO board member
  • Twenty Ironmans
  • Airplane and helicopter pilot
  • Two Pacific crossings as captain

Sold Solis, Japan's third largest domain registrar, to Global Media Online in 2005.

Tokyo listed. Multi-billion dollar. Owns Japan's largest registrar.

The list is all true, and it escalates into the irrelevant on purpose. Two Pacific crossings tell you nothing about which language model to buy. That is the point of putting them on the same slide as the registrar sale.

§4what I actually run

Real operating companies, not slideware. Honolulu, Tokyo, Hanoi. Each mark links to its site.

§5

None of this tells you which model to use.

Credentials are hearsay with a letterhead. This talk argues against them, including my own.

§6built with claude

Two sites

Both are live. Both took many sessions with my eye on every screen.

§7

I'm not a designer.

Six months ago I could not have built either of those. That is the point: the machine closed that gap.

Can it close it eight hundred times, when I am not looking?

That is the whole difficulty. Those two sites took many sessions with my attention on every screen. I cannot give eight hundred sites that kind of attention, and I am not going to pretend otherwise.

§8the point of view

Eight hundred dropped domains.

Real search value. No websites. The play is not flipping them, it is putting real content on them so the value compounds.

Eight hundred domains, and about four hours a week to spend on them. That constraint is what makes this a measurement problem instead of a craft problem.

§9disclosure

I own the gateway.

fastmetal.ai

Every model in this talk runs through it, and I resell all of them. I am also its largest customer. Four of my companies run on it.

So do not take my word for anything today. Take the numbers.

FastMetal home pagefastmetal.aiopen the live site

enabled on my key

  • anthropic-claude-fable-5
  • anthropic-claude-fable-5-1
  • anthropic-claude-haiku-4-5
  • anthropic-claude-opus-4-6
  • anthropic-claude-opus-4-7
  • anthropic-claude-opus-4-8
  • anthropic-claude-opus-5
  • anthropic-claude-sonnet-4-6
  • anthropic-claude-sonnet-5
  • bytedance-seedream-4.5
  • deepseek-v4-flash
  • deepseek-v4-flash-0731
  • deepseek-v4-pro
  • gemini-3.5-flash
  • gemini-3.7-flash
  • gemini-flash-lite-free
  • glm-4.7
  • glm-4.7-flash
  • glm-5
  • glm-5.1
  • glm-5.2
  • glm-5.3
  • glm-5.3-flash
  • google-nano-banana-2
  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra
  • gpt-oss-120b
  • grok-4.5
  • grok-4.6
  • inkling
  • japan-gemma-4-31b
  • japan-kimi-k2.6
  • japan-kimi-k2.7-code
  • japan-qwen3.6-35b
  • kimi-k2.6
  • kimi-k3
  • llm-jp-3.1-8x13b-instruct4
  • mimo-v2.5
  • mimo-v2.5-pro
  • minimax-m2.7
  • minimax-m3
  • mistral-voxtral-mini-3b-2507
  • muse-glimmer-30b
  • muse-spark-1.2
  • qwen3.6-27b
  • qwen3.7-max
  • qwen3.8-27b
  • qwen3.8-2.4t-a95b
  • qwen3.8-max
  • solar-pro4
  • z-image-turbo

The roster is what is enabled on my key, not the whole catalogue — the service advertises far more. Saying this early and at full strength is cheaper than having it surface later.

Part one

How the experiment was set up

Almost everybody picks a model from hearsay. Someone said it was good, so it is the one they use. You cannot evaluate until you can state the job.

§10key point one

You cannot evaluate until you can state the job

Almost everybody picks a model from hearsay. Someone said it was good, so it is the one they use.

This section buys the rest of the talk. If the setup is not believable, no result that comes out of it is worth reading. So it goes first, in full, before a single number.

§11the workload

A five-page site that has to earn its ranking

Real topical content, on a domain that has never had any. Design that holds up next to a working business site. On-page SEO that earns the ranking. Writing a person would actually read.

Five pages that make the domain worth more than it was.

Exhibit A below is the harness itself: what is fixed, what the model decides, and what gets written to disk. Everything tinted in its centre lane is the model's work. Everything else is byte-identical for all five contestants.

Exhibit A

contents ↑

How this actually works

There is no orchestrating model. The harness is a loop making direct HTTP calls to a gateway — no agent, no tool use, no retries on content. A language model appears once more at the end, as the judge, and it shares no family with any contestant. Everything tinted in the centre lane is the model's work; everything else is byte-identical for all five.

One domain, one model
FIXED — identical for every model, domain and stageTHE MODEL — the only thing that changesDISK — written verbatimthe domainwaikiki.jp · build language enfour frozen prompts_shared.md prepended to every calls1 · s2 · s3 · s4 + JSON schemassha recorded in the run manifestthe shared floorbase.css — reset, type & space scalepage.html — required page shapepinned samplingtemperature 0.7 · top_p 1.0max_tokens 128,000one config stringswap the model, change nothing elseapplies to every stageS1 · concept & keyword architecturereads the bare domain nameemits concept, clusters, 5-page sitemapconcept + sitemap, as textS2 · design directionpalette, type, per-page sectionsimage manifest: prompt, placement, altdesign + manifestS3 · buildcontent, design and on-page SEO at onceemits complete HTML for every pagecraft brief injected into the writingthe finished siteS4 · self-auditreports its own violationsfixes nothingpromptsz-image-turbo — fixedsame model, same settings, every candidate1024² PNG · $0.011 · ≤1 per pageverbatimthe cell on diskindex.html + up to 4 pagesimages/*.pngbase.css (copied, unmodified)_artifacts/s1–s4.json_artifacts/meta.json — tokens, cost, TTFTWHAT ONE CALL CONTAINS — every stage, every modeltwo messages. no history.system → _shared.md, byte-identical every timeuser → this stage's prompt + the prior stages' JSON, pasted in as textthere are NO assistant turns — the model is never shown its own repliesso S2 is a brand-new sessionsame model, no memory of S1 — as far as it knows this is the firstthing it has ever seen. the only continuity is the JSON we hand it.nothing is remembered; everything is re-supplied.the harness does exactly four thingssubstitute variables · call the API · write bytes · fetch the images the model asked forit never edits model output — malformed HTML fails a gate, it is not repaired
The boundary is the point. Four calls, each consuming the last one's structured output. The harness substitutes variables into frozen prompts, calls the API, writes what comes back to disk unaltered, and fetches the images the model asked for from a fixed image model. It does not repair malformed HTML, fill missing fields, or retry on a bad answer — a broken page fails a gate, and that failure is the measurement. Transport failures (timeouts, 429s) are retried, twice, and every retry is logged.
Forty cells, one named model
GRADING RUNS CHEAPEST FIRST — nothing paid grades something that failed to parsethe matrix5 models × 8 domains= 40 sites+6 replicate cells for noise40 sitestier 1 — hard gatesfree · deterministic · msbuild / spec / warnpunycode, word bands, one h1survivorstier 3 — the judgeFable · blind · shuffledR2 from screenshots, not sourceevidence required per scorescoresgridsmodels × domains per rubricweighted overall, never averagedW1 / W2 / W3 side by sidethe two plotsevery siteevery model + whiskerPareto frontiertier 2 — SEO agentscut for this run — scheduleR4 falls back to the judgecalibration probes2 hand-written referencesknown-bad stub · known-good pagemeasured: stub 2, pro 8pre-registered decision rulewritten and committed before any data existedhighest W1 median among models passing build ≥ 7/8the choice is arithmetic, not judgement after the factwhat this does NOT controlthe prompt chain is itself a harness — results hold for this rigone run per cell, except the six replicatesone judge, mitigated by calibration and the room vote, not eliminatedtwo permitted normalisations — outer fence, trailing bytes — both loggedprices are this gateway key's rates, not a public price sheetonly two Japanese domains: judge language bias cannot be separated from difficultywhat is held constantthe four prompts, byte for byte, including the craft briefbase.css and the page contracttemperature 0.7, top_p 1.0, max_tokens 128,000the image model, its settings and the one-per-page capthe judge, its rubrics and the deterministic blind shufflethe domain set and which four are shown — committed before the run
Ordering is the methodology lesson. Free deterministic checks run before anything paid, so no money grades a site that failed to parse. Two hand-written reference pages of known quality sit in a blind pool to prove the judge's scale is not compressed — measured at 2 for the deliberate stub and 8 for the professional page. The decision rule was written and committed before any results existed, so the model chosen on stage is arithmetic rather than a judgement made after seeing the data.

Source: talk/methodology.py, generated from harness/config.json.

§12the prompts

Five prompts, frozen, identical for every model

An identical preamble, then concept, design, build and self-audit. Frozen once the first production run started, and hashed into every run's manifest.

In the room this slide was a wall of unreadable type, and that was the point — the audience reads volume, not content. On a page you can do better than volume, so Exhibit B carries all five in full.

Exhibit B

contents ↑

Five prompts, frozen, identical for every model

The complete text of every prompt in the chain, exactly as sent. Their SHA is recorded in each run's manifest, so what is printed here is checkable against what actually ran.

_sharedthe identical preamble

Prepended verbatim to every call, for every model. It states the job, the five mechanical constraints that fail a build outright, and the sentence the whole talk turns on: assume nobody will inspect your output before it goes live.

# Shared preamble — prepended verbatim to every stage, for every model

**FROZEN once the first production run starts. Identical for all contestants.**

---

You are building a small, genuinely useful website for a domain that is currently for
sale. The site's job is to make the domain worth more: real topical content that can rank
in search and be cited by answer engines. It is not a parked page, not a placeholder, and
not a template fill.

The domain owner has roughly 800 such domains and will run this same process on all of
them without reviewing each result. Assume nobody will inspect your output before it goes
live. Build accordingly.

## Hard constraints

These are checked mechanically. Violating any of them fails the build outright.

1. **At most 5 pages.** Fewer is fine if fewer is right.
2. **At most 1 image per page.** Five images total, maximum.
3. **Word and character bands** — every page must fall inside the band for its declared
   page type. Under the floor is thin content; over the ceiling is padding. Both fail.

   | Page type | EN words | JA characters |
   |---|---|---|
   | homepage | 500 – 1,000 | 1,000 – 2,000 |
   | service | 800 – 1,600 | 1,600 – 3,200 |
   | landing | 600 – 1,200 | 1,200 – 2,400 |
   | category | 400 – 800 | 800 – 1,600 |
   | about | 400 – 800 | 800 – 1,600 |
   | faq | 800 – 1,600 | 1,600 – 3,200 |
   | blog | 1,500 – 3,000 | 3,000 – 6,000 |

4. **Every page must be uniquely valuable.** Not the same page with the nouns swapped.
   Google's doorway-page algorithm exists and this owner has 800 domains — near-duplicate
   pages get the whole portfolio deindexed.
5. **The shared floor is fixed.** `base.css` is linked, never modified, never inlined.

## Output

Return **only** valid JSON matching the schema given for your stage. No prose outside the
JSON, no markdown fences, no commentary.

s1-conceptconcept and keywords

Input is the bare domain and the build language. No brief, no market research. Deriving what the site should be is the job, and being honest about an ambiguous name then committing anyway is what scores.

# S1 · Concept & keyword architecture

## Input
    domain:         {{DOMAIN}}
    build_language: {{BUILD_LANGUAGE}}     # en | ja — fixed, applies to the whole site

That is all you get. No brief, no market research, no instructions about what the site
should be about. Deriving that is the job.

`build_language` is settled and not yours to change — plan the sitemap, slugs and working
titles in it from the start. You are still asked below what language you *would* have
chosen; say so plainly even when it differs, and then plan in the language given.

## What to do

**1. Read the domain.** What does this name imply? What would someone typing it expect?
What would a buyer of this domain want to own? Be honest when a domain is ambiguous or
carries little meaning — say so, and then make a defensible choice anyway.

**2. Consider the TLD.** It is part of the name and it carries information. Reason about
what it implies for audience, market and language, and state your reasoning. If it points
somewhere the rest of the name does not, say what you would do about it. Name the language
you would have picked and why — disagreeing with `build_language` is a legitimate answer
and is scored on the quality of the argument, not on agreement.

**3. Design the site.** Concept, audience, and a sitemap of at most five pages. Every page
needs a distinct job and a declared page type from the table in the preamble. Do not pad
to five.

**4. Keyword architecture.** Clusters, not a keyword list. For each cluster: a primary
term, secondary and semantic variations, the search intent behind it, and which page owns
it. Every page in the sitemap must own exactly one cluster.

## What is being judged

Whether you read the domain correctly and designed a site worth building — coherence,
commercial realism, keyword and intent coverage, sitemap logic, and the quality of your
reasoning about the TLD. Everything downstream is built on this, so a weak concept costs
you four times over.

s2-designdesign direction

Palette, typography, per-page section plan and the image manifest. Note the instruction not to reach for the house style — cream, terracotta, serif display — written into the prompt because every model had already converged on it.

# S2 · Design direction & image manifest

## Input
    domain:         {{DOMAIN}}
    build_language: {{BUILD_LANGUAGE}}     # en | ja — final, not a suggestion
    s1:             {{S1_JSON}}

`build_language` is the language the site will be written in. It may differ from what you
recommended in S1; it is fixed now and applies to all page copy, headings, alt text and
metadata.

**If it differs from your S1 recommendation, commit to it fully.** Page slugs, working
titles and section names all follow `build_language` — an English site with romanised
Japanese slugs is worse than either language done properly. Rename anything from S1 that
no longer fits and note the corrected slugs in your page layouts.

## The shared floor

Every contestant links the same `base.css`. It provides a reset, a fluid type scale
(`--step--1` … `--step-5`), a spacing scale (`--space-3xs` … `--space-3xl`), and layout
primitives: `.container`, `.stack` / `.stack-m` / `.stack-l`, `.prose`, `.section`,
`.grid`, `.cluster`, `.switcher`, `.center`, `.visually-hidden`, `.skip-link`.

It sets **no palette, no typeface, no section layout, and no page structure.** Those are
yours, and they are the thing being judged. The fallback palette is pure black on pure
white with no accent — if you skip it, the site looks exactly as unconsidered as it is.

## What to do

**1. Design rationale.** Two or three sentences: what this site should feel like and why
that suits this subject and audience. Not adjectives — an argument.

**2. Palette.** Concrete hex values for all seven `--c-*` tokens. Must pass WCAG AA for
body text on background.

**Do not reach for the house style.** Language models converge hard on one look: a warm
cream background around `#FBF7F0`, a serif display face, and a terracotta or amber accent.
Measured across five models from four different companies on this exact task, every one
produced that palette and every one chose the same typeface. It is the most recognisable
signature of a generated website, and a buyer who has seen two of them recognises the
third.

Derive the palette and the type from **this** subject and **this** audience. If your first
instinct is cream-and-terracotta with a serif display, that is your first instinct, not
your best one — justify it against the subject or choose again. A B2B industrial site, a
health reference and a destination guide should not arrive at the same colours.

**3. Typography.** Heading and body stacks. Google Fonts is permitted via a single
stylesheet link; if you use it, name the exact families and weights. System stacks are a
legitimate choice, not a cop-out — Japanese builds especially.

**4. Section plan per page.** For each page in the sitemap, the ordered sections and what
each is for. This is where design variance actually lives: what a page is made of and in
what order.

**5. Image manifest.** **At most one image per page, five total.** Fewer is a valid
choice. Every image is generated by the same fixed image model at the same settings for
every contestant, so image *quality* is a constant — what is judged is your art
direction: subject choice, prompt quality, placement, and alt text.

**Every image comes back as a 1024×1024 PNG.** The image model ignores size and aspect
requests, so design for square and crop with CSS if you want something else. A layout that
assumes a wide hero will break.

Alt text is written in `build_language` and does real SEO work. Do not describe the
image; say what it contributes.

## What is being judged

The rendered result, scored from desktop and mobile screenshots, plus art direction. Your
rationale here is read as evidence of intent.

s3-buildfive pages of html

Content, design and on-page SEO in one response, as complete final HTML. There is no later pass to fix anything in. This is the stage where every budget failure happened.

# S3 · Build

## Input
    domain:         {{DOMAIN}}
    build_language: {{BUILD_LANGUAGE}}
    s1:             {{S1_JSON}}
    s2:             {{S2_JSON}}

## What to do

Emit the **complete, final HTML for every page** in the sitemap. Content, design and
on-page SEO all at once — there is no later pass to fix things in. Write it correctly the
first time.

Each page conforms to the contract in `page.html`:

- `<html lang>` matches `build_language`
- `<title>` 50–60 characters, unique across the site, primary keyword present
- `<meta name="description">` 150–160 characters, written to earn a click
- self-referencing absolute `<link rel="canonical">`
- Open Graph: `og:title`, `og:description`, `og:type`, `og:url`, `og:image`
- exactly one JSON-LD block; the `@type` is your choice and should suit the page
- `<link rel="stylesheet" href="base.css">` — never modified, never inlined
- one `<style>` block implementing your S2 palette and typography; the seven `--c-*`
  tokens must all be defined
- `<a class="skip-link">`, `<header>`, `<main id="main">`, `<footer>`
- exactly one `<h1>` inside `<main>`; heading levels never skipped
- images referenced at `images/<filename>` exactly as named in your S2 manifest

## On-page craft — how to write for search

Follow these while writing, not afterwards.

**Structure carries intent.** Headings are a map of the page, not decoration. Each H2
answers a distinct question a searcher actually has. If a heading could sit on any page
about anything, it is doing no work.

**Answer first.** Open every section with a direct, self-contained answer in one or two
sentences, then expand. A passage lifted out of context must still make sense and still
be attributable — that is what makes it quotable by an answer engine.

**Keywords are placed, not sprinkled.** Primary term in the title, the H1, the first
hundred words, and naturally thereafter. Semantic and related variations throughout.
Density in the 1–3% range. Anything that reads as stuffing has failed twice: once with the
reader, once with the ranking.

**Specific beats true.** "Athens has a rich history" is true and worthless. Numbers,
names, dates, prices, comparisons, trade-offs, and things only someone who knows the
subject would think to mention. Generic truth is the signature of generated content, and
it is what a reader detects before they can name it.

**Show experience and be checkable.** Concrete detail, honest limitations, and outbound
links to genuinely authoritative sources where a claim needs backing. Never invent a
statistic, a citation, a review, or a credential. On health, legal, safety or financial
subjects, be conservative and say what you do not know — a confident wrong answer is a
liability, not a style problem.

**Scannable and readable.** Short paragraphs. Lists and tables where the content is
genuinely a list or a table, never to fill space. Plain sentences.

**Link with intent.** Body-copy links to other pages on this site using descriptive anchor
text — never "click here", never the bare page title every time. Every page reachable from
the navigation. No orphans.

**Every page must earn its own existence.** If two pages could be swapped by changing a
few nouns, you have built a doorway network and this site is worthless to its owner.

**Stay inside the word band for the declared page type.** Under the floor is thin; over
the ceiling is padding. Both fail mechanically, so count.

## What is being judged

The rendered pages (design), the prose (is this worth reading, is it specific, does it
read as written rather than generated), and the on-page SEO. Reliability counts as much as
quality: a page that fails to parse scores nothing at all.

s4-selfauditgrade your own work

Audit what you just built and report the violations honestly. Fix nothing. What is judged is whether the self-report matches what the deterministic gates find — a model that reports all clear while failing three gates cannot be run unattended at any price.

# S4 · Self-audit and site artifacts

## Input
    domain:         {{DOMAIN}}
    build_language: {{BUILD_LANGUAGE}}
    s1:             {{S1_JSON}}
    s2:             {{S2_JSON}}
    s3:             {{S3_JSON}}

## What to do

**1. Audit your own work.** Go back over what you just built and check it against the hard
constraints and the page contract. Report what you find — honestly.

Nobody is going to review these sites before they go live. The owner is running this
process 800 times unattended. Your self-report is the only signal available that something
went wrong, so its value depends entirely on it being accurate.

Check at minimum:

- page count ≤ 5, and images ≤ 1 per page, ≤ 5 total
- every page inside the word or character band for its declared page type — **count, do
  not estimate**
- exactly one `<h1>` per page, no skipped heading levels
- title lengths 50–60, meta descriptions 150–160
- all seven `--c-*` tokens defined; `base.css` linked and unmodified
- every internal link resolves to a page that exists
- every referenced image is in your S2 manifest
- no placeholder text, lorem, TODO, or unfilled token anywhere
- language of all copy, headings, alt text and metadata matches `build_language`
- no two pages that differ only in their nouns

For each violation: what it is, which page, and how bad. **If you find nothing, say so
explicitly** — but only if you actually checked. Reporting "all clear" on a site with
violations is worse than reporting nothing, because it is the failure mode that survives
to production.

**2. Site-level artifacts.** `sitemap.xml` and `robots.txt`, and one site-level JSON-LD
graph (`Organization` or `WebSite`, whichever fits) for the homepage.

## What is being judged

Whether your self-report matches what the deterministic gates actually find. A model that
correctly flags its own violations can be run unattended. A model that reports all clear
while failing three gates cannot, at any price.

**Do not fix anything.** Report only. S3's output is final.

Source: harness/prompts/*.md, frozen at the first production run.

§13from the shared preamble

Assume nobody will inspect your output before it goes live. Build accordingly.

Prepended verbatim to all four stages, for every model.

That is the whole talk in one sentence, and every contestant got it in the first line of every call.

§14four stateless calls

S1 to S2 to S3 to S4

S1Concept and keywords. Gets the domain. That is all.
S2Design direction and image manifest.
S3Build. Five pages of final HTML at once.
S4Self-audit. Grade your own work, honestly.

SEO is not a stage. It is folded into S3 and graded in tier two.

It is not design, then SEO, then writing. Everything lands in one response at S3, and there is no later pass to fix anything in. That is exactly why every budget failure in both runs happened at S3.

Exhibit C

contents ↑

What actually happens between S1 and S4

Every call is exactly two messages — a system and a user — with no assistant turns and no history. Four independent sessions; the only continuity is the JSON handed forward as text.

  1. 1Call 1 (S1)

    Fresh, stateless. System = _shared.md. User = s1-concept.md with waikiki.jp and en substituted in, plus the S1 JSON schema appended. The model returns the concept and sitemap. We write it to _artifacts/s1.json.

  2. 2Call 2 (S2)

    A completely new, clean call to the same model. The model has no memory of call 1 — as far as it's concerned this is the first thing it's ever seen. System = _shared.md again. User = s2-design.md, with S1's entire JSON pasted in as text where {{S1_JSON}} sits, plus the S2 schema.

  3. 3Call 3 (S3)

    Same again — S1's JSON and S2's JSON pasted in as text.

  4. 4Call 4 (S4)

    S1, S2 and S3 all pasted in.

So it is four independent sessions, and the only continuity is the JSON we hand it. Nothing is remembered; everything is re-supplied.

Why stateless rather than a conversation

Source: talk/stage_chain.py, from talk/stage-chain-explainer.md.

§15sample size

Eight domains, not one

Four models on one domain is n=1. That is an anecdote, and this talk is anti-anecdote.

Eight buys a pass rate and a variance band.

Stratified across niches so topical difficulty varies, and pre-registered before any run: allergy.jp, cheese.jp, iconiq.jp, iphone-repair.jp, solarpanel.jp, waikiki.jp, 中華料理.jp, 田中.jp.

§16the lineup

Five models, four vendors

modelvendorwhy it is here
gpt-5.6-solOpenAItheir flagship
kimi-k3Moonshottheir flagship
glm-5.2Zhiputheir flagship
glm-5.3-flashZhiputhe cheap one, same vendor
deepseek-v4-flash-0731DeepSeekdeliberately their cheapest

Each vendor’s flagship, except DeepSeek, where I deliberately took the cheapest thing on the gateway. I wanted to know whether the floor was already good enough.

No prices here. Cost is the reveal after the blind vote, and showing it now would state the conclusion before the method has been explained. Free tiers were rejected on reliability, not price: a model that fails one build in five is infinitely expensive at eight hundred sites, because the rework was never budgeted.

§17pre-registered

The rule was written before any data existed

Committed at 56f1c12. Never amended.

The rule itself is arithmetic: a model is eligible only if it passes build on at least 7 of 8 domains, and among eligible models the highest median weighted quality wins, with the weights fixed in advance. Putting the hash on screen is what makes “I did not pick the rule to suit the result” checkable rather than claimed.

Part two

What “better” means

Three tiers of checking, ordered by cost. Never pay a judge to grade something that failed to parse.

§18key point two

Three tiers, ordered by cost

Never pay a judge to grade something that failed to parse.

The ordering is the methodology lesson, not the tiers. Free deterministic checks first, cheap agents second, the expensive judge last. Run it the other way round and most of your budget grades rubble.

§19tier one · free, milliseconds

Hard gates

Renders. Every page present. Links resolve. No placeholder text. Not truncated. Title, description, one H1, schema. Word floor. Images exist.

Deterministic, free, and finished in milliseconds. If a candidate dies here it never reaches anything that costs money.

§20tier two · cheap, seconds

My production SEO agents

seo-technical · seo-schema · seo-content · seo-geo · seo-performance

The agents I bill clients with. Not a benchmark invented for this stage.

Dogfooding is why this is defensible. It also means the second tier was cut from run 2 on schedule grounds, and R4 fell back to the judge — which is stated in the methodology rather than quietly dropped.

§21tier three · expensive, slow, last

Judge, then my own eye

Pairwise on what survives. Then: is this useful, or is it slop?

The human eye stays in the loop, and it stays last. It is the most expensive instrument in the rig and the one that scales worst, which is the whole reason the other two tiers exist.

§22the weighting

Concept 20 · Design 10 · Content 30 · SEO 40

Resale-first. Fixed before the run. Two alternative weightings are reported for robustness. They do not select.

Design carries the least weight, and it is the number people argue with. The argument is worth having, so the opposing weighting is reported in full rather than ignored — Exhibit D sets out all three and says which one is allowed to pick the winner.

Exhibit D

contents ↑

What “better” means

Four rubrics, each scored 1–10 from different evidence — the planning artifact, the rendered screenshots, the stripped prose, the markup. Then three weightings collapse those four numbers into one, and only one of them is allowed to pick the winner.

The bar

These sites sit on domains that are for sale. The site exists to raise the price, so a buyer has to read it as real, credible and rankable. It does not have to be a finished business — that is a more honest bar than "would you ship this," and it is the bar most of the room actually has too.

The scale

Identical anchors for every rubric, every domain, every candidate. The judge is told not to cluster: where candidates genuinely tie it must say so rather than invent a gap. Verified against two hand-written references of known quality — the deliberate stub scored 2, the professional page 8.

1–2UnusableWould reduce the domain's sale value.
3–4Below barRecognisably machine-made; a buyer would discount it.
5–6AcceptableShips, adds no lustre.
7–8GoodA buyer reads it as a real site.
9–10ExcellentWould pass as commissioned work.
The four rubrics
Concept & keywordsR1
What it seesOnly the planning artifact — the concept, the keyword clusters, the sitemap and the reasoning about the TLD. Not the finished site.
A 9 looks likeA 9 reads the domain accurately, is honest where the name is ambiguous, then commits anyway. Clusters map to real search intent. Every page has a distinct job.
A 3 looks likeA 3 restates the domain name as a concept, calls a synonym list a cluster, and pads to five pages with overlapping jobs.
Design & art directionR2
What it seesRendered screenshots, desktop and mobile — never the source. Plus the stated design rationale and the image manifest.
A 9 looks likeA 9 has a deliberate palette that suits the subject and holds contrast, type that reads as chosen, visible section rhythm, and mobile that was thought about.
A 3 looks likeA 3 is the fallback palette, one undifferentiated column, headings separated only by size, and an image dropped in because one was allowed.
Content qualityR3
What it seesThe plain text of every page, markup stripped, so prose is judged as prose.
A 9 looks likeA 9 is specific: numbers, names, trade-offs, and details only someone who knows the subject would include. Honest about limits.
A 3 looks likeA 3 is generically true — sentences that survive unchanged if you swap the subject. Invented statistics score below that, because they are a liability.
On-page SEO / AEOR4
What it seesHead elements, heading outlines, schema types and the internal link graph.
A 9 looks likeA 9 has correctly sized unique titles written to earn a click, an outline that maps the page's questions, schema that suits the page, and passages an answer engine can quote out of context.
A 3 looks likeA 3 has stuffed or truncated titles, one generic schema block copied everywhere, and 'click here' anchors.
The three weightings
Concept & keywordsDesign & art directionContent qualityOn-page SEO / AEO
W1 · resale-first20%10%30%40%
W2 · equal25%25%25%25%
W3 · buyer-facing10%40%25%25%
W1 · resale-firstThis is the house position, and it selects.

You are not selling websites. You are selling domains, and the site exists to raise the price. SEO is therefore the product and design is the packaging — which is why design carries the least weight here and it is the number people argue with.

W2 · equalThe neutral check.

Declares no opinion. Included because a result that only holds under a favourable weighting is not a result.

W3 · buyer-facingThe opposing case, argued fairly.

If you believe a buyer's first impression drives what they will pay, design dominates. This is the strongest honest argument against W1, so it is reported rather than ignored.

How the winner is chosen

W1 selects, on its own. It was written and committed before a single result existed. W2 and W3 do not select — they exist so you can see whether the answer survives a change in what you value.

All three agreeThe answer does not depend on what you value. That is a robustness result and a stronger claim than any single number.
They disagreeW1 still selects, because it was committed in advance. But name the weighting that flips the answer and say what someone would have to value for that to be true. "It depends on what you're optimising for" is a real finding, not a hedge.
Before quality is ranked at allA model must pass build on at least 7 of 8 domains. A model that cannot produce a working site 7 times out of 8 is not a candidate at any price.
Excluded from every comparisonThe calibration references, which measure the instrument rather than the contestants, and the replicate cells, which exist only to size run-to-run noise.

Source: talk/scoring.py; weights imported from evals/report.py so they cannot drift.

§23the judge

No family overlap with any contestant

Fable judges a lineup of OpenAI, Moonshot, Zhipu and DeepSeek. Blind, order randomised per domain.

Which is exactly why Opus is not a contestant.

Judge contamination is the first question a technical room asks, so it is answered before it is asked.

§24the known weakness

Judges favour their own family

Self-preference bias is real and measurable. Naming it is what separates this from a vendor benchmark.

It is mitigated, not eliminated: two hand-written reference pages of known quality sit in the blind pool to prove the scale is not compressed, and they measured 2 for the deliberate stub and 8 for the professional page. A room vote against the same four sites is the other check.

Part three

Four sites. No labels.

The room voted before it saw a price. You can do the same here — pick one, then scroll.

§25

Which one would you ship?

Four homepages for the same domain, built by four different models from the same frozen prompt. No labels, no prices.

Commit to one before you scroll.

§26the blind vote

waikiki.jp, four times

Candidate A
A
Candidate B
B
Candidate C
C
Candidate D
D

Desktop renders, cropped to the same height. These are the images the design rubric was judged from.

Which model built which

A deepseek-v4-flash-0731 · B glm-5.2 · C glm-5.3-flash · D gpt-5.6-sol. Exhibit E has the state of all fifty-three sites built across both runs.

Exhibit E

contents ↑

Every site built, both runs

Fifty-three sites are on disk across the two runs. This is the cell-by-cell state of all of them: which cleared the build gates, which produced a site that failed one, and which produced nothing at all.

run 123 sites on disk
run 230 sites on disk
every cellmeasured, not published
Run 2 · 20260902-124555-v2

Money as the constraint. Cells marked open cleared the build gates. Cells marked broken produced a site that failed one or more gates but is still on disk and still openable. Replicates are the ~r columns and never score.

modelallergy.jpcheese.jpiconiq.jpiphone-repair.jpsolarpanel.jpwaikiki.jp中華料理.jp田中.jpcheese.jp~r2cheese.jp~r3 built
deepseek-v4-flash-0731openopenbrokenopenno siteopenbrokenbroken··4
glm-5.2brokenbrokenopenbrokenbrokenopenopenopen··4
glm-5.3-flashbrokenbrokenno siteno siteno sitebrokenno siteno siteno siteno site0
gpt-5.6-solopenno siteopenopenopenopenopenopenbrokenbroken7
kimi-k3no siteopenno sitebrokenno siteopenno siteno site··2
Run 1 · 20260901-161631-clean

The 64,000 token ceiling. Preserved unchanged; nothing here was overwritten by run 2.

modelallergy.jpcheese.jpiconiq.jpiphone-repair.jpsolarpanel.jpwaikiki.jp中華料理.jp田中.jpcheese.jp~r2cheese.jp~r3 built
deepseek-v4-flash-0731openopenno siteno siteopenno siteno siteno site··3
glm-5.2brokenno sitebrokenbrokenopenbrokenno siteopen··2
glm-5.3-flashopenno siteno siteopenno siteno siteno siteno site··2
gpt-5.6-solopenopenopenopenopenopenopenopenopenopen10
kimi-k3openno site·no siteno siteno siteopen···2

Source: talk/matrix.py, from each run's gates.json. The built sites themselves are not published.

§27the reveal

Model, cost, tokens

kimi-k3
$1,348
22.2 tok/word
gpt-5.6-sol
$855
5.0 tok/word
glm-5.2
$340
13.0 tok/word
deepseek-v4-flash-0731
$38
11.4 tok/word
glm-5.3-flash
$28
21.9 tok/word

Estimates based on building 800+ sites. Measured cost per site, multiplied out.

Same run the blind vote came from, so the room votes on one set of sites and prices the same set. The spread does the arguing: about a hundred to one from the top rung to the bottom, which is why the bars are on a log scale — a linear bar would render the cheap end invisible.

§28unprompted

Five models. No shared code. One look.

Every one independently chose the same typeface, the same cream background within two percent, and a warm rust accent.

Template collapse does not need a template.

This is the strongest single argument for benchmarking before you commit a batch. Five models from four companies, no shared code, no communication between the lanes — and a buyer who has seen two of these sites will recognise the third. The design prompt now names the failure mode explicitly, which is a thing you only learn by measuring.

Part four

What happened the first time

A clean result, and then the reason it had to be thrown away.

§29key point three

One model cleared the bar

modelbuild pass
gpt-5.6-sol8/8eligible
deepseek-v4-flash-07313/8no
glm-5.22/8no
glm-5.3-flash2/8no
kimi-k32/6no

Rates are over the cells that actually ran. kimi-k3 ran six of its eight domains in run 1 before the run ended, so its denominator is six — Exhibit F reports the same model as 2/8 against the eight that were planned.

gpt-5.6-sol was the only model that cleared the bar.

That is the clean story, and it is the story I would have told if I had stopped here. Eight of eight against a pre-registered threshold of seven of eight. Everything else well under. Exhibit F is the full run.

Exhibit F

contents ↑

Run 1 · what happened

The run with a 64,000-token ceiling as the constraint. Forty cells, twenty-three sites on disk, and one model that cleared the bar. Read the failure table before believing the pass table.

planned40+4
cells run40
sites built23
build pass19
spend$23.95
spent on failures$5.65
git56f1c12
The verdict

gpt-5.6-sol is the only model that cleared the eligibility bar. It built 8/8 against a threshold of 7/8, pre-registered in results/decision-rule.md step 1. It is selected because it was the only model that could reliably build a website — not because it won a quality contest.

By model

Build pass is the eligibility number — the site works: it parses, every page exists, links resolve, images load, no placeholder text shipped. Spec pass is the stricter "did it follow the brief". Tok/word is output tokens spent per word that reached a page. Wasted is money spent on cells that produced no site.

domains run build passspec passreplicatestok/wordspend wasted
gpt-5.6-sol88/81/824.7$10.91$0.000ELIGIBLE
deepseek-v4-flash-073183/81/8018.7$0.58$0.26no
glm-5.282/84/8012.2$2.78$0.55no
glm-5.3-flash82/81/807.4$0.40$0.28no
kimi-k36 (+2 never run)2/81/8025.0$9.28$4.57no
Why things failed

Every failure classified from artifacts on disk. The blame column is the point: a failure is only a fact about the model if it cannot be explained by our network, our timeouts or our bugs — and results/pre-run-findings.md §4 is a list of times it was ours and looked exactly like a finding.

nwhat happened whose faultmeaning
8returned the wrong shapemodelComplete, valid JSON — but not the shape the schema asked for. A dropped connection cannot produce this, so it is the model's.
5spent the whole budget, said nothingmodelConsumed all 128,000 tokens and returned zero bytes. Nothing usable came back at all.
5ran out of budget mid-answermodelCut off mid-string at the pinned 128,000-token ceiling. The budget is identical for every contestant.
3returned nothingambiguousZero bytes, without hitting the ceiling. Cause not established — not counted against the model.
1malformed, cause unclearambiguousUnparseable well below the ceiling. Malformed output and an early-ended stream look identical here — not counted against the model.

18 failure(s) attributable to the model · 0 to us · 4 ambiguous. Not one failure in this run was caused by our infrastructure, so every build failure is a property of the model. Ambiguous failures are never counted against a model.

Every cell

All 40 cells, with the failing stage and the reason. Shots is screenshots rendered — the images R2 is judged from.

Read the token columns carefully. A site is four separate stateless calls — S1 concept, S2 design, S3 build, S4 self-audit — and the 64,000-token cap applies to each call, not to the site. Site total out routinely exceeds the cap with every individual call comfortably under it. The column that the ceiling actually governs is biggest single call; red means it hit the cap and was cut off. In this run every budget failure was S3, the build stage, where five complete HTML pages have to fit inside one JSON response.

domain statusstagesite total outbiggest single call costshots what happened
gpt-5.6-solallergy.jpbuilt · gates pass23,399S3 15,169 (24%)$1.0010/10
gpt-5.6-solcheese.jpbuilt · gates pass23,461S3 15,412 (24%)$1.0010/10
gpt-5.6-solcheese.jp~r2built · gates pass25,541S3 16,775 (26%)$1.0910/10
gpt-5.6-solcheese.jp~r3built · gates pass24,992S3 16,475 (26%)$1.0610/10
gpt-5.6-soliconiq.jpbuilt · gates pass22,438S3 14,458 (23%)$0.9710/10
gpt-5.6-soliphone-repair.jpbuilt · gates pass25,115S3 16,931 (26%)$1.0610/10
gpt-5.6-solsolarpanel.jpbuilt · gates pass25,809S3 16,649 (26%)$1.0810/10
gpt-5.6-solwaikiki.jpbuilt · gates pass23,204S3 14,912 (23%)$0.9910/10
gpt-5.6-sol中華料理.jpbuilt · gates pass33,960S3 22,451 (35%)$1.4010/10
gpt-5.6-sol田中.jpbuilt · gates pass30,973S3 19,796 (31%)$1.2710/10
kimi-k3allergy.jpbuilt · gates pass90,770S3 63,636 (99%)$1.7010/10
kimi-k3cheese.jpno sites379,422S3 64,000 (100%)$1.41ran out of budget mid-answer model
parse failed (Unterminated string starting at: line 31) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget
kimi-k3iphone-repair.jpno sites378,785S3 64,000 (100%)$1.40ran out of budget mid-answer model
parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget
kimi-k3solarpanel.jpno sites378,819S3 64,000 (100%)$1.42ran out of budget mid-answer model
parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget
kimi-k3waikiki.jpno sites316,370S2 9,929 (16%)$0.34returned nothing ambiguous
returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
kimi-k3中華料理.jpbuilt · gates passs4169,415S3 84,473 (132%)$3.0110/10spent the whole budget, said nothing model
spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2
glm-5.2allergy.jpbuilt · gates FAILs461,237S3 44,570 (70%)$0.3910/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this.
glm-5.2cheese.jpno sites312,231S2 7,910 (12%)$0.13returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['slug', 'page_type', 'filename', 'html', 'body_word_count', 'body_char_count']. A dropped connection cannot produce this.
glm-5.2iconiq.jpbuilt · gates FAILgates67,452S3 54,584 (85%)$0.4210/10site built but failed build gates: internal_links_resolve
glm-5.2iphone-repair.jpbuilt · gates FAILgates71,405S3 60,809 (95%)$0.4510/10site built but failed build gates: internal_links_resolve
glm-5.2solarpanel.jpbuilt · gates pass48,207S3 35,633 (56%)$0.3210/10
glm-5.2waikiki.jpbuilt · gates FAILs428,773S3 24,848 (39%)$0.236/6returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['$schema', 'title', 'type', 'additionalProperties', 'required', 'properties']. A dropped connection cannot produce this.
glm-5.2中華料理.jpno sites370,386S3 64,000 (100%)$0.42ran out of budget mid-answer model
parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget
glm-5.2田中.jpbuilt · gates pass69,003S3 35,668 (56%)$0.4210/10
deepseek-v4-flash-0731allergy.jpbuilt · gates passs466,412S3 45,323 (71%)$0.1110/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys []. A dropped connection cannot produce this.
deepseek-v4-flash-0731cheese.jpbuilt · gates passs437,600S3 13,140 (21%)$0.08810/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys []. A dropped connection cannot produce this.
deepseek-v4-flash-0731iconiq.jpno sites376,081S3 64,000 (100%)$0.11spent the whole budget, said nothing model
spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2
deepseek-v4-flash-0731iphone-repair.jpno sites220,193S2 14,891 (23%)$0.016returned the wrong shape model
valid complete JSON, wrong shape — top-level keys [':']. A dropped connection cannot produce this.
deepseek-v4-flash-0731solarpanel.jpbuilt · gates pass82,248S3 52,860 (83%)$0.1210/10
deepseek-v4-flash-0731waikiki.jpno sites15,194S1 5,194 (8%)$0.004returned the wrong shape model
valid complete JSON, wrong shape — top-level keys [':']. A dropped connection cannot produce this.
deepseek-v4-flash-0731中華料理.jpno sites216,661S2 11,047 (17%)$0.014returned the wrong shape model
valid complete JSON, wrong shape — top-level keys [': ']. A dropped connection cannot produce this.
deepseek-v4-flash-0731田中.jpno sites391,449S3 64,000 (100%)$0.11ran out of budget mid-answer model
parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget
glm-5.3-flashallergy.jpbuilt · gates pass29,228S3 17,197 (27%)$0.06310/10
glm-5.3-flashcheese.jpno sites339,174S3 21,947 (34%)$0.054returned nothing ambiguous
returned ZERO bytes after burning 21,947 completion tokens (34% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
glm-5.3-flashiconiq.jpno sites384,986S3 64,000 (100%)$0.066spent the whole budget, said nothing model
spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2
glm-5.3-flashiphone-repair.jpbuilt · gates pass33,226S3 18,951 (30%)$0.06410/10
glm-5.3-flashsolarpanel.jpno sites16,419S1 6,419 (10%)$0.002malformed, cause unclear ambiguous
parse failed (Expecting ',' delimiter: line 124 column) on 7,431 bytes; 6,419 completion tokens, 10.0% of cap — well below the ceiling, so not budget exhaustion; malformed output and an earl
glm-5.3-flashwaikiki.jpno sites389,348S3 64,000 (100%)$0.078spent the whole budget, said nothing model
spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2
glm-5.3-flash中華料理.jpno sites212,689S1 8,329 (13%)$0.004returned nothing ambiguous
returned ZERO bytes after burning 4,360 completion tokens (7% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
glm-5.3-flash田中.jpno sites368,463S3 64,000 (100%)$0.072spent the whole budget, said nothing model
spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2

Source: evals/summary.py, run 20260901-161631-clean. This page deliberately needs no judge data.

§30the turn

Then I looked at why they failed

8returned the wrong shape
5spent the whole budget, said nothing
5ran out of budget mid-answer
3returned nothing
1malformed, cause unclear

18 the model · 0 us · 4 ambiguous

This is the most important beat in the talk. The pass table above is a fact about the models only if the failures cannot be explained by my network, my timeouts or my bugs. So every failure was classified from the artifacts on disk, and the blame column is the part that matters.

§31the instrument

I capped every model at the same number of tokens

modelcost of hitting the same wall
kimi-k3$1.16
glm-5.2$0.32
deepseek-v4-flash$0.049
glm-5.3-flash$0.018

Same event. Sixty-four times the spread. Tokens are a proxy for money, and a bad one.

Ten of the twenty-two failures in run 1 were calls my own ceiling cut off. A constraint that looked identical for every contestant was not identical at all — it was a different amount of money for each of them, and the cheap models were being handed sixty-four times more rope than the expensive one.

§32and one of them

This one was mine.

One broken line in my own template dropped every navigation item after the first. Ten of twenty-three sites, three models, and not one worked around it.

I was about to score my own bug as their taste.

Stated flatly, because it is rigor rather than confession. There was a second one too: a prompt pointed at a schema file on disk. I was sure that explained the malformed JSON, removed it for run 2 — and malformed JSON went from eight to ten. gpt-5.6-sol was never affected by it and still built eight of eight in run 1.

§33

I was measuring my instrument.

So I threw it out and re-ran it.

§34run two

Money as the constraint

kimi-k3
$1,348
22.2 tok/word2/8
gpt-5.6-sol
$855
5.0 tok/word7/8
glm-5.2
$340
13.0 tok/word4/8
deepseek-v4-flash-0731
$38
11.4 tok/word4/8
glm-5.3-flash
$28
21.9 tok/word0/8

Estimates based on building 800+ sites. Measured cost per site, multiplied out. The pass rate is beside each rung.

A cost governor replaces the token ceiling: every model gets the same money, not the same tokens. Same domains, same models, same rubrics, same judge, same weights. Three things changed in the harness, so this is v1 against v2 as a package rather than a controlled single change — and saying so is part of the result.

§35the finding

Pass rate, not mean score

At eight hundred sites you cannot inspect them all. Excellent four times in five is not cheap. It is an unbudgeted rework project.

The question is not which is best. It is which fails least.

This is where the two tables stop agreeing. Under all three weightings the highest median quality belongs to kimi-k3. It built two of eight. gpt-5.6-sol came second on quality and built seven of eight, and the pre-registered rule selects on eligibility first — so the model that wins the quality contest is not the model you buy.

§36every model, plotted

Quality against what you actually spend

4.55.46.37.18.08.9$0.12$0.31$0.62$1.25$3.12$6.25$12.50median weighted quality (W1)cost per passing sitecheaper & better ↘gpt-5.6-sol median quality 7.70 (worst 6.70, best 8.50) pass 7/8 · $1.26 per passing sitegptkimi-k3 median quality 7.85 (worst 7.40, best 8.30) pass 2/8 · $6.87 per passing sitekimiglm-5.2 median quality 5.20 (worst 4.90, best 5.80) pass 4/8 · $0.94 per passing siteglmdeepseek-v4-flash-0731 median quality 5.55 (worst 5.10, best 7.10) pass 4/8 · $0.15 per passing sitedeepseek
gpt-5.6-solkimi-k3glm-5.2deepseek-v4-flash-0731glm-5.3-flash

Each dot is a model. The whisker runs from its worst site to its best, so a wide whisker is the variance argument made visible. Cost is per passing site on a log axis, because failures get re-run and the ladder spans two decades. Down and to the right wins. You are not buying the dot, you are buying the whisker.

Exhibit G below is the full scorecard behind this plot: the four rubric grids, the three weightings side by side, whether each rubric is still separating the models, every gate result, and what it cost.

Exhibit G

contents ↑

Run 2 · the scorecard

The graded report: four rubric grids, the three weightings side by side, whether each rubric is still separating the models, the three plots, every gate result, and what it all cost. Read the weighting table and the pass rate together — they do not point at the same model, and that disagreement is the finding.

models5
domains8
sites built30/44
judgements32
generation$21.60
judging$12.51
gitf0c5162
Scores

Rows are models, columns are domains, cells are the judge's 1–10 score, coloured by the shared anchor bands. Overall is weighted, never averaged — the three weightings answer different questions, and a model that wins under all three is a robustness result rather than one number.

Every median on this page is over build-passing sites only, per step 2 of results/decision-rule.md. A site that did not build is struck through and left out of the median — never scored zero and averaged in, which would destroy the information the pass rate exists to carry. Spec failures do not affect the median; they carry separately as the spec pass rate in the Gates table below.

Concept & keywordsR1
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol98789999
kimi-k38898
glm-5.2675876586
deepseek-v4-flash-073156767856
glm-5.3-flash899
Design & art directionR2
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol57788767
kimi-k38938.5
glm-5.2668865876.5
deepseek-v4-flash-073167757446.5
glm-5.3-flash459
Content qualityR3
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol45779797
kimi-k39788
glm-5.2564656866
deepseek-v4-flash-073164645655
glm-5.3-flash888
On-page SEO / AEOR4
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol89788988
kimi-k38777.5
glm-5.2484964434
deepseek-v4-flash-073165867876
glm-5.3-flash468
Weighted overallW1 · resale-first
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol6.707.407.007.708.508.208.307.70
kimi-k38.307.407.307.85
glm-5.24.907.004.607.805.905.105.805.305.20
deepseek-v4-flash-07315.805.107.105.306.407.005.705.55
glm-5.3-flash6.007.108.30
Weighted overallW2 · equal
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol6.507.257.007.758.508.008.007.75
kimi-k38.257.756.758.00
glm-5.25.256.755.257.756.005.256.256.005.62
deepseek-v4-flash-07315.755.507.005.256.506.505.255.62
glm-5.3-flash6.007.008.50
Weighted overallW3 · buyer-facing
waikiki.jpcheese.jpiphone-repair.jpallergy.jpsolarpanel.jp中華料理.jpiconiq.jp田中.jpmedian
gpt-5.6-sol5.907.107.007.758.357.707.557.55
kimi-k38.257.905.858.07
glm-5.25.256.605.707.755.855.106.705.855.55
deepseek-v4-flash-07315.905.657.005.106.505.905.105.78
glm-5.3-flash5.406.408.50
Do the three weightings agree?

Each weighting collapses the four rubric scores differently. W1 selects the winner — it was committed before any data existed. W2 and W3 are reported so you can see whether the answer survives a change in what you value.

W1 · resale-firstW2 · equalW3 · buyer-facing
medianrankmedianrankmedianrank
gpt-5.6-sol7.70#27.75#27.55#2
kimi-k37.85#18.00#18.07#1
glm-5.25.20#45.62#35.55#4
deepseek-v4-flash-07315.55#35.62#45.78#3
glm-5.3-flash

All three weightings pick kimi-k3. The answer does not depend on what you value — that is a robustness result, and a stronger claim than any single number.

Is each rubric still discriminating?

Spread is the distance from the highest model median to the lowest, on that rubric. A rubric where every model lands within 2 points has stopped separating them, and a ranking built from it is a ranking of rounding.

gpt-5.6-solkimi-k3glm-5.2deepseek-v4-flash-0731glm-5.3-flashcell rangespreadverdict
Concept & keywordsR19.08.06.06.05–93.0discriminating
Design & art directionR27.08.56.56.55–92.0discriminating
Content qualityR37.08.06.05.04–93.0discriminating
On-page SEO / AEOR48.07.54.06.03–94.0discriminating

All four rubrics discriminate. Every one spans at least 2 points across model medians, so each is carrying information rather than agreeing with itself.

How big is a gap that means nothing?

The same model, the same domain, the same frozen prompt, run again. Whatever that moves the score by is the floor under every comparison on this page — read the grid against it, not against zero. 2 replicate cell(s) failed build gates and are excluded on the same basis as the grid — a cell that did not build measures failure, not variance.

No judged replicate cells.

No judged replicate cells yet — run-to-run noise is unmeasured, so no gap in the grid can be called real.

The plots

Order is the argument. Plot 2 collapses each model to one dot, and that collapse is exactly the mistake this benchmark exists to warn about — a model that is brilliant six times and unusable twice averages into a dot sitting comfortably beside a boring reliable one. Plot 1 earns plot 2. Cost is on a log axis because the ladder spans two decades; a linear axis flattens four of five models into the floor. The score grids above are the table view of the same numbers.

Plot 1 — every sitequality vs cost
4.15.16.17.08.09.0$0.031$0.062$0.12$0.31$0.62$1.25$3.12weighted quality (W1)cost per sitecheaper & better ↘gpt-5.6-sol · waikiki.jp quality 6.70 · $0.99gpt-5.6-sol · iphone-repair.jp quality 7.40 · $1.05gpt-5.6-sol · allergy.jp quality 7.00 · $1.02gpt-5.6-sol · solarpanel.jp quality 7.70 · $1.07gpt-5.6-sol · 中華料理.jp quality 8.50 · $1.31gpt-5.6-sol · iconiq.jp quality 8.20 · $0.94gpt-5.6-sol · 田中.jp quality 8.30 · $1.35kimi-k3 · waikiki.jp quality 8.30 · $1.85kimi-k3 · cheese.jp quality 7.40 · $1.85kimi-k3 · iphone-repair.jp quality 7.30 · $1.45glm-5.2 · waikiki.jp quality 4.90 · $0.38glm-5.2 · cheese.jp quality 7.00 · $0.56glm-5.2 · iphone-repair.jp quality 4.60 · $0.39glm-5.2 · allergy.jp quality 7.80 · $0.42glm-5.2 · solarpanel.jp quality 5.90 · $0.30glm-5.2 · 中華料理.jp quality 5.10 · $0.70glm-5.2 · iconiq.jp quality 5.80 · $0.45glm-5.2 · 田中.jp quality 5.30 · $0.54deepseek-v4-flash-0731 · waikiki.jp quality 5.80 · $0.090deepseek-v4-flash-0731 · cheese.jp quality 5.10 · $0.14deepseek-v4-flash-0731 · iphone-repair.jp quality 7.10 · $0.080deepseek-v4-flash-0731 · allergy.jp quality 5.30 · $0.093deepseek-v4-flash-0731 · 中華料理.jp quality 6.40 · $0.047deepseek-v4-flash-0731 · iconiq.jp quality 7.00 · $0.035deepseek-v4-flash-0731 · 田中.jp quality 5.70 · $0.061glm-5.3-flash · waikiki.jp quality 6.00 · $0.040glm-5.3-flash · cheese.jp quality 7.10 · $0.031glm-5.3-flash · allergy.jp quality 8.30 · $0.034
gpt-5.6-solkimi-k3glm-5.2deepseek-v4-flash-0731glm-5.3-flash

Shapes to read: a tight cluster is a model you can run unattended 800 times. A wide horizontal smear — same cost, wildly different quality — is one you cannot. Vertical spread means token usage swings by domain, so you cannot budget it.

Plot 2 — every modelmedian quality vs cost per passing site
4.55.46.37.18.08.9$0.12$0.31$0.62$1.25$3.12$6.25$12.50median weighted quality (W1)cost per passing sitecheaper & better ↘gpt-5.6-sol median quality 7.70 (worst 6.70, best 8.50) pass 7/8 · $1.26 per passing sitegptkimi-k3 median quality 7.85 (worst 7.40, best 8.30) pass 2/8 · $6.87 per passing sitekimiglm-5.2 median quality 5.20 (worst 4.90, best 5.80) pass 4/8 · $0.94 per passing siteglmdeepseek-v4-flash-0731 median quality 5.55 (worst 5.10, best 7.10) pass 4/8 · $0.15 per passing sitedeepseek
gpt-5.6-solkimi-k3glm-5.2deepseek-v4-flash-0731glm-5.3-flash
Plot 3 · what a token actually costs

Every built site: output tokens on X, what they cost on Y, both axes measuring the same generation. Run 1 capped every model at 64,000 completion tokens and treated that as one budget for everyone. This plot is why that was the wrong variable — find two dots at the same horizontal position and read off the vertical gap. A call that spends exactly 64,000 tokens costs $1.16 on kimi-k3 and $0.018 on glm-5.3-flash: the same event, 64× apart. Tokens are a proxy for money and a bad one. Note the axis direction — here fewer tokens is better, so the good corner is bottom-left, not bottom-right.

9k37k64k91k119k146k$0.031$0.062$0.12$0.31$0.62$1.25$3.12output tokens per sitecost per site↙ cheaper & leanergpt-5.6-sol · waikiki.jp 23,347 output tokens · $0.99 $0.042 per 1,000 tokensgpt-5.6-sol · iphone-repair.jp 24,802 output tokens · $1.05 $0.043 per 1,000 tokensgpt-5.6-sol · allergy.jp 24,014 output tokens · $1.02 $0.043 per 1,000 tokensgpt-5.6-sol · solarpanel.jp 25,829 output tokens · $1.07 $0.042 per 1,000 tokensgpt-5.6-sol · 中華料理.jp 31,375 output tokens · $1.31 $0.042 per 1,000 tokensgpt-5.6-sol · iconiq.jp 22,548 output tokens · $0.94 $0.042 per 1,000 tokensgpt-5.6-sol · 田中.jp 33,308 output tokens · $1.35 $0.041 per 1,000 tokenskimi-k3 · waikiki.jp 99,877 output tokens · $1.85 $0.019 per 1,000 tokenskimi-k3 · cheese.jp 100,494 output tokens · $1.85 $0.018 per 1,000 tokenskimi-k3 · iphone-repair.jp 84,342 output tokens · $1.45 $0.017 per 1,000 tokensglm-5.2 · waikiki.jp 56,712 output tokens · $0.38 $0.007 per 1,000 tokensglm-5.2 · cheese.jp 94,583 output tokens · $0.56 $0.006 per 1,000 tokensglm-5.2 · iphone-repair.jp 62,853 output tokens · $0.39 $0.006 per 1,000 tokensglm-5.2 · allergy.jp 62,410 output tokens · $0.42 $0.007 per 1,000 tokensglm-5.2 · solarpanel.jp 38,807 output tokens · $0.30 $0.008 per 1,000 tokensglm-5.2 · 中華料理.jp 120,118 output tokens · $0.70 $0.006 per 1,000 tokensglm-5.2 · iconiq.jp 69,105 output tokens · $0.45 $0.007 per 1,000 tokensglm-5.2 · 田中.jp 93,752 output tokens · $0.54 $0.006 per 1,000 tokensdeepseek-v4-flash-0731 · waikiki.jp 39,085 output tokens · $0.090 $0.002 per 1,000 tokensdeepseek-v4-flash-0731 · cheese.jp 102,063 output tokens · $0.14 $0.001 per 1,000 tokensdeepseek-v4-flash-0731 · iphone-repair.jp 39,750 output tokens · $0.080 $0.002 per 1,000 tokensdeepseek-v4-flash-0731 · allergy.jp 58,230 output tokens · $0.093 $0.002 per 1,000 tokensdeepseek-v4-flash-0731 · 中華料理.jp 24,818 output tokens · $0.047 $0.002 per 1,000 tokensdeepseek-v4-flash-0731 · iconiq.jp 37,490 output tokens · $0.035 $0.001 per 1,000 tokensdeepseek-v4-flash-0731 · 田中.jp 71,094 output tokens · $0.061 $0.001 per 1,000 tokensglm-5.3-flash · waikiki.jp 132,689 output tokens · $0.040 $0.000 per 1,000 tokensglm-5.3-flash · cheese.jp 110,708 output tokens · $0.031 $0.000 per 1,000 tokensglm-5.3-flash · allergy.jp 109,957 output tokens · $0.034 $0.000 per 1,000 tokens
gpt-5.6-solkimi-k3glm-5.2deepseek-v4-flash-0731glm-5.3-flash

Whiskers run worst site to best. Cost is mean spend divided by pass rate, because failures get re-run. You are not buying the dot, you are buying the whisker. The dashed line is the Pareto frontier — anything above and left of it is dominated.

Gates

Build means the site works. Spec means it followed the brief. Warn is a real quality signal that is not a correctness failure. Count err is the model's own word or character count against the measured one — how well it knows what it just wrote, which is exactly what unattended operation depends on.

domainbuild specwarncheckscount errfailures
gpt-5.6-solwaikiki.jpPASSFAIL097+35.2%
5 failing
  • specpalette_tokens_defined airport-and-transport.html missing ['--c-fg-muted']
  • specpalette_tokens_defined costs-payments-and-tipping.html missing ['--c-fg-muted']
  • specpalette_tokens_defined first-trip-itinerary.html missing ['--c-fg-muted']
  • specpalette_tokens_defined index.html missing ['--c-fg-muted']
  • specpalette_tokens_defined where-to-stay.html missing ['--c-fg-muted']
gpt-5.6-solcheese.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
gpt-5.6-soliphone-repair.jpPASSFAIL097+31.8%
6 failing
  • specpalette_tokens_defined authorized-vs-independent.html missing ['--c-fg-muted']
  • specword_band authorized-vs-independent.html 1472 words, band 1500-3000 (blog, measured as en)
  • specpalette_tokens_defined before-repair.html missing ['--c-fg-muted']
  • specpalette_tokens_defined index.html missing ['--c-fg-muted']
  • specpalette_tokens_defined repair-costs.html missing ['--c-fg-muted']
  • specpalette_tokens_defined repair-faq.html missing ['--c-fg-muted']
gpt-5.6-solallergy.jpPASSPASS097+39.5%clean
gpt-5.6-solsolarpanel.jpPASSFAIL097+36.9%
5 failing
  • specpalette_tokens_defined choose-solar-installer.html missing ['--c-fg-muted']
  • specpalette_tokens_defined home-solar-planning.html missing ['--c-fg-muted']
  • specpalette_tokens_defined index.html missing ['--c-fg-muted']
  • specpalette_tokens_defined solar-panel-cost-japan.html missing ['--c-fg-muted']
  • specpalette_tokens_defined solar-subsidies-japan.html missing ['--c-fg-muted']
gpt-5.6-sol中華料理.jpPASSFAIL198+28.8%
2 failing
  • warnlinks_portable site 38 root-relative links (valid at domain root, unnavigable locally): ['chiiki-ryori.html->/', 'chiiki-ryori.html->/chiiki-ryori.htm
  • specidn_punycode_in_urls site raw IDN in 5 canonical(s); expected xn--fiqw48ctgkppp.jp
gpt-5.6-soliconiq.jpPASSFAIL079+32.5%
4 failing
  • specpalette_tokens_defined app-icon-localization.html missing ['--c-fg-muted']
  • specpalette_tokens_defined icon-localization-checklist.html missing ['--c-fg-muted']
  • specpalette_tokens_defined index.html missing ['--c-fg-muted']
  • specpalette_tokens_defined japanese-symbols.html missing ['--c-fg-muted']
gpt-5.6-sol田中.jpPASSFAIL098+19.1%
2 failing
  • specidn_punycode_in_urls site raw IDN in 5 canonical(s); expected xn--fiqx46g.jp
  • specword_band tanaka-distribution.html 1757 chars, band 800-1600 (category, measured as ja)
kimi-k3waikiki.jpPASSFAIL097+21.8%
2 failing
  • specword_band things-to-do.html 991 words, band 400-800 (category, measured as en)
  • specword_band waikiki-itinerary.html 1197 words, band 1500-3000 (blog, measured as en)
kimi-k3cheese.jpPASSFAIL097+32.2%
1 failing
  • specword_band japanese-cheese.html 1207 words, band 1500-3000 (blog, measured as en)
kimi-k3iphone-repair.jpFAILPASS197+21.8%
5 failing
  • warnlinks_portable site 62 root-relative links (valid at domain root, unnavigable locally): ['faq.html->/', 'faq.html->/repair-options.html']
  • buildimages_exist faq.html missing ['images/iphone-backup-before-repair.png']
  • buildimages_exist index.html missing ['images/iphone-repair-japan-hero.png']
  • buildimages_exist iphone-repair-tokyo.html missing ['images/tokyo-ginza-apple-store-street.png']
  • buildimages_exist repair-prices.html missing ['images/iphone-screen-battery-parts.png']
kimi-k3allergy.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s1', 's2', 's3', 's4']
kimi-k3solarpanel.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
kimi-k3中華料理.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
kimi-k3iconiq.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s1', 's2', 's3', 's4']
kimi-k3田中.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s1', 's2', 's3', 's4']
glm-5.2waikiki.jpPASSFAIL097-4.2%
5 failing
  • spechas_valid_jsonld about.html missing or unparseable
  • spechas_valid_jsonld index.html missing or unparseable
  • spechas_valid_jsonld waikiki-beaches.html missing or unparseable
  • spechas_valid_jsonld waikiki-dining.html missing or unparseable
  • spechas_valid_jsonld waikiki-travel-tips.html missing or unparseable
glm-5.2cheese.jpFAILFAIL197-13.9%
3 failing
  • buildinternal_links_resolve site 71 broken: ['cheese-pairing-japan.html->/japanese-cheesemakers', 'cheese-pairing-japan.html->/where-to-buy', 'cheese-pairing-japan
  • warnlinks_portable site 86 root-relative links (valid at domain root, unnavigable locally): ['cheese-pairing-japan.html->/', 'cheese-pairing-japan.html->/
  • specword_band japanese-cheesemakers.html 926 words, band 400-800 (category, measured as en)
glm-5.2iphone-repair.jpFAILFAIL097-10.4%
11 failing
  • spechas_valid_jsonld battery-replacement.html missing or unparseable
  • buildno_placeholder_text battery-replacement.html XXXX
  • spechas_valid_jsonld faq.html missing or unparseable
  • buildno_placeholder_text faq.html XXXX
  • spechas_valid_jsonld index.html missing or unparseable
  • specno_skipped_headings index.html [(1, 3)]
  • buildno_placeholder_text index.html XXXX
  • spechas_valid_jsonld screen-repair.html missing or unparseable
  • buildno_placeholder_text screen-repair.html XXXX
  • spechas_valid_jsonld water-damage-repair.html missing or unparseable
  • buildno_placeholder_text water-damage-repair.html XXXX
glm-5.2allergy.jpFAILFAIL197-14.5%
4 failing
  • buildinternal_links_resolve site 64 broken: ['allergy-treatment-japan.html->/cedar-pollen-allergy-japan', 'allergy-treatment-japan.html->/food-allergies-japan', 'a
  • warnlinks_portable site 73 root-relative links (valid at domain root, unnavigable locally): ['allergy-treatment-japan.html->/', 'allergy-treatment-japan.h
  • specword_band faq.html 1817 words, band 800-1600 (faq, measured as en)
  • specword_band food-allergies-japan.html 1405 words, band 600-1200 (landing, measured as en)
glm-5.2solarpanel.jpFAILFAIL197-13.4%
7 failing
  • buildinternal_links_resolve site 49 broken: ['index.html->/residential-solar-installation-japan', 'index.html->/japan-solar-subsidies-fit-fip', 'index.html->/solar
  • warnlinks_portable site 59 root-relative links (valid at domain root, unnavigable locally): ['index.html->/', 'index.html->/']
  • spechas_valid_jsonld index.html missing or unparseable
  • spechas_valid_jsonld japan-solar-subsidies-fit-fip.html missing or unparseable
  • spechas_valid_jsonld residential-solar-installation-japan.html missing or unparseable
  • spechas_valid_jsonld solar-panel-cost-japan.html missing or unparseable
  • spechas_valid_jsonld solar-panel-japan-faq.html missing or unparseable
glm-5.2中華料理.jpPASSFAIL098-3.0%
6 failing
  • specidn_punycode_in_urls site raw IDN in 5 canonical(s); expected xn--fiqw48ctgkppp.jp
  • spechas_valid_jsonld index.html missing or unparseable
  • spechas_valid_jsonld ingredients.html missing or unparseable
  • spechas_valid_jsonld recipes.html missing or unparseable
  • spechas_valid_jsonld regional-cuisines.html missing or unparseable
  • spechas_valid_jsonld restaurant-guide.html missing or unparseable
glm-5.2iconiq.jpPASSFAIL097-17.1%
5 failing
  • specno_skipped_headings about.html [(2, 4)]
  • specword_band about.html 948 words, band 400-800 (about, measured as en)
  • specno_skipped_headings architecture.html [(2, 4)]
  • specno_skipped_headings design.html [(2, 4)]
  • specno_skipped_headings food.html [(2, 4)]
glm-5.2田中.jpPASSFAIL098+6.9%
7 failing
  • specidn_punycode_in_urls site raw IDN in 5 canonical(s); expected xn--fiqx46g.jp
  • spechas_valid_jsonld family-crest.html missing or unparseable
  • spechas_valid_jsonld famous-tanaka.html missing or unparseable
  • spechas_valid_jsonld faq.html missing or unparseable
  • spechas_valid_jsonld index.html missing or unparseable
  • spechas_valid_jsonld name-origin.html missing or unparseable
  • specword_band name-origin.html 2929 chars, band 3000-6000 (blog, measured as ja)
deepseek-v4-flash-0731waikiki.jpPASSFAIL097+12.9%
1 failing
  • specword_band faq.html 680 words, band 800-1600 (faq, measured as en)
deepseek-v4-flash-0731cheese.jpPASSPASS097-2.5%clean
deepseek-v4-flash-0731iphone-repair.jpPASSFAIL097+11.7%
3 failing
  • specword_band battery-repair.html 747 words, band 800-1600 (service, measured as en)
  • specword_band faq.html 742 words, band 800-1600 (faq, measured as en)
  • specword_band screen-repair.html 796 words, band 800-1600 (service, measured as en)
deepseek-v4-flash-0731allergy.jpPASSPASS097-3.2%clean
deepseek-v4-flash-0731solarpanel.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
deepseek-v4-flash-0731中華料理.jpFAILFAIL098+63.5%
6 failing
  • specword_band index.html 874 chars, band 1000-2000 (homepage, measured as ja)
  • specword_band reshipi.html 1063 chars, band 3000-6000 (blog, measured as ja)
  • buildimages_exist reshipi.html missing ['images/chuka-stirfry.png']
  • specword_band taberu.html 1020 chars, band 1200-2400 (landing, measured as ja)
  • buildimages_exist taberu.html missing ['images/chuka-tabehodai.png']
  • buildimages_exist zairyo.html missing ['images/chuka-seasonings.png']
deepseek-v4-flash-0731iconiq.jpFAILFAIL097+17.5%
7 failing
  • specword_band brand-design.html 706 words, band 800-1600 (service, measured as en)
  • buildimages_exist brand-design.html missing ['images/brand-design-process.png']
  • buildimages_exist index.html missing ['images/index-hero.png']
  • buildimages_exist japanese-logos.html missing ['images/logos-hanko.png']
  • specword_band japanese-minimalism.html 739 words, band 1500-3000 (blog, measured as en)
  • buildimages_exist japanese-minimalism.html missing ['images/minimalism-tea.png']
  • buildimages_exist japanese-products.html missing ['images/products-still-life.png']
deepseek-v4-flash-0731田中.jpFAILFAIL098-100.0%
5 failing
  • buildimages_exist inaka-guide.html missing ['images/inaka-guide-family-walk.png']
  • buildimages_exist index.html missing ['images/tanaka-hero-ricefield.png']
  • specword_band nouhaku.html 1336 chars, band 1600-3200 (service, measured as ja)
  • buildimages_exist nouhaku.html missing ['images/nouhaku-farmstay-table.png']
  • buildimages_exist tanada.html missing ['images/tanada-autumn-harvest.png']
glm-5.3-flashwaikiki.jpFAILFAIL097+4.8%
10 failing
  • spechas_valid_jsonld index.html missing or unparseable
  • buildimages_exist index.html missing ['images/waikiki-diamond-head-golden-hour.png']
  • spechas_valid_jsonld waikiki-beaches.html missing or unparseable
  • buildimages_exist waikiki-beaches.html missing ['images/waikiki-beginner-surf-lesson.png']
  • spechas_valid_jsonld waikiki-dining.html missing or unparseable
  • buildimages_exist waikiki-dining.html missing ['images/waikiki-japanese-counter-sushi.png']
  • spechas_valid_jsonld waikiki-faq.html missing or unparseable
  • buildimages_exist waikiki-faq.html missing ['images/waikiki-beach-early-morning-canoes.png']
  • spechas_valid_jsonld waikiki-hotels.html missing or unparseable
  • buildimages_exist waikiki-hotels.html missing ['images/waikiki-hotel-lanai-view.png']
glm-5.3-flashcheese.jpFAILFAIL097-0.4%
15 failing
  • spechas_valid_jsonld cheese-in-japan-faq.html missing or unparseable
  • specpalette_tokens_defined cheese-in-japan-faq.html missing ['--c-fg-muted']
  • buildimages_exist cheese-in-japan-faq.html missing ['images/sliced-cheese-stack.png']
  • spechas_valid_jsonld cheese-shops-tokyo.html missing or unparseable
  • specpalette_tokens_defined cheese-shops-tokyo.html missing ['--c-fg-muted']
  • buildimages_exist cheese-shops-tokyo.html missing ['images/cheesemonger-wrapping.png']
  • spechas_valid_jsonld cheese-tourism-japan.html missing or unparseable
  • specpalette_tokens_defined cheese-tourism-japan.html missing ['--c-fg-muted']
  • buildimages_exist cheese-tourism-japan.html missing ['images/hokkaido-factory-window.png']
  • spechas_valid_jsonld index.html missing or unparseable
  • specpalette_tokens_defined index.html missing ['--c-fg-muted']
  • buildimages_exist index.html missing ['images/tokyo-cheese-counter.png']
  • spechas_valid_jsonld japanese-cheese-makers.html missing or unparseable
  • specpalette_tokens_defined japanese-cheese-makers.html missing ['--c-fg-muted']
  • buildimages_exist japanese-cheese-makers.html missing ['images/hokkaido-dairy-dawn.png']
glm-5.3-flashiphone-repair.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
glm-5.3-flashallergy.jpFAILPASS197-0.6%
5 failing
  • warnlinks_portable site 70 root-relative links (valid at domain root, unnavigable locally): ['faq.html->/', 'faq.html->/']
  • buildimages_exist faq.html missing ['images/japanese-drugstore-aisle.png']
  • buildimages_exist japan-pollen-season.html missing ['images/sugi-cedar-pollen-spring.png']
  • buildimages_exist japan-travel-food-allergies.html missing ['images/conbini-packaged-food-labels.png']
  • buildimages_exist japanese-allergy-card.html missing ['images/allergy-card-restaurant-table.png']
glm-5.3-flashsolarpanel.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
glm-5.3-flash中華料理.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
glm-5.3-flashiconiq.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
glm-5.3-flash田中.jpFAILFAIL01
1 failing
  • buildstages_completed site no site built; missing ['s3', 's4']
Cost

Measured, not estimated. ×800 uses the median built-site cost, converted at ¥160 to $1.

built median $/site×800tokens intokens out
gpt-5.6-sol7/8$1.05$843233,312185,223
kimi-k33/8$1.85$1,47785,319284,713
glm-5.28/8$0.44$349295,870598,340
deepseek-v4-flash-07317/8$0.080$64232,875372,530
glm-5.3-flash3/8$0.034$2794,687353,354
Evidence

Every score cites what the judge actually observed. Blind and order-randomised — it did not know which model produced which candidate. Of the 112 written judgements this run produced, the highest- and lowest-scored cell on each rubric is reproduced here: the two ends of the scale are what show whether the scale is being used at all.

Concept & keywordsR1
9gpt-5.6-sol田中.jp

evidenceSame surname-reference read as B, but executed with more discipline. The plan explicitly refuses to invent where the record is ambiguous: '由来を単一の物語として断定せず' and the research page's premise that '同姓だけでは血縁を判断できない' — precisely the epistemic honesty the rubric rewards. Cluster separation is the sharpest of the three: origin, prefecture-level distribution (including why counts differ by source), genealogy how-to (戸籍, 菩提寺, 旧土地台帳), and a romanization page (田中 ローマ字, パスポート 表記) that captures a genuinely distinct practical intent B misses entirely. Buyer types are named concretely (系譜調査サービス, 田中を屋号に持つ企業). Written natively in Japanese, matching its own language recommendation.

weaknessesTLD reasoning is correct but thinner than A's — it gestures at an English expansion without weighing it. All five clusters are informational; no commercial-intent page, which slightly limits the 'plausible business' impression for a buyer.

Concept & keywordsR1
5glm-5.2iconiq.jp

evidenceCompetent surface reasoning ('the stylized spelling suggests an international or cosmopolitan positioning... a defensible niche on a .jp domain') but the concept collapses into a generic 'iconic Japan' culture guide — landmarks, food, architecture, design — which is closer to restating the domain broadly than building a differentiated brand. Cluster c5 is padding dressed as strategy: 'about ICONIQ Japan guide' with secondary keywords like 'editorial standards' is not a real search cluster, and it exists only to justify a fifth page. Three editorial features are mislabeled page_type 'service', suggesting the schema was filled mechanically. The food and architecture sections are massive, hyper-competitive query spaces where a five-page site has no chance of authority.

weaknessesNo commercial angle anywhere — pure informational publication with no monetization or lead-gen story for a buyer. About-page keyword cluster is invented to hit five pages. Topical breadth (food + architecture + design + landmarks) dilutes rather than builds authority.

Design & art directionR2
9kimi-k3cheese.jp

evidenceThe indigo system is executed end to end: accent #2B3A9E appears in the logo, link arrows, the callout's left rule, and a full-bleed indigo band ('Sixty years of cheese in Japan') whose four-column timeline (1870s / 1964 / 1970s / 2000s–now) gives the page real rhythm between white sections. The art direction is the strongest of the set — the hero board is shot on indigo-dyed linen so the commissioned image literally states the palette, exactly as the manifest claims ('doubles as the palette statement for the whole site'). Japanese glyphs render cleanly in card headings ('Sakura(さくら)', '燻製チーズ'), proving the Noto Sans JP decision. Pill badges under the hero ('~50% of Japan's milk from Hokkaido') and the boxed 'What is cheese in Japan actually like?' callout show section-level variety, and mobile restacks hero → pills → image → callout in a considered order.

weaknessesThe three hero pills and three 'Start where you are' cards are near-duplicate content, and the bordered callout box's long paragraphs run a little dense on mobile. Space Grotesk headings at mid weights ('Three cheeses to try first') sit close to body color and could carry more contrast.

Design & art directionR2
3kimi-k3iphone-repair.jp

evidenceThe hero image is broken in both desktop and mobile screenshots — the page renders a bordered empty box with the raw alt text ('A cracked iPhone awaiting screen replacement…') where the photograph should be. This is the single most prominent element on the page and it is visibly failed, which is exactly what makes a buyer discount a site as machine-generated. Beneath that, the page is competent but plain: three channel cards and a pricing table have clear labels ('Apple Store (Genius Bar)', 'Screen: ~¥25,000–¥45,000 · 1–3 days'), the blue #0B5FFF CTA reads deliberately, and the tabular yen pricing matches the rationale's 'tabular numerals' claim. The manifest itself is thoughtful — the FAQ backup image ('back up before handing over the phone') is genuinely purposeful art direction — but none of it survives the broken render on the one page shown.

weaknessesBroken hero image on the homepage in both viewports; dense, small-set body text with weak sectional rhythm below the fold; mobile is a single reflowed column with the failed image box occupying a large dead area.

Content qualityR3
9gpt-5.6-sol田中.jp

evidenceThis reads as commissioned work by someone who knows Japanese genealogy. The research-methods page walks through 除籍 and 改製原戸籍, 旧土地台帳 and 地籍図, warns that 本籍 is not the same as residence, tells the reader to annotate inferred locations as '比定', and gives a four-axis cross-checking table (氏名/年代/続柄/地名) with realistic failure modes (数え年 vs. dates, 養子 breaking the surname/bloodline link). The distribution page refuses to invent a ranking and instead explains *why* rankings differ (集計年, 母集団, 電話帳の偏り, 異体字の統合) — exactly the honesty-about-limits the rubric rewards. The romanization page is practically useful: Surname-field guidance, the passport-match-over-'correct'-spelling principle, and the TANAKA, Taro index format. Every page answers a distinct question; zero near-duplication.

weaknessesDense and austere — it withholds even the satisfying basic facts (never states the well-known rank or population figure), which is epistemically defensible but may frustrate casual readers. The romanization checklist verges on over-explaining trivial cases (Taanaka/Tanakaa).

Content qualityR3
4deepseek-v4-flash-0731allergy.jp

evidenceReads as competent but generic — sentences like 'Eating in Japan with a food allergy is manageable, but it requires preparation' and 'If you are still struggling, see a doctor. Allergic rhinitis is treatable' would survive with the country swapped. The restaurant risk table is the one genuinely useful artifact. Factual problems: the OTC table lists Allegra as '120 mg, once daily' (Japanese OTC Allegra FX is 60 mg twice daily) and includes Desalex/desloratadine as OTC, which is prescription-only in Japan. 'Tonkatsu sauce contains wheat and sometimes peanut' looks invented. The forecast page attributes the daily pollen count to the 'Japan Meteorological Association' (it is the Japan Weather Association) and hard-codes '2025 season' claims ('predicted higher-than-average sugi counts for Kanto and Chubu') that read as fabricated forecast specifics and will date instantly. The clinics page is the thinnest: 'look for clinics near major stations like Shinjuku, Shibuya, or Tokyo Station that advertise English support' is padding, not guidance.

weaknessesRecognisable model cadence throughout, several wrong medication facts stated without hedging, invented-looking seasonal forecast specifics, and a clinics page that says almost nothing actionable.

On-page SEO / AEOR4
9gpt-5.6-soliphone-repair.jp

evidenceThe strongest answer-engine execution of the four. Nearly every H2 is phrased as a directly answerable question ('Why should coverage and repair history be checked first?', 'When should Find My be turned off?'), so each section can stand alone as a quotable passage. Schema is varied and page-appropriate: HowTo on the before-repair checklist, FAQPage on the FAQ, Article on the comparison, WebSite on index. The FAQ has 8 topical H2 groupings with 28 specific H3 questions ('Should I put a wet iPhone in rice?', 'Do quoted prices include Japanese consumption tax?') plus in-page anchor navigation (#damage, #liquid) and body-copy cross-links with descriptive, varied anchors ('backup and privacy preparation sequence', 'repair cost comparison worksheet'). All five pages interlink; no orphans, no duplicate titles.

weaknessesTwo truncated anchors on index ('authorized versus independent repair com', 'Japan repair cost and turnaround workshe'). Title 'Authorized vs Independent iPhone Repair Japan Guide' reads slightly keyword-ordered rather than natural. FAQ page's link section repeats the same three targets with two anchor variants each.

On-page SEO / AEOR4
3glm-5.2田中.jp

evidenceEvery title is grossly overstuffed and will truncate hard, e.g. family-crest.html: 「田中家の家紋一覧|木の字紋・桔梗紋・片喰紋など代表的な紋章とその意味・歴史的背景を詳しく解説する家紋参考事典」(~55+ chars). Meta descriptions run 130–160+ characters and read as essays restating the page (「本ページでは…詳しく解説します」boilerplate on multiple pages). schema_type is null on all five pages — even faq.html, whose h3s are literally formatted 「Q:田中姓は日本で何番目に多い名字ですか?」, an obvious missed FAQPage. internal_links is an empty array on every page, so despite 「関連ページ」 h2s appearing in outlines, the delivered link graph makes all five pages orphans.

weaknessesZero schema, zero internal links, uniformly oversized titles/descriptions — the three core brief requirements are all failed. Heading outlines are shallow but serviceable; that is the only element executed to bar.

Source: evals/report.py, run 20260902-124555-v2. Weighted quality is defined in evals/, which is authoritative on measurement.

§37token utilization

A cheap model that rambles is not cheap

kimi-k3
22.2 tok/word
$1,348 for 800
gpt-5.6-sol
5.0 tok/word
$855 for 800
glm-5.2
13.0 tok/word
$340 for 800
deepseek-v4-flash-0731
11.4 tok/word
$38 for 800
glm-5.3-flash
21.9 tok/word
$28 for 800

Bar is tokens spent per word that reached a page. The projected 800-site cost is beside it.

The bar is ramble here, not cost. The rambliest models are the cheap ones, and their spend per useful word is where cheap stops being cheap: glm-5.3-flash burns four and a half times as many tokens per delivered word as gpt-5.6-sol. Exhibit H carries the per-model token and word counts these ratios come from.

Exhibit H

contents ↑

Run 2 · what happened

The re-run with money as the constraint instead of tokens. Same domains, same models, same rubrics, same judge, same weights — three things changed in the harness, so this is v1 against v2 as a package, not a controlled single change.

planned40+4
cells run44
sites built30
build pass17
spend$21.60
spent on failures$2.19
gitf0c5162
The verdict

gpt-5.6-sol is the only model that cleared the eligibility bar. It built 7/8 against a threshold of 7/8, pre-registered in results/decision-rule.md step 1. It is selected because it was the only model that could reliably build a website — not because it won a quality contest.

By model

Build pass is the eligibility number — the site works: it parses, every page exists, links resolve, images load, no placeholder text shipped. Spec pass is the stricter "did it follow the brief". Tok/word is output tokens spent per word that reached a page. Wasted is money spent on cells that produced no site.

domains run build passspec passreplicatestok/wordspend wasted
gpt-5.6-sol87/81/825.0$10.15$0.30ELIGIBLE
deepseek-v4-flash-073184/82/8011.4$0.60$0.059no
glm-5.284/80/8013.0$3.76$0.000no
kimi-k382/81/8022.2$6.88$1.73no
glm-5.3-flash80/81/8221.9$0.21$0.11no
Why things failed

Every failure classified from artifacts on disk. The blame column is the point: a failure is only a fact about the model if it cannot be explained by our network, our timeouts or our bugs — and results/pre-run-findings.md §4 is a list of times it was ours and looked exactly like a finding.

nwhat happened whose faultmeaning
10returned the wrong shapemodelComplete, valid JSON — but not the shape the schema asked for. A dropped connection cannot produce this, so it is the model's.
4returned nothingambiguousZero bytes, without hitting the ceiling. Cause not established — not counted against the model.
4malformed, cause unclearambiguousUnparseable well below the ceiling. Malformed output and an early-ended stream look identical here — not counted against the model.
4budget_governor?budget_governor
3answered_nothing?answered_nothing

10 failure(s) attributable to the model · 0 to us · 8 ambiguous. Not one failure in this run was caused by our infrastructure, so every build failure is a property of the model. Ambiguous failures are never counted against a model.

Every cell

All 44 cells, with the failing stage and the reason. Shots is screenshots rendered — the images R2 is judged from.

Read the token columns carefully. A site is four separate stateless calls — S1 concept, S2 design, S3 build, S4 self-audit — and the 128,000-token cap applies to each call, not to the site. Site total out routinely exceeds the cap with every individual call comfortably under it. The column that the ceiling actually governs is biggest single call; red means it hit the cap and was cut off. In this run every budget failure was S3, the build stage, where five complete HTML pages have to fit inside one JSON response.

domain statusstagesite total outbiggest single call costshots what happened
gpt-5.6-solallergy.jpbuilt · gates pass24,014S3 16,352 (13%)$1.0210/10
gpt-5.6-solcheese.jpno sites35,771S2 3,382 (3%)$0.30returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['$schema', 'title', 'type', 'additionalProperties', 'required', 'properties']. A dropped connection cannot produce this.
gpt-5.6-solcheese.jp~r2built · gates FAILgates24,410S3 16,627 (13%)$1.0010/10site built but failed build gates: images_exist
gpt-5.6-solcheese.jp~r3built · gates FAILgates27,458S3 17,783 (14%)$1.1110/10site built but failed build gates: images_exist
gpt-5.6-soliconiq.jpbuilt · gates pass22,548S3 14,947 (12%)$0.948/8
gpt-5.6-soliphone-repair.jpbuilt · gates pass24,802S3 16,162 (13%)$1.0510/10
gpt-5.6-solsolarpanel.jpbuilt · gates pass25,829S3 15,968 (12%)$1.0710/10
gpt-5.6-solwaikiki.jpbuilt · gates pass23,347S3 14,909 (12%)$0.9910/10
gpt-5.6-sol中華料理.jpbuilt · gates pass31,375S3 19,348 (15%)$1.3110/10
gpt-5.6-sol田中.jpbuilt · gates pass33,308S3 20,779 (16%)$1.3510/10
kimi-k3allergy.jpno sites17,345S1 7,345 (6%)$0.13returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['domain_reading', 'concept', 'audience', 'tld_reasoning', 'keyword_clusters']. A dropped connection cannot produce this.
kimi-k3cheese.jpbuilt · gates pass100,494S3 60,662 (47%)$1.8510/10
kimi-k3iconiq.jpno site0$0.000budget_governor ?
kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable
kimi-k3iphone-repair.jpbuilt · gates FAILs484,342S3 70,005 (55%)$1.4510/10budget_governor ?
kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable
kimi-k3solarpanel.jpno sites360,088S3 43,682 (34%)$1.09malformed, cause unclear ambiguous
parse failed (Invalid control character at: line 23 co) on 61,822 bytes; 43,682 completion tokens, 34.1% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ea
kimi-k3waikiki.jpbuilt · gates pass99,877S3 62,958 (49%)$1.8510/10
kimi-k3中華料理.jpno sites325,664S2 19,620 (15%)$0.50budget_governor ?
kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable
kimi-k3田中.jpno site0$0.000budget_governor ?
kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable
glm-5.2allergy.jpbuilt · gates FAILgates62,410S3 26,199 (20%)$0.4210/10site built but failed build gates: internal_links_resolve
glm-5.2cheese.jpbuilt · gates FAILs494,583S3 51,526 (40%)$0.5610/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this.
glm-5.2iconiq.jpbuilt · gates pass69,105S4 41,974 (33%)$0.4510/10
glm-5.2iphone-repair.jpbuilt · gates FAILgates62,853S3 49,629 (39%)$0.3910/10site built but failed build gates: no_placeholder_text
glm-5.2solarpanel.jpbuilt · gates FAILs438,807S3 32,936 (26%)$0.3010/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['site_jsonld']. A dropped connection cannot produce this.
glm-5.2waikiki.jpbuilt · gates passs456,712S3 31,417 (25%)$0.3810/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this.
glm-5.2中華料理.jpbuilt · gates passs4120,118S3 97,748 (76%)$0.7010/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this.
glm-5.2田中.jpbuilt · gates passs493,752S3 87,670 (68%)$0.5410/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['refusal']. A dropped connection cannot produce this.
deepseek-v4-flash-0731allergy.jpbuilt · gates pass58,230S3 42,590 (33%)$0.09310/10
deepseek-v4-flash-0731cheese.jpbuilt · gates pass102,063S3 57,469 (45%)$0.1410/10
deepseek-v4-flash-0731iconiq.jpbuilt · gates FAILs437,490S2 16,396 (13%)$0.03510/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations']. A dropped connection cannot produce this.
deepseek-v4-flash-0731iphone-repair.jpbuilt · gates pass39,750S3 17,372 (14%)$0.08010/10
deepseek-v4-flash-0731solarpanel.jpno sites36,878S1 4,404 (3%)$0.059returned nothing ambiguous
returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
deepseek-v4-flash-0731waikiki.jpbuilt · gates pass39,085S3 17,535 (14%)$0.09010/10
deepseek-v4-flash-0731中華料理.jpbuilt · gates FAILs424,818S3 14,902 (12%)$0.04710/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations']. A dropped connection cannot produce this.
deepseek-v4-flash-0731田中.jpbuilt · gates FAILgates71,094S4 48,966 (38%)$0.06110/10site built but failed build gates: images_exist
glm-5.3-flashallergy.jpbuilt · gates FAILs4109,957S3 73,432 (57%)$0.03410/10returned the wrong shape model
valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this.
glm-5.3-flashcheese.jpbuilt · gates FAILs4110,708S3 93,095 (73%)$0.03110/10returned nothing ambiguous
returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
glm-5.3-flashcheese.jp~r2no sites323,215S2 13,043 (10%)$0.007malformed, cause unclear ambiguous
parse failed (Unterminated string starting at: line 33) on 70,732 bytes; 0 completion tokens, 0.0% of cap — well below the ceiling, so not budget exhaustion; malformed output and an early-en
glm-5.3-flashcheese.jp~r3no sites322,522S2 16,731 (13%)$0.007answered_nothing ?
returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned, decided it was finished, and emitted nothing. Run 2 keeps
glm-5.3-flashiconiq.jpno sites323,585S2 15,229 (12%)$0.007returned nothing ambiguous
returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
glm-5.3-flashiphone-repair.jpno sites394,875S3 76,134 (59%)$0.027malformed, cause unclear ambiguous
parse failed (Expecting value: line 1 column 1 (char 0) on 75,399 bytes; 76,134 completion tokens, 59.5% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ea
glm-5.3-flashsolarpanel.jpno sites364,477S3 39,724 (31%)$0.019answered_nothing ?
returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned for 39,724 completion tokens, decided it was finished, and
glm-5.3-flashwaikiki.jpbuilt · gates FAILs4132,689S3 97,249 (76%)$0.04010/10malformed, cause unclear ambiguous
parse failed (Extra data: line 78 column 1 (char 3378)) on 3,380 bytes; 14,259 completion tokens, 11.1% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ear
glm-5.3-flash中華料理.jpno sites320,780S2 13,729 (11%)$0.006returned nothing ambiguous
returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established
glm-5.3-flash田中.jpno sites3127,411S3 104,562 (82%)$0.036answered_nothing ?
returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned for 104,562 completion tokens, decided it was finished, and

Source: evals/summary.py, run 20260902-124555-v2. This page deliberately needs no judge data.

§38the shape of the operation

Five independent calls. No conductor.

the samefrozen promptgpt-5.6-soldeepseek-v4-flash-0731glm-5.2glm-5.3-flashkimi-k3the sameshape back

No lane can see any other lane. They converge anyway.

This one diagram argues both halves of the talk. It is why fan-out works — no shared state, no orchestration, each lane restartable on its own — and it is why eight hundred sites come back looking like one operation. Nobody up there is being told what to play, and they still play the same tune.

§39the failure mode

And this one stopped checking in.

One call ran for forty minutes, spent about sixty dollars, and returned nothing at all.

Nobody was watching. At eight hundred sites, nobody ever is.

The provider reported finish_reason=stop. Not truncation, not a dropped connection. It reasoned, decided it was finished, and emitted zero bytes of answer. Reasoning is billed as completion tokens and is not the answer. This is the cost of unattended fan-out stated plainly, and it is why the thing you are actually buying is a failure rate you can catch.

Part five

The decision

What committing now actually means, and the five steps that outlive the answer.

§40the decision

What committing now means

kimi-k3
$1,348
22.2 tok/word2/8
gpt-5.6-sol
$855
5.0 tok/word7/8
glm-5.2
$340
13.0 tok/word4/8
deepseek-v4-flash-0731
$38
11.4 tok/word4/8
glm-5.3-flash
$28
21.9 tok/word0/8

The eligible rung is lit. The rest are dimmed because the pre-registered rule removed them before quality was ranked at all.

gpt-5.6-sol, at a projected $855 to build eight hundred sites, accepting a one-in-eight build failure rate, with a free deterministic gate that catches every one of those failures before it reaches a live domain. That is the decision, stated with the failure rate in it rather than left out.

It is not the model that scored highest on quality. That was kimi-k3, under every one of the three weightings, and it built two sites in eight. Buying it would have been buying a rework project I had not budgeted.

§41

The method

  1. State the job before you evaluate it.
  2. Write the decision rule down before you have data.
  3. Order your checks by cost. Gates first, judges last.
  4. Sample enough for a pass rate, not an anecdote.
  5. Buy on pass rate, not on average score.

The answer above has a shelf life measured in weeks — a new model ships and the ladder changes. The five steps do not. The rig outlives the talk, and re-running it against a new lineup is a morning's work rather than a project.

§42

Thank you.

Questions, the raw runs, or the harness — that address reaches me.


Appendix · questions from the room

A1 · Methodology

How a site gets built. Each stage is a stateless call.

Inline above ↓

A2 · Scoring

Four rubrics, three weightings, how the winner is chosen.

Inline above ↓

A3 · Stage chain

The four independent calls, and why they are not a conversation.

Inline above ↓

A4 · Full run summary

Every cell, every failure, every token count.

Inline above ↓

A5 · Judge self-preference

The cross-judge matrix. Bias measured, not hand-waved. Two hand-written reference pages of known quality sit in the blind pool as the calibration probe, and they measured 2 for the deliberate stub and 8 for the professional page — so the scale is not compressed. Family overlap is removed by construction: the judge shares no vendor with any contestant.

A6 · Regional pricing

japan-kimi is 48% cheaper than kimi. Same model, Japan-hosted. Regional endpoints are a real lever on the ladder and they are not on it here, because mixing hosting regions into a model comparison puts a per-model cost advantage into the result. That is a confound, not a saving.

A7 · Version pinning

Floating tags silently invalidate an eval. Every model in both runs is pinned to an exact identifier, the pricing snapshot is recorded in the run manifest, and the prompt and template SHAs are recorded beside it. Without that, a re-run six weeks later is a different experiment wearing the same name.

Colophon

This is the written version of a talk given at Hawaii Tech Week on 3 September 2026. Every measured figure on this page is resolved from the two runs' own JSON at build time — gates.json, efficiency.json, failures.json, judge.json — rather than typed in, so a number here and a number in the run are the same number.

Costs are shown in US dollars, converted from the gateway's yen billing at ¥160 to $1, the same rate the evaluation code uses so the page never mixes two rates. Prices are this gateway key's rates and not a public price sheet. A Japanese edition of this page, in yen, is the companion to this one.

Two runs, forty-four cells each, five models, eight domains, one judge that shares no family with any contestant, and one decision rule written before any of it existed.

Try GPT-5.6 Sol right now

GPT-5.6 Sol is available on FastMetal through one API key. Start in the browser, or call it from the OpenAI SDK.