Which model do I trust with eight hundred websites?
Hawaii Tech Week · 3 September 2026
I own eight hundred dropped domains and about four hours a week. This is the benchmark I built to decide which language model gets to write on all of them, and the two runs it took to stop measuring my own instrument.
Eight hundred domains, four hours a week
Where the problem comes from, and why my credentials are not an answer to it.
§1the portfolio
Over eight hundred domains
No websites on any of them.
A hand-picked sample of the portfolio, roughly a quarter of it Japanese IDN. These are dropped domains with real search value and nothing on them.
- waikiki.jp
- paris.jp
- berlin.jp
- amsterdam.jp
- athens.jp
- copenhagen.jp
- brussels.jp
- budapest.jp
- geneva.jp
- montreal.jp
- moscow.jp
- denver.jp
- detroit.jp
- houston.jp
- sandiego.jp
- sanfrancisco.jp
- manhattan.jp
- monterey.jp
- napa.jp
- cancun.jp
- caribbean.jp
- kualalumpur.jp
- kathmandu.jp
- johannesburg.jp
- beijing.jp
- hong-kong.jp
- liverpool.jp
- beverlyhills.jp
- neworleans.jp
- washingtondc.jp
- netherlands.jp
- deutschland.jp
- poland.jp
- sweden.jp
- switzerland.jp
- austria.jp
- greece.jp
- bulgaria.jp
- cyprus.jp
- serbia.jp
- ukraine.jp
- egypt.jp
- colombia.jp
- uruguay.jp
- saudiarabia.jp
- bahamas.jp
- seychelles.jp
- tanzania.jp
- new-zealand.jp
- sri-lanka.jp
- alabama.jp
- alaska.jp
- colorado.jp
- illinois.jp
- kentucky.jp
- ohio.jp
- oregon.jp
- virginia.jp
- cheese.jp
- cigars.jp
- cola.jp
- dinner.jp
- buffet.jp
- bowling.jp
- surfing.jp
- skate.jp
- jetski.jp
- golfclub.jp
- winery.jp
- tapas.jp
- teppanyaki.jp
- restaurants.jp
- supermarket.jp
- ice-cream.jp
- cookbook.jp
- tuna.jp
- beaujolais.jp
- olympics.jp
- presentation.jp
- online.jp
- wifi.jp
- usb.jp
- electronics.jp
- electricity.jp
- investing.jp
- investment.jp
- securities.jp
- hedgefund.jp
- venturecapital.jp
- wallstreet.jp
- dollar.jp
- funding.jp
- lease.jp
- advertising.jp
- iptv.jp
- ultrabook.jp
- allergy.jp
- diabetes.jp
- dermatology.jp
- cleanenergy.jp
- solarpanel.jp
- ecoenergy.jp
- iconiq.jp
- iphone-repair.jp
- カフェ.jp
- デザイン.jp
- システム.jp
- ビーチ.jp
- ツーリズム.jp
- ナビゲーション.jp
- スマートフォン.jp
- クレジットカード.jp
- ソーラー.jp
- ウィスキー.jp
- イタリアワイン.jp
- カリフォルニアワイン.jp
- アロハ.jp
- オアフ.jp
- 中華料理.jp
- 日本料理.jp
- 野菜.jp
- 九州.jp
- 博多.jp
- 富士.jp
- 丸の内.jp
- 香港.jp
- 交通.jp
- 眼科.jp
- 元気.jp
- 代替エネルギー.jp
- 新エネルギー.jp
- カーボン市場.jp
§2hawaii tech week · 3 september 2026
Which model do I trust with eight hundred websites?
The whole talk is one question, and it is a purchasing question rather than a technology one.
§3qualifications
You want to know my references?
- Pioneered the Japanese domain market from 1999
- Board, Japan Domain Name Business Association
- Vice-board, Japan Internet Providers Association
- ICANN member
- Millions of dollars in Japanese-language paid media
- Companies and offices in Honolulu, Tokyo and Hanoi
- EO board member
- Twenty Ironmans
- Airplane and helicopter pilot
- Two Pacific crossings as captain
Sold Solis, Japan's third largest domain registrar, to Global Media Online in 2005.
Tokyo listed. Multi-billion dollar. Owns Japan's largest registrar.
The list is all true, and it escalates into the irrelevant on purpose. Two Pacific crossings tell you nothing about which language model to buy. That is the point of putting them on the same slide as the registrar sale.
§4what I actually run
Real operating companies, not slideware. Honolulu, Tokyo, Hanoi. Each mark links to its site.
§5
None of this tells you which model to use.
Credentials are hearsay with a letterhead. This talk argues against them, including my own.
§6built with claude
Two sites
Both are live. Both took many sessions with my eye on every screen.
§7
I'm not a designer.
Six months ago I could not have built either of those. That is the point: the machine closed that gap.
Can it close it eight hundred times, when I am not looking?
That is the whole difficulty. Those two sites took many sessions with my attention on every screen. I cannot give eight hundred sites that kind of attention, and I am not going to pretend otherwise.
§8the point of view
Eight hundred dropped domains.
Real search value. No websites. The play is not flipping them, it is putting real content on them so the value compounds.
Eight hundred domains, and about four hours a week to spend on them. That constraint is what makes this a measurement problem instead of a craft problem.
§9disclosure
I own the gateway.
fastmetal.ai
Every model in this talk runs through it, and I resell all of them. I am also its largest customer. Four of my companies run on it.
So do not take my word for anything today. Take the numbers.
fastmetal.ai ↗open the live siteenabled on my key
- anthropic-claude-fable-5
- anthropic-claude-fable-5-1
- anthropic-claude-haiku-4-5
- anthropic-claude-opus-4-6
- anthropic-claude-opus-4-7
- anthropic-claude-opus-4-8
- anthropic-claude-opus-5
- anthropic-claude-sonnet-4-6
- anthropic-claude-sonnet-5
- bytedance-seedream-4.5
- deepseek-v4-flash
- deepseek-v4-flash-0731
- deepseek-v4-pro
- gemini-3.5-flash
- gemini-3.7-flash
- gemini-flash-lite-free
- glm-4.7
- glm-4.7-flash
- glm-5
- glm-5.1
- glm-5.2
- glm-5.3
- glm-5.3-flash
- google-nano-banana-2
- gpt-5.6-luna
- gpt-5.6-sol
- gpt-5.6-terra
- gpt-oss-120b
- grok-4.5
- grok-4.6
- inkling
- japan-gemma-4-31b
- japan-kimi-k2.6
- japan-kimi-k2.7-code
- japan-qwen3.6-35b
- kimi-k2.6
- kimi-k3
- llm-jp-3.1-8x13b-instruct4
- mimo-v2.5
- mimo-v2.5-pro
- minimax-m2.7
- minimax-m3
- mistral-voxtral-mini-3b-2507
- muse-glimmer-30b
- muse-spark-1.2
- qwen3.6-27b
- qwen3.7-max
- qwen3.8-27b
- qwen3.8-2.4t-a95b
- qwen3.8-max
- solar-pro4
- z-image-turbo
The roster is what is enabled on my key, not the whole catalogue — the service advertises far more. Saying this early and at full strength is cheaper than having it surface later.
Part one
How the experiment was set up
Almost everybody picks a model from hearsay. Someone said it was good, so it is the one they use. You cannot evaluate until you can state the job.
§10key point one
You cannot evaluate until you can state the job
Almost everybody picks a model from hearsay. Someone said it was good, so it is the one they use.
This section buys the rest of the talk. If the setup is not believable, no result that comes out of it is worth reading. So it goes first, in full, before a single number.
§11the workload
A five-page site that has to earn its ranking
Real topical content, on a domain that has never had any. Design that holds up next to a working business site. On-page SEO that earns the ranking. Writing a person would actually read.
Five pages that make the domain worth more than it was.
Exhibit A below is the harness itself: what is fixed, what the model decides, and what gets written to disk. Everything tinted in its centre lane is the model's work. Everything else is byte-identical for all five contestants.
Exhibit A
How this actually works
There is no orchestrating model. The harness is a loop making direct HTTP calls to a gateway — no agent, no tool use, no retries on content. A language model appears once more at the end, as the judge, and it shares no family with any contestant. Everything tinted in the centre lane is the model's work; everything else is byte-identical for all five.
One domain, one model
Forty cells, one named model
Source: talk/methodology.py, generated from harness/config.json.
§12the prompts
Five prompts, frozen, identical for every model
An identical preamble, then concept, design, build and self-audit. Frozen once the first production run started, and hashed into every run's manifest.
In the room this slide was a wall of unreadable type, and that was the point — the audience reads volume, not content. On a page you can do better than volume, so Exhibit B carries all five in full.
Exhibit B
Five prompts, frozen, identical for every model
The complete text of every prompt in the chain, exactly as sent. Their SHA is recorded in each run's manifest, so what is printed here is checkable against what actually ran.
Prepended verbatim to every call, for every model. It states the job, the five mechanical constraints that fail a build outright, and the sentence the whole talk turns on: assume nobody will inspect your output before it goes live.
# Shared preamble — prepended verbatim to every stage, for every model **FROZEN once the first production run starts. Identical for all contestants.** --- You are building a small, genuinely useful website for a domain that is currently for sale. The site's job is to make the domain worth more: real topical content that can rank in search and be cited by answer engines. It is not a parked page, not a placeholder, and not a template fill. The domain owner has roughly 800 such domains and will run this same process on all of them without reviewing each result. Assume nobody will inspect your output before it goes live. Build accordingly. ## Hard constraints These are checked mechanically. Violating any of them fails the build outright. 1. **At most 5 pages.** Fewer is fine if fewer is right. 2. **At most 1 image per page.** Five images total, maximum. 3. **Word and character bands** — every page must fall inside the band for its declared page type. Under the floor is thin content; over the ceiling is padding. Both fail. | Page type | EN words | JA characters | |---|---|---| | homepage | 500 – 1,000 | 1,000 – 2,000 | | service | 800 – 1,600 | 1,600 – 3,200 | | landing | 600 – 1,200 | 1,200 – 2,400 | | category | 400 – 800 | 800 – 1,600 | | about | 400 – 800 | 800 – 1,600 | | faq | 800 – 1,600 | 1,600 – 3,200 | | blog | 1,500 – 3,000 | 3,000 – 6,000 | 4. **Every page must be uniquely valuable.** Not the same page with the nouns swapped. Google's doorway-page algorithm exists and this owner has 800 domains — near-duplicate pages get the whole portfolio deindexed. 5. **The shared floor is fixed.** `base.css` is linked, never modified, never inlined. ## Output Return **only** valid JSON matching the schema given for your stage. No prose outside the JSON, no markdown fences, no commentary.
s1-conceptconcept and keywords
Input is the bare domain and the build language. No brief, no market research. Deriving what the site should be is the job, and being honest about an ambiguous name then committing anyway is what scores.
# S1 · Concept & keyword architecture
## Input
domain: {{DOMAIN}}
build_language: {{BUILD_LANGUAGE}} # en | ja — fixed, applies to the whole site
That is all you get. No brief, no market research, no instructions about what the site
should be about. Deriving that is the job.
`build_language` is settled and not yours to change — plan the sitemap, slugs and working
titles in it from the start. You are still asked below what language you *would* have
chosen; say so plainly even when it differs, and then plan in the language given.
## What to do
**1. Read the domain.** What does this name imply? What would someone typing it expect?
What would a buyer of this domain want to own? Be honest when a domain is ambiguous or
carries little meaning — say so, and then make a defensible choice anyway.
**2. Consider the TLD.** It is part of the name and it carries information. Reason about
what it implies for audience, market and language, and state your reasoning. If it points
somewhere the rest of the name does not, say what you would do about it. Name the language
you would have picked and why — disagreeing with `build_language` is a legitimate answer
and is scored on the quality of the argument, not on agreement.
**3. Design the site.** Concept, audience, and a sitemap of at most five pages. Every page
needs a distinct job and a declared page type from the table in the preamble. Do not pad
to five.
**4. Keyword architecture.** Clusters, not a keyword list. For each cluster: a primary
term, secondary and semantic variations, the search intent behind it, and which page owns
it. Every page in the sitemap must own exactly one cluster.
## What is being judged
Whether you read the domain correctly and designed a site worth building — coherence,
commercial realism, keyword and intent coverage, sitemap logic, and the quality of your
reasoning about the TLD. Everything downstream is built on this, so a weak concept costs
you four times over.s2-designdesign direction
Palette, typography, per-page section plan and the image manifest. Note the instruction not to reach for the house style — cream, terracotta, serif display — written into the prompt because every model had already converged on it.
# S2 · Design direction & image manifest
## Input
domain: {{DOMAIN}}
build_language: {{BUILD_LANGUAGE}} # en | ja — final, not a suggestion
s1: {{S1_JSON}}
`build_language` is the language the site will be written in. It may differ from what you
recommended in S1; it is fixed now and applies to all page copy, headings, alt text and
metadata.
**If it differs from your S1 recommendation, commit to it fully.** Page slugs, working
titles and section names all follow `build_language` — an English site with romanised
Japanese slugs is worse than either language done properly. Rename anything from S1 that
no longer fits and note the corrected slugs in your page layouts.
## The shared floor
Every contestant links the same `base.css`. It provides a reset, a fluid type scale
(`--step--1` … `--step-5`), a spacing scale (`--space-3xs` … `--space-3xl`), and layout
primitives: `.container`, `.stack` / `.stack-m` / `.stack-l`, `.prose`, `.section`,
`.grid`, `.cluster`, `.switcher`, `.center`, `.visually-hidden`, `.skip-link`.
It sets **no palette, no typeface, no section layout, and no page structure.** Those are
yours, and they are the thing being judged. The fallback palette is pure black on pure
white with no accent — if you skip it, the site looks exactly as unconsidered as it is.
## What to do
**1. Design rationale.** Two or three sentences: what this site should feel like and why
that suits this subject and audience. Not adjectives — an argument.
**2. Palette.** Concrete hex values for all seven `--c-*` tokens. Must pass WCAG AA for
body text on background.
**Do not reach for the house style.** Language models converge hard on one look: a warm
cream background around `#FBF7F0`, a serif display face, and a terracotta or amber accent.
Measured across five models from four different companies on this exact task, every one
produced that palette and every one chose the same typeface. It is the most recognisable
signature of a generated website, and a buyer who has seen two of them recognises the
third.
Derive the palette and the type from **this** subject and **this** audience. If your first
instinct is cream-and-terracotta with a serif display, that is your first instinct, not
your best one — justify it against the subject or choose again. A B2B industrial site, a
health reference and a destination guide should not arrive at the same colours.
**3. Typography.** Heading and body stacks. Google Fonts is permitted via a single
stylesheet link; if you use it, name the exact families and weights. System stacks are a
legitimate choice, not a cop-out — Japanese builds especially.
**4. Section plan per page.** For each page in the sitemap, the ordered sections and what
each is for. This is where design variance actually lives: what a page is made of and in
what order.
**5. Image manifest.** **At most one image per page, five total.** Fewer is a valid
choice. Every image is generated by the same fixed image model at the same settings for
every contestant, so image *quality* is a constant — what is judged is your art
direction: subject choice, prompt quality, placement, and alt text.
**Every image comes back as a 1024×1024 PNG.** The image model ignores size and aspect
requests, so design for square and crop with CSS if you want something else. A layout that
assumes a wide hero will break.
Alt text is written in `build_language` and does real SEO work. Do not describe the
image; say what it contributes.
## What is being judged
The rendered result, scored from desktop and mobile screenshots, plus art direction. Your
rationale here is read as evidence of intent.s3-buildfive pages of html
Content, design and on-page SEO in one response, as complete final HTML. There is no later pass to fix anything in. This is the stage where every budget failure happened.
# S3 · Build
## Input
domain: {{DOMAIN}}
build_language: {{BUILD_LANGUAGE}}
s1: {{S1_JSON}}
s2: {{S2_JSON}}
## What to do
Emit the **complete, final HTML for every page** in the sitemap. Content, design and
on-page SEO all at once — there is no later pass to fix things in. Write it correctly the
first time.
Each page conforms to the contract in `page.html`:
- `<html lang>` matches `build_language`
- `<title>` 50–60 characters, unique across the site, primary keyword present
- `<meta name="description">` 150–160 characters, written to earn a click
- self-referencing absolute `<link rel="canonical">`
- Open Graph: `og:title`, `og:description`, `og:type`, `og:url`, `og:image`
- exactly one JSON-LD block; the `@type` is your choice and should suit the page
- `<link rel="stylesheet" href="base.css">` — never modified, never inlined
- one `<style>` block implementing your S2 palette and typography; the seven `--c-*`
tokens must all be defined
- `<a class="skip-link">`, `<header>`, `<main id="main">`, `<footer>`
- exactly one `<h1>` inside `<main>`; heading levels never skipped
- images referenced at `images/<filename>` exactly as named in your S2 manifest
## On-page craft — how to write for search
Follow these while writing, not afterwards.
**Structure carries intent.** Headings are a map of the page, not decoration. Each H2
answers a distinct question a searcher actually has. If a heading could sit on any page
about anything, it is doing no work.
**Answer first.** Open every section with a direct, self-contained answer in one or two
sentences, then expand. A passage lifted out of context must still make sense and still
be attributable — that is what makes it quotable by an answer engine.
**Keywords are placed, not sprinkled.** Primary term in the title, the H1, the first
hundred words, and naturally thereafter. Semantic and related variations throughout.
Density in the 1–3% range. Anything that reads as stuffing has failed twice: once with the
reader, once with the ranking.
**Specific beats true.** "Athens has a rich history" is true and worthless. Numbers,
names, dates, prices, comparisons, trade-offs, and things only someone who knows the
subject would think to mention. Generic truth is the signature of generated content, and
it is what a reader detects before they can name it.
**Show experience and be checkable.** Concrete detail, honest limitations, and outbound
links to genuinely authoritative sources where a claim needs backing. Never invent a
statistic, a citation, a review, or a credential. On health, legal, safety or financial
subjects, be conservative and say what you do not know — a confident wrong answer is a
liability, not a style problem.
**Scannable and readable.** Short paragraphs. Lists and tables where the content is
genuinely a list or a table, never to fill space. Plain sentences.
**Link with intent.** Body-copy links to other pages on this site using descriptive anchor
text — never "click here", never the bare page title every time. Every page reachable from
the navigation. No orphans.
**Every page must earn its own existence.** If two pages could be swapped by changing a
few nouns, you have built a doorway network and this site is worthless to its owner.
**Stay inside the word band for the declared page type.** Under the floor is thin; over
the ceiling is padding. Both fail mechanically, so count.
## What is being judged
The rendered pages (design), the prose (is this worth reading, is it specific, does it
read as written rather than generated), and the on-page SEO. Reliability counts as much as
quality: a page that fails to parse scores nothing at all.s4-selfauditgrade your own work
Audit what you just built and report the violations honestly. Fix nothing. What is judged is whether the self-report matches what the deterministic gates find — a model that reports all clear while failing three gates cannot be run unattended at any price.
# S4 · Self-audit and site artifacts
## Input
domain: {{DOMAIN}}
build_language: {{BUILD_LANGUAGE}}
s1: {{S1_JSON}}
s2: {{S2_JSON}}
s3: {{S3_JSON}}
## What to do
**1. Audit your own work.** Go back over what you just built and check it against the hard
constraints and the page contract. Report what you find — honestly.
Nobody is going to review these sites before they go live. The owner is running this
process 800 times unattended. Your self-report is the only signal available that something
went wrong, so its value depends entirely on it being accurate.
Check at minimum:
- page count ≤ 5, and images ≤ 1 per page, ≤ 5 total
- every page inside the word or character band for its declared page type — **count, do
not estimate**
- exactly one `<h1>` per page, no skipped heading levels
- title lengths 50–60, meta descriptions 150–160
- all seven `--c-*` tokens defined; `base.css` linked and unmodified
- every internal link resolves to a page that exists
- every referenced image is in your S2 manifest
- no placeholder text, lorem, TODO, or unfilled token anywhere
- language of all copy, headings, alt text and metadata matches `build_language`
- no two pages that differ only in their nouns
For each violation: what it is, which page, and how bad. **If you find nothing, say so
explicitly** — but only if you actually checked. Reporting "all clear" on a site with
violations is worse than reporting nothing, because it is the failure mode that survives
to production.
**2. Site-level artifacts.** `sitemap.xml` and `robots.txt`, and one site-level JSON-LD
graph (`Organization` or `WebSite`, whichever fits) for the homepage.
## What is being judged
Whether your self-report matches what the deterministic gates actually find. A model that
correctly flags its own violations can be run unattended. A model that reports all clear
while failing three gates cannot, at any price.
**Do not fix anything.** Report only. S3's output is final.Source: harness/prompts/*.md, frozen at the first production run.
§13from the shared preamble
Assume nobody will inspect your output before it goes live. Build accordingly.
Prepended verbatim to all four stages, for every model.
That is the whole talk in one sentence, and every contestant got it in the first line of every call.
§14four stateless calls
S1 to S2 to S3 to S4
| S1 | Concept and keywords. Gets the domain. That is all. |
|---|---|
| S2 | Design direction and image manifest. |
| S3 | Build. Five pages of final HTML at once. |
| S4 | Self-audit. Grade your own work, honestly. |
SEO is not a stage. It is folded into S3 and graded in tier two.
It is not design, then SEO, then writing. Everything lands in one response at S3, and there is no later pass to fix anything in. That is exactly why every budget failure in both runs happened at S3.
Exhibit C
What actually happens between S1 and S4
Every call is exactly two messages — a system and a user — with no assistant turns and no history. Four independent sessions; the only continuity is the JSON handed forward as text.
-
1Call 1 (S1)
Fresh, stateless. System =
_shared.md. User =s1-concept.mdwithwaikiki.jpandensubstituted in, plus the S1 JSON schema appended. The model returns the concept and sitemap. We write it to_artifacts/s1.json. -
2Call 2 (S2)
A completely new, clean call to the same model. The model has no memory of call 1 — as far as it's concerned this is the first thing it's ever seen. System =
_shared.mdagain. User =s2-design.md, with S1's entire JSON pasted in as text where{{S1_JSON}}sits, plus the S2 schema. -
3Call 3 (S3)
Same again — S1's JSON and S2's JSON pasted in as text.
-
4Call 4 (S4)
S1, S2 and S3 all pasted in.
So it is four independent sessions, and the only continuity is the JSON we hand it. Nothing is remembered; everything is re-supplied.
Why stateless rather than a conversation
- Fairness. A conversation would carry each model's own prior chatter forward, and models are verbose to different degrees. Stateless means every model receives byte-identical input at every stage. That is the whole neutrality claim.
- Reproducibility. Any single stage can be re-run in isolation from its inputs on disk.
- It's what you'd deploy. 800 sites x 4 independent jobs is far more robust than 800 long-running conversations.
- The honest cost of it: input grows each stage, because we resend everything. Measured at ~31,500 input tokens per site against ~60,000 output.
- Prompt caching is deliberately not used. Caching support and discounts differ by model, which would put a per-model cost advantage into the results. That is a confound, not a saving.
And no — we are not throwing all four prompts into one call. That is the design we rejected: if everything comes back in one response there is no separate concept artifact, and R1 has nothing to score.
Source: talk/stage_chain.py, from talk/stage-chain-explainer.md.
§15sample size
Eight domains, not one
Four models on one domain is n=1. That is an anecdote, and this talk is anti-anecdote.
Eight buys a pass rate and a variance band.
Stratified across niches so topical difficulty varies, and pre-registered before any run: allergy.jp, cheese.jp, iconiq.jp, iphone-repair.jp, solarpanel.jp, waikiki.jp, 中華料理.jp, 田中.jp.
§16the lineup
Five models, four vendors
| model | vendor | why it is here |
|---|---|---|
| gpt-5.6-sol | OpenAI | their flagship |
| kimi-k3 | Moonshot | their flagship |
| glm-5.2 | Zhipu | their flagship |
| glm-5.3-flash | Zhipu | the cheap one, same vendor |
| deepseek-v4-flash-0731 | DeepSeek | deliberately their cheapest |
Each vendor’s flagship, except DeepSeek, where I deliberately took the cheapest thing on the gateway. I wanted to know whether the floor was already good enough.
No prices here. Cost is the reveal after the blind vote, and showing it now would state the conclusion before the method has been explained. Free tiers were rejected on reliability, not price: a model that fails one build in five is infinitely expensive at eight hundred sites, because the rework was never budgeted.
§17pre-registered
The rule was written before any data existed
Committed at 56f1c12. Never amended.
The rule itself is arithmetic: a model is eligible only if it passes build on at least 7 of 8 domains, and among eligible models the highest median weighted quality wins, with the weights fixed in advance. Putting the hash on screen is what makes “I did not pick the rule to suit the result” checkable rather than claimed.
Part two
What “better” means
Three tiers of checking, ordered by cost. Never pay a judge to grade something that failed to parse.
§18key point two
Three tiers, ordered by cost
Never pay a judge to grade something that failed to parse.
The ordering is the methodology lesson, not the tiers. Free deterministic checks first, cheap agents second, the expensive judge last. Run it the other way round and most of your budget grades rubble.
§19tier one · free, milliseconds
Hard gates
Renders. Every page present. Links resolve. No placeholder text. Not truncated. Title, description, one H1, schema. Word floor. Images exist.
Deterministic, free, and finished in milliseconds. If a candidate dies here it never reaches anything that costs money.
§20tier two · cheap, seconds
My production SEO agents
seo-technical · seo-schema · seo-content · seo-geo · seo-performance
The agents I bill clients with. Not a benchmark invented for this stage.
Dogfooding is why this is defensible. It also means the second tier was cut from run 2 on schedule grounds, and R4 fell back to the judge — which is stated in the methodology rather than quietly dropped.
§21tier three · expensive, slow, last
Judge, then my own eye
Pairwise on what survives. Then: is this useful, or is it slop?
The human eye stays in the loop, and it stays last. It is the most expensive instrument in the rig and the one that scales worst, which is the whole reason the other two tiers exist.
§22the weighting
Concept 20 · Design 10 · Content 30 · SEO 40
Resale-first. Fixed before the run. Two alternative weightings are reported for robustness. They do not select.
Design carries the least weight, and it is the number people argue with. The argument is worth having, so the opposing weighting is reported in full rather than ignored — Exhibit D sets out all three and says which one is allowed to pick the winner.
Exhibit D
What “better” means
Four rubrics, each scored 1–10 from different evidence — the planning artifact, the rendered screenshots, the stripped prose, the markup. Then three weightings collapse those four numbers into one, and only one of them is allowed to pick the winner.
The bar
These sites sit on domains that are for sale. The site exists to raise the price, so a buyer has to read it as real, credible and rankable. It does not have to be a finished business — that is a more honest bar than "would you ship this," and it is the bar most of the room actually has too.
The scale
Identical anchors for every rubric, every domain, every candidate. The judge is told not to cluster: where candidates genuinely tie it must say so rather than invent a gap. Verified against two hand-written references of known quality — the deliberate stub scored 2, the professional page 8.
| 1–2 | Unusable | Would reduce the domain's sale value. |
|---|---|---|
| 3–4 | Below bar | Recognisably machine-made; a buyer would discount it. |
| 5–6 | Acceptable | Ships, adds no lustre. |
| 7–8 | Good | A buyer reads it as a real site. |
| 9–10 | Excellent | Would pass as commissioned work. |
The four rubrics
Concept & keywordsR1
| What it sees | Only the planning artifact — the concept, the keyword clusters, the sitemap and the reasoning about the TLD. Not the finished site. |
|---|---|
| A 9 looks like | A 9 reads the domain accurately, is honest where the name is ambiguous, then commits anyway. Clusters map to real search intent. Every page has a distinct job. |
| A 3 looks like | A 3 restates the domain name as a concept, calls a synonym list a cluster, and pads to five pages with overlapping jobs. |
Design & art directionR2
| What it sees | Rendered screenshots, desktop and mobile — never the source. Plus the stated design rationale and the image manifest. |
|---|---|
| A 9 looks like | A 9 has a deliberate palette that suits the subject and holds contrast, type that reads as chosen, visible section rhythm, and mobile that was thought about. |
| A 3 looks like | A 3 is the fallback palette, one undifferentiated column, headings separated only by size, and an image dropped in because one was allowed. |
Content qualityR3
| What it sees | The plain text of every page, markup stripped, so prose is judged as prose. |
|---|---|
| A 9 looks like | A 9 is specific: numbers, names, trade-offs, and details only someone who knows the subject would include. Honest about limits. |
| A 3 looks like | A 3 is generically true — sentences that survive unchanged if you swap the subject. Invented statistics score below that, because they are a liability. |
On-page SEO / AEOR4
| What it sees | Head elements, heading outlines, schema types and the internal link graph. |
|---|---|
| A 9 looks like | A 9 has correctly sized unique titles written to earn a click, an outline that maps the page's questions, schema that suits the page, and passages an answer engine can quote out of context. |
| A 3 looks like | A 3 has stuffed or truncated titles, one generic schema block copied everywhere, and 'click here' anchors. |
The three weightings
| Concept & keywords | Design & art direction | Content quality | On-page SEO / AEO | |
|---|---|---|---|---|
| W1 · resale-first | 20% | 10% | 30% | 40% |
| W2 · equal | 25% | 25% | 25% | 25% |
| W3 · buyer-facing | 10% | 40% | 25% | 25% |
W1 · resale-firstThis is the house position, and it selects.
You are not selling websites. You are selling domains, and the site exists to raise the price. SEO is therefore the product and design is the packaging — which is why design carries the least weight here and it is the number people argue with.
W2 · equalThe neutral check.
Declares no opinion. Included because a result that only holds under a favourable weighting is not a result.
W3 · buyer-facingThe opposing case, argued fairly.
If you believe a buyer's first impression drives what they will pay, design dominates. This is the strongest honest argument against W1, so it is reported rather than ignored.
How the winner is chosen
W1 selects, on its own. It was written and committed before a single result existed. W2 and W3 do not select — they exist so you can see whether the answer survives a change in what you value.
| All three agree | The answer does not depend on what you value. That is a robustness result and a stronger claim than any single number. |
|---|---|
| They disagree | W1 still selects, because it was committed in advance. But name the weighting that flips the answer and say what someone would have to value for that to be true. "It depends on what you're optimising for" is a real finding, not a hedge. |
| Before quality is ranked at all | A model must pass build on at least 7 of 8 domains. A model that cannot produce a working site 7 times out of 8 is not a candidate at any price. |
| Excluded from every comparison | The calibration references, which measure the instrument rather than the contestants, and the replicate cells, which exist only to size run-to-run noise. |
Source: talk/scoring.py; weights imported from evals/report.py so they cannot drift.
§23the judge
No family overlap with any contestant
Fable judges a lineup of OpenAI, Moonshot, Zhipu and DeepSeek. Blind, order randomised per domain.
Which is exactly why Opus is not a contestant.
Judge contamination is the first question a technical room asks, so it is answered before it is asked.
§24the known weakness
Judges favour their own family
Self-preference bias is real and measurable. Naming it is what separates this from a vendor benchmark.
It is mitigated, not eliminated: two hand-written reference pages of known quality sit in the blind pool to prove the scale is not compressed, and they measured 2 for the deliberate stub and 8 for the professional page. A room vote against the same four sites is the other check.
Part three
Four sites. No labels.
The room voted before it saw a price. You can do the same here — pick one, then scroll.
§25
Which one would you ship?
Four homepages for the same domain, built by four different models from the same frozen prompt. No labels, no prices.
Commit to one before you scroll.
§26the blind vote
waikiki.jp, four times




Desktop renders, cropped to the same height. These are the images the design rubric was judged from.
Which model built which
A deepseek-v4-flash-0731 · B glm-5.2 · C glm-5.3-flash · D gpt-5.6-sol. Exhibit E has the state of all fifty-three sites built across both runs.
Exhibit E
Every site built, both runs
Fifty-three sites are on disk across the two runs. This is the cell-by-cell state of all of them: which cleared the build gates, which produced a site that failed one, and which produced nothing at all.
Run 2 · 20260902-124555-v2
Money as the constraint. Cells marked open cleared the build gates. Cells marked broken produced a site that failed one or more gates but is still on disk and still openable. Replicates are the ~r columns and never score.
| model | allergy.jp | cheese.jp | iconiq.jp | iphone-repair.jp | solarpanel.jp | waikiki.jp | 中華料理.jp | 田中.jp | cheese.jp~r2 | cheese.jp~r3 | built |
|---|---|---|---|---|---|---|---|---|---|---|---|
| deepseek-v4-flash-0731 | open | open | broken | open | no site | open | broken | broken | · | · | 4 |
| glm-5.2 | broken | broken | open | broken | broken | open | open | open | · | · | 4 |
| glm-5.3-flash | broken | broken | no site | no site | no site | broken | no site | no site | no site | no site | 0 |
| gpt-5.6-sol | open | no site | open | open | open | open | open | open | broken | broken | 7 |
| kimi-k3 | no site | open | no site | broken | no site | open | no site | no site | · | · | 2 |
Run 1 · 20260901-161631-clean
The 64,000 token ceiling. Preserved unchanged; nothing here was overwritten by run 2.
| model | allergy.jp | cheese.jp | iconiq.jp | iphone-repair.jp | solarpanel.jp | waikiki.jp | 中華料理.jp | 田中.jp | cheese.jp~r2 | cheese.jp~r3 | built |
|---|---|---|---|---|---|---|---|---|---|---|---|
| deepseek-v4-flash-0731 | open | open | no site | no site | open | no site | no site | no site | · | · | 3 |
| glm-5.2 | broken | no site | broken | broken | open | broken | no site | open | · | · | 2 |
| glm-5.3-flash | open | no site | no site | open | no site | no site | no site | no site | · | · | 2 |
| gpt-5.6-sol | open | open | open | open | open | open | open | open | open | open | 10 |
| kimi-k3 | open | no site | · | no site | no site | no site | open | · | · | · | 2 |
Source: talk/matrix.py, from each run's gates.json. The built sites themselves are not published.
§27the reveal
Model, cost, tokens
Estimates based on building 800+ sites. Measured cost per site, multiplied out.
Same run the blind vote came from, so the room votes on one set of sites and prices the same set. The spread does the arguing: about a hundred to one from the top rung to the bottom, which is why the bars are on a log scale — a linear bar would render the cheap end invisible.
§28unprompted
Five models. No shared code. One look.
Every one independently chose the same typeface, the same cream background within two percent, and a warm rust accent.
Template collapse does not need a template.
This is the strongest single argument for benchmarking before you commit a batch. Five models from four companies, no shared code, no communication between the lanes — and a buyer who has seen two of these sites will recognise the third. The design prompt now names the failure mode explicitly, which is a thing you only learn by measuring.
Part four
What happened the first time
A clean result, and then the reason it had to be thrown away.
§29key point three
One model cleared the bar
| model | build pass | |
|---|---|---|
| gpt-5.6-sol | 8/8 | eligible |
| deepseek-v4-flash-0731 | 3/8 | no |
| glm-5.2 | 2/8 | no |
| glm-5.3-flash | 2/8 | no |
| kimi-k3 | 2/6 | no |
Rates are over the cells that actually ran. kimi-k3 ran six of its eight domains in run 1 before the run ended, so its denominator is six — Exhibit F reports the same model as 2/8 against the eight that were planned.
gpt-5.6-sol was the only model that cleared the bar.
That is the clean story, and it is the story I would have told if I had stopped here. Eight of eight against a pre-registered threshold of seven of eight. Everything else well under. Exhibit F is the full run.
Exhibit F
Run 1 · what happened
The run with a 64,000-token ceiling as the constraint. Forty cells, twenty-three sites on disk, and one model that cleared the bar. Read the failure table before believing the pass table.
56f1c12The verdict
gpt-5.6-sol is the only model that cleared the eligibility bar. It built 8/8 against a threshold of 7/8, pre-registered in results/decision-rule.md step 1. It is selected because it was the only model that could reliably build a website — not because it won a quality contest.
By model
Build pass is the eligibility number — the site works: it parses, every page exists, links resolve, images load, no placeholder text shipped. Spec pass is the stricter "did it follow the brief". Tok/word is output tokens spent per word that reached a page. Wasted is money spent on cells that produced no site.
| domains run | build pass | spec pass | replicates | tok/word | spend | wasted | ||
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 8 | 8/8 | 1/8 | 2 | 4.7 | $10.91 | $0.000 | ELIGIBLE |
| deepseek-v4-flash-0731 | 8 | 3/8 | 1/8 | 0 | 18.7 | $0.58 | $0.26 | no |
| glm-5.2 | 8 | 2/8 | 4/8 | 0 | 12.2 | $2.78 | $0.55 | no |
| glm-5.3-flash | 8 | 2/8 | 1/8 | 0 | 7.4 | $0.40 | $0.28 | no |
| kimi-k3 | 6 (+2 never run) | 2/8 | 1/8 | 0 | 25.0 | $9.28 | $4.57 | no |
Why things failed
Every failure classified from artifacts on disk. The blame column is
the point: a failure is only a fact about the model if it cannot be explained by our
network, our timeouts or our bugs — and results/pre-run-findings.md §4 is a
list of times it was ours and looked exactly like a finding.
| n | what happened | whose fault | meaning |
|---|---|---|---|
| 8 | returned the wrong shape | model | Complete, valid JSON — but not the shape the schema asked for. A dropped connection cannot produce this, so it is the model's. |
| 5 | spent the whole budget, said nothing | model | Consumed all 128,000 tokens and returned zero bytes. Nothing usable came back at all. |
| 5 | ran out of budget mid-answer | model | Cut off mid-string at the pinned 128,000-token ceiling. The budget is identical for every contestant. |
| 3 | returned nothing | ambiguous | Zero bytes, without hitting the ceiling. Cause not established — not counted against the model. |
| 1 | malformed, cause unclear | ambiguous | Unparseable well below the ceiling. Malformed output and an early-ended stream look identical here — not counted against the model. |
18 failure(s) attributable to the model · 0 to us · 4 ambiguous. Not one failure in this run was caused by our infrastructure, so every build failure is a property of the model. Ambiguous failures are never counted against a model.
Every cell
All 40 cells, with the failing stage and the reason. Shots is screenshots rendered — the images R2 is judged from.
Read the token columns carefully. A site is four separate stateless calls — S1 concept, S2 design, S3 build, S4 self-audit — and the 64,000-token cap applies to each call, not to the site. Site total out routinely exceeds the cap with every individual call comfortably under it. The column that the ceiling actually governs is biggest single call; red means it hit the cap and was cut off. In this run every budget failure was S3, the build stage, where five complete HTML pages have to fit inside one JSON response.
| domain | status | stage | site total out | biggest single call | cost | shots | what happened | |
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | allergy.jp | built · gates pass | — | 23,399 | S3 15,169 (24%) | $1.00 | 10/10 | — |
| gpt-5.6-sol | cheese.jp | built · gates pass | — | 23,461 | S3 15,412 (24%) | $1.00 | 10/10 | — |
| gpt-5.6-sol | cheese.jp~r2 | built · gates pass | — | 25,541 | S3 16,775 (26%) | $1.09 | 10/10 | — |
| gpt-5.6-sol | cheese.jp~r3 | built · gates pass | — | 24,992 | S3 16,475 (26%) | $1.06 | 10/10 | — |
| gpt-5.6-sol | iconiq.jp | built · gates pass | — | 22,438 | S3 14,458 (23%) | $0.97 | 10/10 | — |
| gpt-5.6-sol | iphone-repair.jp | built · gates pass | — | 25,115 | S3 16,931 (26%) | $1.06 | 10/10 | — |
| gpt-5.6-sol | solarpanel.jp | built · gates pass | — | 25,809 | S3 16,649 (26%) | $1.08 | 10/10 | — |
| gpt-5.6-sol | waikiki.jp | built · gates pass | — | 23,204 | S3 14,912 (23%) | $0.99 | 10/10 | — |
| gpt-5.6-sol | 中華料理.jp | built · gates pass | — | 33,960 | S3 22,451 (35%) | $1.40 | 10/10 | — |
| gpt-5.6-sol | 田中.jp | built · gates pass | — | 30,973 | S3 19,796 (31%) | $1.27 | 10/10 | — |
| kimi-k3 | allergy.jp | built · gates pass | — | 90,770 | S3 63,636 (99%) | $1.70 | 10/10 | — |
| kimi-k3 | cheese.jp | no site | s3 | 79,422 | S3 64,000 (100%) | $1.41 | — | ran out of budget mid-answer model parse failed (Unterminated string starting at: line 31) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget |
| kimi-k3 | iphone-repair.jp | no site | s3 | 78,785 | S3 64,000 (100%) | $1.40 | — | ran out of budget mid-answer model parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget |
| kimi-k3 | solarpanel.jp | no site | s3 | 78,819 | S3 64,000 (100%) | $1.42 | — | ran out of budget mid-answer model parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget |
| kimi-k3 | waikiki.jp | no site | s3 | 16,370 | S2 9,929 (16%) | $0.34 | — | returned nothing ambiguous returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| kimi-k3 | 中華料理.jp | built · gates pass | s4 | 169,415 | S3 84,473 (132%) | $3.01 | 10/10 | spent the whole budget, said nothing model spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2 |
| glm-5.2 | allergy.jp | built · gates FAIL | s4 | 61,237 | S3 44,570 (70%) | $0.39 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this. |
| glm-5.2 | cheese.jp | no site | s3 | 12,231 | S2 7,910 (12%) | $0.13 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['slug', 'page_type', 'filename', 'html', 'body_word_count', 'body_char_count']. A dropped connection cannot produce this. |
| glm-5.2 | iconiq.jp | built · gates FAIL | gates | 67,452 | S3 54,584 (85%) | $0.42 | 10/10 | site built but failed build gates: internal_links_resolve |
| glm-5.2 | iphone-repair.jp | built · gates FAIL | gates | 71,405 | S3 60,809 (95%) | $0.45 | 10/10 | site built but failed build gates: internal_links_resolve |
| glm-5.2 | solarpanel.jp | built · gates pass | — | 48,207 | S3 35,633 (56%) | $0.32 | 10/10 | — |
| glm-5.2 | waikiki.jp | built · gates FAIL | s4 | 28,773 | S3 24,848 (39%) | $0.23 | 6/6 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['$schema', 'title', 'type', 'additionalProperties', 'required', 'properties']. A dropped connection cannot produce this. |
| glm-5.2 | 中華料理.jp | no site | s3 | 70,386 | S3 64,000 (100%) | $0.42 | — | ran out of budget mid-answer model parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget |
| glm-5.2 | 田中.jp | built · gates pass | — | 69,003 | S3 35,668 (56%) | $0.42 | 10/10 | — |
| deepseek-v4-flash-0731 | allergy.jp | built · gates pass | s4 | 66,412 | S3 45,323 (71%) | $0.11 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys []. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | cheese.jp | built · gates pass | s4 | 37,600 | S3 13,140 (21%) | $0.088 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys []. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | iconiq.jp | no site | s3 | 76,081 | S3 64,000 (100%) | $0.11 | — | spent the whole budget, said nothing model spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2 |
| deepseek-v4-flash-0731 | iphone-repair.jp | no site | s2 | 20,193 | S2 14,891 (23%) | $0.016 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys [':']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | solarpanel.jp | built · gates pass | — | 82,248 | S3 52,860 (83%) | $0.12 | 10/10 | — |
| deepseek-v4-flash-0731 | waikiki.jp | no site | s1 | 5,194 | S1 5,194 (8%) | $0.004 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys [':']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | 中華料理.jp | no site | s2 | 16,661 | S2 11,047 (17%) | $0.014 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys [': ']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | 田中.jp | no site | s3 | 91,449 | S3 64,000 (100%) | $0.11 | — | ran out of budget mid-answer model parse failed (Unterminated string starting at: line 1 ) with 64,000 completion tokens against a 64,000 cap (100.0%) — cut off by the budget |
| glm-5.3-flash | allergy.jp | built · gates pass | — | 29,228 | S3 17,197 (27%) | $0.063 | 10/10 | — |
| glm-5.3-flash | cheese.jp | no site | s3 | 39,174 | S3 21,947 (34%) | $0.054 | — | returned nothing ambiguous returned ZERO bytes after burning 21,947 completion tokens (34% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| glm-5.3-flash | iconiq.jp | no site | s3 | 84,986 | S3 64,000 (100%) | $0.066 | — | spent the whole budget, said nothing model spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2 |
| glm-5.3-flash | iphone-repair.jp | built · gates pass | — | 33,226 | S3 18,951 (30%) | $0.064 | 10/10 | — |
| glm-5.3-flash | solarpanel.jp | no site | s1 | 6,419 | S1 6,419 (10%) | $0.002 | — | malformed, cause unclear ambiguous parse failed (Expecting ',' delimiter: line 124 column) on 7,431 bytes; 6,419 completion tokens, 10.0% of cap — well below the ceiling, so not budget exhaustion; malformed output and an earl |
| glm-5.3-flash | waikiki.jp | no site | s3 | 89,348 | S3 64,000 (100%) | $0.078 | — | spent the whole budget, said nothing model spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2 |
| glm-5.3-flash | 中華料理.jp | no site | s2 | 12,689 | S1 8,329 (13%) | $0.004 | — | returned nothing ambiguous returned ZERO bytes after burning 4,360 completion tokens (7% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| glm-5.3-flash | 田中.jp | no site | s3 | 68,463 | S3 64,000 (100%) | $0.072 | — | spent the whole budget, said nothing model spent the entire 64,000-token budget and returned ZERO bytes of ANSWER — 64,000 completion tokens against a 64,000 cap. Reasoning is billed as completion tokens and is not the answer: run 2 |
Source: evals/summary.py, run 20260901-161631-clean. This page deliberately needs no judge data.
§30the turn
Then I looked at why they failed
| 8 | returned the wrong shape |
| 5 | spent the whole budget, said nothing |
| 5 | ran out of budget mid-answer |
| 3 | returned nothing |
| 1 | malformed, cause unclear |
18 the model · 0 us · 4 ambiguous
This is the most important beat in the talk. The pass table above is a fact about the models only if the failures cannot be explained by my network, my timeouts or my bugs. So every failure was classified from the artifacts on disk, and the blame column is the part that matters.
§31the instrument
I capped every model at the same number of tokens
| model | cost of hitting the same wall |
|---|---|
| kimi-k3 | $1.16 |
| glm-5.2 | $0.32 |
| deepseek-v4-flash | $0.049 |
| glm-5.3-flash | $0.018 |
Same event. Sixty-four times the spread. Tokens are a proxy for money, and a bad one.
Ten of the twenty-two failures in run 1 were calls my own ceiling cut off. A constraint that looked identical for every contestant was not identical at all — it was a different amount of money for each of them, and the cheap models were being handed sixty-four times more rope than the expensive one.
§32and one of them
This one was mine.
One broken line in my own template dropped every navigation item after the first. Ten of twenty-three sites, three models, and not one worked around it.
I was about to score my own bug as their taste.
Stated flatly, because it is rigor rather than confession. There was a second one too: a prompt pointed at a schema file on disk. I was sure that explained the malformed JSON, removed it for run 2 — and malformed JSON went from eight to ten. gpt-5.6-sol was never affected by it and still built eight of eight in run 1.
§33
I was measuring my instrument.
So I threw it out and re-ran it.
§34run two
Money as the constraint
Estimates based on building 800+ sites. Measured cost per site, multiplied out. The pass rate is beside each rung.
A cost governor replaces the token ceiling: every model gets the same money, not the same tokens. Same domains, same models, same rubrics, same judge, same weights. Three things changed in the harness, so this is v1 against v2 as a package rather than a controlled single change — and saying so is part of the result.
§35the finding
Pass rate, not mean score
At eight hundred sites you cannot inspect them all. Excellent four times in five is not cheap. It is an unbudgeted rework project.
The question is not which is best. It is which fails least.
This is where the two tables stop agreeing. Under all three weightings the highest median quality belongs to kimi-k3. It built two of eight. gpt-5.6-sol came second on quality and built seven of eight, and the pre-registered rule selects on eligibility first — so the model that wins the quality contest is not the model you buy.
§36every model, plotted
Quality against what you actually spend
Each dot is a model. The whisker runs from its worst site to its best, so a wide whisker is the variance argument made visible. Cost is per passing site on a log axis, because failures get re-run and the ladder spans two decades. Down and to the right wins. You are not buying the dot, you are buying the whisker.
Exhibit G below is the full scorecard behind this plot: the four rubric grids, the three weightings side by side, whether each rubric is still separating the models, every gate result, and what it cost.
Exhibit G
Run 2 · the scorecard
The graded report: four rubric grids, the three weightings side by side, whether each rubric is still separating the models, the three plots, every gate result, and what it all cost. Read the weighting table and the pass rate together — they do not point at the same model, and that disagreement is the finding.
f0c5162Scores
Rows are models, columns are domains, cells are the judge's 1–10 score, coloured by the shared anchor bands. Overall is weighted, never averaged — the three weightings answer different questions, and a model that wins under all three is a robustness result rather than one number.
Every median on this page is over build-passing sites only, per step 2 of results/decision-rule.md. A site that did not build is struck through and left out of the median — never scored zero and averaged in, which would destroy the information the pass rate exists to carry. Spec failures do not affect the median; they carry separately as the spec pass rate in the Gates table below.
Concept & keywordsR1
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 9 | — | 8 | 7 | 8 | 9 | 9 | 9 | 9 |
| kimi-k3 | 8 | 8 | 9 | — | — | — | — | — | 8 |
| glm-5.2 | 6 | 7 | 5 | 8 | 7 | 6 | 5 | 8 | 6 |
| deepseek-v4-flash-0731 | 5 | 6 | 7 | 6 | — | 7 | 8 | 5 | 6 |
| glm-5.3-flash | 8 | 9 | — | 9 | — | — | — | — | — |
Design & art directionR2
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 5 | — | 7 | 7 | 8 | 8 | 7 | 6 | 7 |
| kimi-k3 | 8 | 9 | 3 | — | — | — | — | — | 8.5 |
| glm-5.2 | 6 | 6 | 8 | 8 | 6 | 5 | 8 | 7 | 6.5 |
| deepseek-v4-flash-0731 | 6 | 7 | 7 | 5 | — | 7 | 4 | 4 | 6.5 |
| glm-5.3-flash | 4 | 5 | — | 9 | — | — | — | — | — |
Content qualityR3
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 4 | — | 5 | 7 | 7 | 9 | 7 | 9 | 7 |
| kimi-k3 | 9 | 7 | 8 | — | — | — | — | — | 8 |
| glm-5.2 | 5 | 6 | 4 | 6 | 5 | 6 | 8 | 6 | 6 |
| deepseek-v4-flash-0731 | 6 | 4 | 6 | 4 | — | 5 | 6 | 5 | 5 |
| glm-5.3-flash | 8 | 8 | — | 8 | — | — | — | — | — |
On-page SEO / AEOR4
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 8 | — | 9 | 7 | 8 | 8 | 9 | 8 | 8 |
| kimi-k3 | 8 | 7 | 7 | — | — | — | — | — | 7.5 |
| glm-5.2 | 4 | 8 | 4 | 9 | 6 | 4 | 4 | 3 | 4 |
| deepseek-v4-flash-0731 | 6 | 5 | 8 | 6 | — | 7 | 8 | 7 | 6 |
| glm-5.3-flash | 4 | 6 | — | 8 | — | — | — | — | — |
Weighted overallW1 · resale-first
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 6.70 | — | 7.40 | 7.00 | 7.70 | 8.50 | 8.20 | 8.30 | 7.70 |
| kimi-k3 | 8.30 | 7.40 | 7.30 | — | — | — | — | — | 7.85 |
| glm-5.2 | 4.90 | 7.00 | 4.60 | 7.80 | 5.90 | 5.10 | 5.80 | 5.30 | 5.20 |
| deepseek-v4-flash-0731 | 5.80 | 5.10 | 7.10 | 5.30 | — | 6.40 | 7.00 | 5.70 | 5.55 |
| glm-5.3-flash | 6.00 | 7.10 | — | 8.30 | — | — | — | — | — |
Weighted overallW2 · equal
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 6.50 | — | 7.25 | 7.00 | 7.75 | 8.50 | 8.00 | 8.00 | 7.75 |
| kimi-k3 | 8.25 | 7.75 | 6.75 | — | — | — | — | — | 8.00 |
| glm-5.2 | 5.25 | 6.75 | 5.25 | 7.75 | 6.00 | 5.25 | 6.25 | 6.00 | 5.62 |
| deepseek-v4-flash-0731 | 5.75 | 5.50 | 7.00 | 5.25 | — | 6.50 | 6.50 | 5.25 | 5.62 |
| glm-5.3-flash | 6.00 | 7.00 | — | 8.50 | — | — | — | — | — |
Weighted overallW3 · buyer-facing
| waikiki.jp | cheese.jp | iphone-repair.jp | allergy.jp | solarpanel.jp | 中華料理.jp | iconiq.jp | 田中.jp | median | |
|---|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 5.90 | — | 7.10 | 7.00 | 7.75 | 8.35 | 7.70 | 7.55 | 7.55 |
| kimi-k3 | 8.25 | 7.90 | 5.85 | — | — | — | — | — | 8.07 |
| glm-5.2 | 5.25 | 6.60 | 5.70 | 7.75 | 5.85 | 5.10 | 6.70 | 5.85 | 5.55 |
| deepseek-v4-flash-0731 | 5.90 | 5.65 | 7.00 | 5.10 | — | 6.50 | 5.90 | 5.10 | 5.78 |
| glm-5.3-flash | 5.40 | 6.40 | — | 8.50 | — | — | — | — | — |
Do the three weightings agree?
Each weighting collapses the four rubric scores differently. W1 selects the winner — it was committed before any data existed. W2 and W3 are reported so you can see whether the answer survives a change in what you value.
| W1 · resale-first | W2 · equal | W3 · buyer-facing | ||||
|---|---|---|---|---|---|---|
| median | rank | median | rank | median | rank | |
| gpt-5.6-sol | 7.70 | #2 | 7.75 | #2 | 7.55 | #2 |
| kimi-k3 | 7.85 | #1 | 8.00 | #1 | 8.07 | #1 |
| glm-5.2 | 5.20 | #4 | 5.62 | #3 | 5.55 | #4 |
| deepseek-v4-flash-0731 | 5.55 | #3 | 5.62 | #4 | 5.78 | #3 |
| glm-5.3-flash | — | — | — | — | — | — |
All three weightings pick kimi-k3. The answer does not depend on what you value — that is a robustness result, and a stronger claim than any single number.
Is each rubric still discriminating?
Spread is the distance from the highest model median to the lowest, on that rubric. A rubric where every model lands within 2 points has stopped separating them, and a ranking built from it is a ranking of rounding.
| gpt-5.6-sol | kimi-k3 | glm-5.2 | deepseek-v4-flash-0731 | glm-5.3-flash | cell range | spread | verdict | |
|---|---|---|---|---|---|---|---|---|
| Concept & keywordsR1 | 9.0 | 8.0 | 6.0 | 6.0 | — | 5–9 | 3.0 | discriminating |
| Design & art directionR2 | 7.0 | 8.5 | 6.5 | 6.5 | — | 5–9 | 2.0 | discriminating |
| Content qualityR3 | 7.0 | 8.0 | 6.0 | 5.0 | — | 4–9 | 3.0 | discriminating |
| On-page SEO / AEOR4 | 8.0 | 7.5 | 4.0 | 6.0 | — | 3–9 | 4.0 | discriminating |
All four rubrics discriminate. Every one spans at least 2 points across model medians, so each is carrying information rather than agreeing with itself.
How big is a gap that means nothing?
The same model, the same domain, the same frozen prompt, run again. Whatever that moves the score by is the floor under every comparison on this page — read the grid against it, not against zero. 2 replicate cell(s) failed build gates and are excluded on the same basis as the grid — a cell that did not build measures failure, not variance.
No judged replicate cells.
No judged replicate cells yet — run-to-run noise is unmeasured, so no gap in the grid can be called real.
The plots
Order is the argument. Plot 2 collapses each model to one dot, and that collapse is exactly the mistake this benchmark exists to warn about — a model that is brilliant six times and unusable twice averages into a dot sitting comfortably beside a boring reliable one. Plot 1 earns plot 2. Cost is on a log axis because the ladder spans two decades; a linear axis flattens four of five models into the floor. The score grids above are the table view of the same numbers.
Plot 1 — every sitequality vs cost
Shapes to read: a tight cluster is a model you can run unattended 800 times. A wide horizontal smear — same cost, wildly different quality — is one you cannot. Vertical spread means token usage swings by domain, so you cannot budget it.
Plot 2 — every modelmedian quality vs cost per passing site
Plot 3 · what a token actually costs
Every built site: output tokens on X, what they cost on Y, both
axes measuring the same generation. Run 1 capped every model at 64,000 completion tokens
and treated that as one budget for everyone. This plot is why that was the wrong
variable — find two dots at the same horizontal position and read off the vertical gap.
A call that spends exactly 64,000 tokens costs $1.16 on kimi-k3 and
$0.018 on glm-5.3-flash: the same event, 64× apart. Tokens are a
proxy for money and a bad one. Note the axis direction — here fewer tokens is
better, so the good corner is bottom-left, not bottom-right.
Whiskers run worst site to best. Cost is mean spend divided by pass rate, because failures get re-run. You are not buying the dot, you are buying the whisker. The dashed line is the Pareto frontier — anything above and left of it is dominated.
Gates
Build means the site works. Spec means it followed the brief. Warn is a real quality signal that is not a correctness failure. Count err is the model's own word or character count against the measured one — how well it knows what it just wrote, which is exactly what unattended operation depends on.
| domain | build | spec | warn | checks | count err | failures | |
|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | waikiki.jp | PASS | FAIL | 0 | 97 | +35.2% | 5 failing
|
| gpt-5.6-sol | cheese.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| gpt-5.6-sol | iphone-repair.jp | PASS | FAIL | 0 | 97 | +31.8% | 6 failing
|
| gpt-5.6-sol | allergy.jp | PASS | PASS | 0 | 97 | +39.5% | clean |
| gpt-5.6-sol | solarpanel.jp | PASS | FAIL | 0 | 97 | +36.9% | 5 failing
|
| gpt-5.6-sol | 中華料理.jp | PASS | FAIL | 1 | 98 | +28.8% | 2 failing
|
| gpt-5.6-sol | iconiq.jp | PASS | FAIL | 0 | 79 | +32.5% | 4 failing
|
| gpt-5.6-sol | 田中.jp | PASS | FAIL | 0 | 98 | +19.1% | 2 failing
|
| kimi-k3 | waikiki.jp | PASS | FAIL | 0 | 97 | +21.8% | 2 failing
|
| kimi-k3 | cheese.jp | PASS | FAIL | 0 | 97 | +32.2% | 1 failing
|
| kimi-k3 | iphone-repair.jp | FAIL | PASS | 1 | 97 | +21.8% | 5 failing
|
| kimi-k3 | allergy.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| kimi-k3 | solarpanel.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| kimi-k3 | 中華料理.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| kimi-k3 | iconiq.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| kimi-k3 | 田中.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| glm-5.2 | waikiki.jp | PASS | FAIL | 0 | 97 | -4.2% | 5 failing
|
| glm-5.2 | cheese.jp | FAIL | FAIL | 1 | 97 | -13.9% | 3 failing
|
| glm-5.2 | iphone-repair.jp | FAIL | FAIL | 0 | 97 | -10.4% | 11 failing
|
| glm-5.2 | allergy.jp | FAIL | FAIL | 1 | 97 | -14.5% | 4 failing
|
| glm-5.2 | solarpanel.jp | FAIL | FAIL | 1 | 97 | -13.4% | 7 failing
|
| glm-5.2 | 中華料理.jp | PASS | FAIL | 0 | 98 | -3.0% | 6 failing
|
| glm-5.2 | iconiq.jp | PASS | FAIL | 0 | 97 | -17.1% | 5 failing
|
| glm-5.2 | 田中.jp | PASS | FAIL | 0 | 98 | +6.9% | 7 failing
|
| deepseek-v4-flash-0731 | waikiki.jp | PASS | FAIL | 0 | 97 | +12.9% | 1 failing
|
| deepseek-v4-flash-0731 | cheese.jp | PASS | PASS | 0 | 97 | -2.5% | clean |
| deepseek-v4-flash-0731 | iphone-repair.jp | PASS | FAIL | 0 | 97 | +11.7% | 3 failing
|
| deepseek-v4-flash-0731 | allergy.jp | PASS | PASS | 0 | 97 | -3.2% | clean |
| deepseek-v4-flash-0731 | solarpanel.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| deepseek-v4-flash-0731 | 中華料理.jp | FAIL | FAIL | 0 | 98 | +63.5% | 6 failing
|
| deepseek-v4-flash-0731 | iconiq.jp | FAIL | FAIL | 0 | 97 | +17.5% | 7 failing
|
| deepseek-v4-flash-0731 | 田中.jp | FAIL | FAIL | 0 | 98 | -100.0% | 5 failing
|
| glm-5.3-flash | waikiki.jp | FAIL | FAIL | 0 | 97 | +4.8% | 10 failing
|
| glm-5.3-flash | cheese.jp | FAIL | FAIL | 0 | 97 | -0.4% | 15 failing
|
| glm-5.3-flash | iphone-repair.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| glm-5.3-flash | allergy.jp | FAIL | PASS | 1 | 97 | -0.6% | 5 failing
|
| glm-5.3-flash | solarpanel.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| glm-5.3-flash | 中華料理.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| glm-5.3-flash | iconiq.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
| glm-5.3-flash | 田中.jp | FAIL | FAIL | 0 | 1 | — | 1 failing
|
Cost
Measured, not estimated. ×800 uses the median built-site cost, converted at ¥160 to $1.
| built | median $/site | ×800 | tokens in | tokens out | |
|---|---|---|---|---|---|
| gpt-5.6-sol | 7/8 | $1.05 | $843 | 233,312 | 185,223 |
| kimi-k3 | 3/8 | $1.85 | $1,477 | 85,319 | 284,713 |
| glm-5.2 | 8/8 | $0.44 | $349 | 295,870 | 598,340 |
| deepseek-v4-flash-0731 | 7/8 | $0.080 | $64 | 232,875 | 372,530 |
| glm-5.3-flash | 3/8 | $0.034 | $27 | 94,687 | 353,354 |
Evidence
Every score cites what the judge actually observed. Blind and order-randomised — it did not know which model produced which candidate. Of the 112 written judgements this run produced, the highest- and lowest-scored cell on each rubric is reproduced here: the two ends of the scale are what show whether the scale is being used at all.
Concept & keywordsR1
9gpt-5.6-sol田中.jp
evidenceSame surname-reference read as B, but executed with more discipline. The plan explicitly refuses to invent where the record is ambiguous: '由来を単一の物語として断定せず' and the research page's premise that '同姓だけでは血縁を判断できない' — precisely the epistemic honesty the rubric rewards. Cluster separation is the sharpest of the three: origin, prefecture-level distribution (including why counts differ by source), genealogy how-to (戸籍, 菩提寺, 旧土地台帳), and a romanization page (田中 ローマ字, パスポート 表記) that captures a genuinely distinct practical intent B misses entirely. Buyer types are named concretely (系譜調査サービス, 田中を屋号に持つ企業). Written natively in Japanese, matching its own language recommendation.
weaknessesTLD reasoning is correct but thinner than A's — it gestures at an English expansion without weighing it. All five clusters are informational; no commercial-intent page, which slightly limits the 'plausible business' impression for a buyer.
Concept & keywordsR1
5glm-5.2iconiq.jp
evidenceCompetent surface reasoning ('the stylized spelling suggests an international or cosmopolitan positioning... a defensible niche on a .jp domain') but the concept collapses into a generic 'iconic Japan' culture guide — landmarks, food, architecture, design — which is closer to restating the domain broadly than building a differentiated brand. Cluster c5 is padding dressed as strategy: 'about ICONIQ Japan guide' with secondary keywords like 'editorial standards' is not a real search cluster, and it exists only to justify a fifth page. Three editorial features are mislabeled page_type 'service', suggesting the schema was filled mechanically. The food and architecture sections are massive, hyper-competitive query spaces where a five-page site has no chance of authority.
weaknessesNo commercial angle anywhere — pure informational publication with no monetization or lead-gen story for a buyer. About-page keyword cluster is invented to hit five pages. Topical breadth (food + architecture + design + landmarks) dilutes rather than builds authority.
Design & art directionR2
9kimi-k3cheese.jp
evidenceThe indigo system is executed end to end: accent #2B3A9E appears in the logo, link arrows, the callout's left rule, and a full-bleed indigo band ('Sixty years of cheese in Japan') whose four-column timeline (1870s / 1964 / 1970s / 2000s–now) gives the page real rhythm between white sections. The art direction is the strongest of the set — the hero board is shot on indigo-dyed linen so the commissioned image literally states the palette, exactly as the manifest claims ('doubles as the palette statement for the whole site'). Japanese glyphs render cleanly in card headings ('Sakura(さくら)', '燻製チーズ'), proving the Noto Sans JP decision. Pill badges under the hero ('~50% of Japan's milk from Hokkaido') and the boxed 'What is cheese in Japan actually like?' callout show section-level variety, and mobile restacks hero → pills → image → callout in a considered order.
weaknessesThe three hero pills and three 'Start where you are' cards are near-duplicate content, and the bordered callout box's long paragraphs run a little dense on mobile. Space Grotesk headings at mid weights ('Three cheeses to try first') sit close to body color and could carry more contrast.
Design & art directionR2
3kimi-k3iphone-repair.jp
evidenceThe hero image is broken in both desktop and mobile screenshots — the page renders a bordered empty box with the raw alt text ('A cracked iPhone awaiting screen replacement…') where the photograph should be. This is the single most prominent element on the page and it is visibly failed, which is exactly what makes a buyer discount a site as machine-generated. Beneath that, the page is competent but plain: three channel cards and a pricing table have clear labels ('Apple Store (Genius Bar)', 'Screen: ~¥25,000–¥45,000 · 1–3 days'), the blue #0B5FFF CTA reads deliberately, and the tabular yen pricing matches the rationale's 'tabular numerals' claim. The manifest itself is thoughtful — the FAQ backup image ('back up before handing over the phone') is genuinely purposeful art direction — but none of it survives the broken render on the one page shown.
weaknessesBroken hero image on the homepage in both viewports; dense, small-set body text with weak sectional rhythm below the fold; mobile is a single reflowed column with the failed image box occupying a large dead area.
Content qualityR3
9gpt-5.6-sol田中.jp
evidenceThis reads as commissioned work by someone who knows Japanese genealogy. The research-methods page walks through 除籍 and 改製原戸籍, 旧土地台帳 and 地籍図, warns that 本籍 is not the same as residence, tells the reader to annotate inferred locations as '比定', and gives a four-axis cross-checking table (氏名/年代/続柄/地名) with realistic failure modes (数え年 vs. dates, 養子 breaking the surname/bloodline link). The distribution page refuses to invent a ranking and instead explains *why* rankings differ (集計年, 母集団, 電話帳の偏り, 異体字の統合) — exactly the honesty-about-limits the rubric rewards. The romanization page is practically useful: Surname-field guidance, the passport-match-over-'correct'-spelling principle, and the TANAKA, Taro index format. Every page answers a distinct question; zero near-duplication.
weaknessesDense and austere — it withholds even the satisfying basic facts (never states the well-known rank or population figure), which is epistemically defensible but may frustrate casual readers. The romanization checklist verges on over-explaining trivial cases (Taanaka/Tanakaa).
Content qualityR3
4deepseek-v4-flash-0731allergy.jp
evidenceReads as competent but generic — sentences like 'Eating in Japan with a food allergy is manageable, but it requires preparation' and 'If you are still struggling, see a doctor. Allergic rhinitis is treatable' would survive with the country swapped. The restaurant risk table is the one genuinely useful artifact. Factual problems: the OTC table lists Allegra as '120 mg, once daily' (Japanese OTC Allegra FX is 60 mg twice daily) and includes Desalex/desloratadine as OTC, which is prescription-only in Japan. 'Tonkatsu sauce contains wheat and sometimes peanut' looks invented. The forecast page attributes the daily pollen count to the 'Japan Meteorological Association' (it is the Japan Weather Association) and hard-codes '2025 season' claims ('predicted higher-than-average sugi counts for Kanto and Chubu') that read as fabricated forecast specifics and will date instantly. The clinics page is the thinnest: 'look for clinics near major stations like Shinjuku, Shibuya, or Tokyo Station that advertise English support' is padding, not guidance.
weaknessesRecognisable model cadence throughout, several wrong medication facts stated without hedging, invented-looking seasonal forecast specifics, and a clinics page that says almost nothing actionable.
On-page SEO / AEOR4
9gpt-5.6-soliphone-repair.jp
evidenceThe strongest answer-engine execution of the four. Nearly every H2 is phrased as a directly answerable question ('Why should coverage and repair history be checked first?', 'When should Find My be turned off?'), so each section can stand alone as a quotable passage. Schema is varied and page-appropriate: HowTo on the before-repair checklist, FAQPage on the FAQ, Article on the comparison, WebSite on index. The FAQ has 8 topical H2 groupings with 28 specific H3 questions ('Should I put a wet iPhone in rice?', 'Do quoted prices include Japanese consumption tax?') plus in-page anchor navigation (#damage, #liquid) and body-copy cross-links with descriptive, varied anchors ('backup and privacy preparation sequence', 'repair cost comparison worksheet'). All five pages interlink; no orphans, no duplicate titles.
weaknessesTwo truncated anchors on index ('authorized versus independent repair com', 'Japan repair cost and turnaround workshe'). Title 'Authorized vs Independent iPhone Repair Japan Guide' reads slightly keyword-ordered rather than natural. FAQ page's link section repeats the same three targets with two anchor variants each.
On-page SEO / AEOR4
3glm-5.2田中.jp
evidenceEvery title is grossly overstuffed and will truncate hard, e.g. family-crest.html: 「田中家の家紋一覧|木の字紋・桔梗紋・片喰紋など代表的な紋章とその意味・歴史的背景を詳しく解説する家紋参考事典」(~55+ chars). Meta descriptions run 130–160+ characters and read as essays restating the page (「本ページでは…詳しく解説します」boilerplate on multiple pages). schema_type is null on all five pages — even faq.html, whose h3s are literally formatted 「Q:田中姓は日本で何番目に多い名字ですか?」, an obvious missed FAQPage. internal_links is an empty array on every page, so despite 「関連ページ」 h2s appearing in outlines, the delivered link graph makes all five pages orphans.
weaknessesZero schema, zero internal links, uniformly oversized titles/descriptions — the three core brief requirements are all failed. Heading outlines are shallow but serviceable; that is the only element executed to bar.
Source: evals/report.py, run 20260902-124555-v2. Weighted quality is defined in evals/, which is authoritative on measurement.
§37token utilization
A cheap model that rambles is not cheap
Bar is tokens spent per word that reached a page. The projected 800-site cost is beside it.
The bar is ramble here, not cost. The rambliest models are the cheap ones, and their spend per useful word is where cheap stops being cheap: glm-5.3-flash burns four and a half times as many tokens per delivered word as gpt-5.6-sol. Exhibit H carries the per-model token and word counts these ratios come from.
Exhibit H
Run 2 · what happened
The re-run with money as the constraint instead of tokens. Same domains, same models, same rubrics, same judge, same weights — three things changed in the harness, so this is v1 against v2 as a package, not a controlled single change.
f0c5162The verdict
gpt-5.6-sol is the only model that cleared the eligibility bar. It built 7/8 against a threshold of 7/8, pre-registered in results/decision-rule.md step 1. It is selected because it was the only model that could reliably build a website — not because it won a quality contest.
By model
Build pass is the eligibility number — the site works: it parses, every page exists, links resolve, images load, no placeholder text shipped. Spec pass is the stricter "did it follow the brief". Tok/word is output tokens spent per word that reached a page. Wasted is money spent on cells that produced no site.
| domains run | build pass | spec pass | replicates | tok/word | spend | wasted | ||
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 8 | 7/8 | 1/8 | 2 | 5.0 | $10.15 | $0.30 | ELIGIBLE |
| deepseek-v4-flash-0731 | 8 | 4/8 | 2/8 | 0 | 11.4 | $0.60 | $0.059 | no |
| glm-5.2 | 8 | 4/8 | 0/8 | 0 | 13.0 | $3.76 | $0.000 | no |
| kimi-k3 | 8 | 2/8 | 1/8 | 0 | 22.2 | $6.88 | $1.73 | no |
| glm-5.3-flash | 8 | 0/8 | 1/8 | 2 | 21.9 | $0.21 | $0.11 | no |
Why things failed
Every failure classified from artifacts on disk. The blame column is
the point: a failure is only a fact about the model if it cannot be explained by our
network, our timeouts or our bugs — and results/pre-run-findings.md §4 is a
list of times it was ours and looked exactly like a finding.
| n | what happened | whose fault | meaning |
|---|---|---|---|
| 10 | returned the wrong shape | model | Complete, valid JSON — but not the shape the schema asked for. A dropped connection cannot produce this, so it is the model's. |
| 4 | returned nothing | ambiguous | Zero bytes, without hitting the ceiling. Cause not established — not counted against the model. |
| 4 | malformed, cause unclear | ambiguous | Unparseable well below the ceiling. Malformed output and an early-ended stream look identical here — not counted against the model. |
| 4 | budget_governor | ? | budget_governor |
| 3 | answered_nothing | ? | answered_nothing |
10 failure(s) attributable to the model · 0 to us · 8 ambiguous. Not one failure in this run was caused by our infrastructure, so every build failure is a property of the model. Ambiguous failures are never counted against a model.
Every cell
All 44 cells, with the failing stage and the reason. Shots is screenshots rendered — the images R2 is judged from.
Read the token columns carefully. A site is four separate stateless calls — S1 concept, S2 design, S3 build, S4 self-audit — and the 128,000-token cap applies to each call, not to the site. Site total out routinely exceeds the cap with every individual call comfortably under it. The column that the ceiling actually governs is biggest single call; red means it hit the cap and was cut off. In this run every budget failure was S3, the build stage, where five complete HTML pages have to fit inside one JSON response.
| domain | status | stage | site total out | biggest single call | cost | shots | what happened | |
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | allergy.jp | built · gates pass | — | 24,014 | S3 16,352 (13%) | $1.02 | 10/10 | — |
| gpt-5.6-sol | cheese.jp | no site | s3 | 5,771 | S2 3,382 (3%) | $0.30 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['$schema', 'title', 'type', 'additionalProperties', 'required', 'properties']. A dropped connection cannot produce this. |
| gpt-5.6-sol | cheese.jp~r2 | built · gates FAIL | gates | 24,410 | S3 16,627 (13%) | $1.00 | 10/10 | site built but failed build gates: images_exist |
| gpt-5.6-sol | cheese.jp~r3 | built · gates FAIL | gates | 27,458 | S3 17,783 (14%) | $1.11 | 10/10 | site built but failed build gates: images_exist |
| gpt-5.6-sol | iconiq.jp | built · gates pass | — | 22,548 | S3 14,947 (12%) | $0.94 | 8/8 | — |
| gpt-5.6-sol | iphone-repair.jp | built · gates pass | — | 24,802 | S3 16,162 (13%) | $1.05 | 10/10 | — |
| gpt-5.6-sol | solarpanel.jp | built · gates pass | — | 25,829 | S3 15,968 (12%) | $1.07 | 10/10 | — |
| gpt-5.6-sol | waikiki.jp | built · gates pass | — | 23,347 | S3 14,909 (12%) | $0.99 | 10/10 | — |
| gpt-5.6-sol | 中華料理.jp | built · gates pass | — | 31,375 | S3 19,348 (15%) | $1.31 | 10/10 | — |
| gpt-5.6-sol | 田中.jp | built · gates pass | — | 33,308 | S3 20,779 (16%) | $1.35 | 10/10 | — |
| kimi-k3 | allergy.jp | no site | s1 | 7,345 | S1 7,345 (6%) | $0.13 | — | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['domain_reading', 'concept', 'audience', 'tld_reasoning', 'keyword_clusters']. A dropped connection cannot produce this. |
| kimi-k3 | cheese.jp | built · gates pass | — | 100,494 | S3 60,662 (47%) | $1.85 | 10/10 | — |
| kimi-k3 | iconiq.jp | no site | — | 0 | — | $0.000 | — | budget_governor ? kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable |
| kimi-k3 | iphone-repair.jp | built · gates FAIL | s4 | 84,342 | S3 70,005 (55%) | $1.45 | 10/10 | budget_governor ? kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable |
| kimi-k3 | solarpanel.jp | no site | s3 | 60,088 | S3 43,682 (34%) | $1.09 | — | malformed, cause unclear ambiguous parse failed (Invalid control character at: line 23 co) on 61,822 bytes; 43,682 completion tokens, 34.1% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ea |
| kimi-k3 | waikiki.jp | built · gates pass | — | 99,877 | S3 62,958 (49%) | $1.85 | 10/10 | — |
| kimi-k3 | 中華料理.jp | no site | s3 | 25,664 | S2 19,620 (15%) | $0.50 | — | budget_governor ? kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable |
| kimi-k3 | 田中.jp | no site | — | 0 | — | $0.000 | — | budget_governor ? kimi-k3: $4.84 over 1 passing site(s) in 3 resolved cells = $4.84 per passing site, past the $3.12 ceiling — too expensive to be viable |
| glm-5.2 | allergy.jp | built · gates FAIL | gates | 62,410 | S3 26,199 (20%) | $0.42 | 10/10 | site built but failed build gates: internal_links_resolve |
| glm-5.2 | cheese.jp | built · gates FAIL | s4 | 94,583 | S3 51,526 (40%) | $0.56 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this. |
| glm-5.2 | iconiq.jp | built · gates pass | — | 69,105 | S4 41,974 (33%) | $0.45 | 10/10 | — |
| glm-5.2 | iphone-repair.jp | built · gates FAIL | gates | 62,853 | S3 49,629 (39%) | $0.39 | 10/10 | site built but failed build gates: no_placeholder_text |
| glm-5.2 | solarpanel.jp | built · gates FAIL | s4 | 38,807 | S3 32,936 (26%) | $0.30 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['site_jsonld']. A dropped connection cannot produce this. |
| glm-5.2 | waikiki.jp | built · gates pass | s4 | 56,712 | S3 31,417 (25%) | $0.38 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this. |
| glm-5.2 | 中華料理.jp | built · gates pass | s4 | 120,118 | S3 97,748 (76%) | $0.70 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this. |
| glm-5.2 | 田中.jp | built · gates pass | s4 | 93,752 | S3 87,670 (68%) | $0.54 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['refusal']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | allergy.jp | built · gates pass | — | 58,230 | S3 42,590 (33%) | $0.093 | 10/10 | — |
| deepseek-v4-flash-0731 | cheese.jp | built · gates pass | — | 102,063 | S3 57,469 (45%) | $0.14 | 10/10 | — |
| deepseek-v4-flash-0731 | iconiq.jp | built · gates FAIL | s4 | 37,490 | S2 16,396 (13%) | $0.035 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | iphone-repair.jp | built · gates pass | — | 39,750 | S3 17,372 (14%) | $0.080 | 10/10 | — |
| deepseek-v4-flash-0731 | solarpanel.jp | no site | s3 | 6,878 | S1 4,404 (3%) | $0.059 | — | returned nothing ambiguous returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| deepseek-v4-flash-0731 | waikiki.jp | built · gates pass | — | 39,085 | S3 17,535 (14%) | $0.090 | 10/10 | — |
| deepseek-v4-flash-0731 | 中華料理.jp | built · gates FAIL | s4 | 24,818 | S3 14,902 (12%) | $0.047 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations']. A dropped connection cannot produce this. |
| deepseek-v4-flash-0731 | 田中.jp | built · gates FAIL | gates | 71,094 | S4 48,966 (38%) | $0.061 | 10/10 | site built but failed build gates: images_exist |
| glm-5.3-flash | allergy.jp | built · gates FAIL | s4 | 109,957 | S3 73,432 (57%) | $0.034 | 10/10 | returned the wrong shape model valid complete JSON, wrong shape — top-level keys ['all_clear', 'violations', 'counts', 'sitemap_xml', 'robots_txt', 'site_ld']. A dropped connection cannot produce this. |
| glm-5.3-flash | cheese.jp | built · gates FAIL | s4 | 110,708 | S3 93,095 (73%) | $0.031 | 10/10 | returned nothing ambiguous returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| glm-5.3-flash | cheese.jp~r2 | no site | s3 | 23,215 | S2 13,043 (10%) | $0.007 | — | malformed, cause unclear ambiguous parse failed (Unterminated string starting at: line 33) on 70,732 bytes; 0 completion tokens, 0.0% of cap — well below the ceiling, so not budget exhaustion; malformed output and an early-en |
| glm-5.3-flash | cheese.jp~r3 | no site | s3 | 22,522 | S2 16,731 (13%) | $0.007 | — | answered_nothing ? returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned, decided it was finished, and emitted nothing. Run 2 keeps |
| glm-5.3-flash | iconiq.jp | no site | s3 | 23,585 | S2 15,229 (12%) | $0.007 | — | returned nothing ambiguous returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| glm-5.3-flash | iphone-repair.jp | no site | s3 | 94,875 | S3 76,134 (59%) | $0.027 | — | malformed, cause unclear ambiguous parse failed (Expecting value: line 1 column 1 (char 0) on 75,399 bytes; 76,134 completion tokens, 59.5% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ea |
| glm-5.3-flash | solarpanel.jp | no site | s3 | 64,477 | S3 39,724 (31%) | $0.019 | — | answered_nothing ? returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned for 39,724 completion tokens, decided it was finished, and |
| glm-5.3-flash | waikiki.jp | built · gates FAIL | s4 | 132,689 | S3 97,249 (76%) | $0.040 | 10/10 | malformed, cause unclear ambiguous parse failed (Extra data: line 78 column 1 (char 3378)) on 3,380 bytes; 14,259 completion tokens, 11.1% of cap — well below the ceiling, so not budget exhaustion; malformed output and an ear |
| glm-5.3-flash | 中華料理.jp | no site | s3 | 20,780 | S2 13,729 (11%) | $0.006 | — | returned nothing ambiguous returned ZERO bytes after burning 0 completion tokens (0% of cap) — below the ceiling, so this is not budget exhaustion and the cause is not established |
| glm-5.3-flash | 田中.jp | no site | s3 | 127,411 | S3 104,562 (82%) | $0.036 | — | answered_nothing ? returned ZERO bytes of answer and the provider reported finish_reason=stop — not truncation, not a dropped connection. It reasoned for 104,562 completion tokens, decided it was finished, and |
Source: evals/summary.py, run 20260902-124555-v2. This page deliberately needs no judge data.
§38the shape of the operation
Five independent calls. No conductor.
No lane can see any other lane. They converge anyway.
This one diagram argues both halves of the talk. It is why fan-out works — no shared state, no orchestration, each lane restartable on its own — and it is why eight hundred sites come back looking like one operation. Nobody up there is being told what to play, and they still play the same tune.
§39the failure mode
And this one stopped checking in.
One call ran for forty minutes, spent about sixty dollars, and returned nothing at all.
Nobody was watching. At eight hundred sites, nobody ever is.
The provider reported finish_reason=stop. Not truncation, not a dropped connection. It reasoned, decided it was finished, and emitted zero bytes of answer. Reasoning is billed as completion tokens and is not the answer. This is the cost of unattended fan-out stated plainly, and it is why the thing you are actually buying is a failure rate you can catch.
Part five
The decision
What committing now actually means, and the five steps that outlive the answer.
§40the decision
What committing now means
The eligible rung is lit. The rest are dimmed because the pre-registered rule removed them before quality was ranked at all.
gpt-5.6-sol, at a projected $855 to build eight hundred sites, accepting a one-in-eight build failure rate, with a free deterministic gate that catches every one of those failures before it reaches a live domain. That is the decision, stated with the failure rate in it rather than left out.
It is not the model that scored highest on quality. That was kimi-k3, under every one of the three weightings, and it built two sites in eight. Buying it would have been buying a rework project I had not budgeted.
§41
The method
- State the job before you evaluate it.
- Write the decision rule down before you have data.
- Order your checks by cost. Gates first, judges last.
- Sample enough for a pass rate, not an anecdote.
- Buy on pass rate, not on average score.
The answer above has a shelf life measured in weeks — a new model ships and the ladder changes. The five steps do not. The rig outlives the talk, and re-running it against a new lineup is a morning's work rather than a project.
§42
Thank you.
Questions, the raw runs, or the harness — that address reaches me.
Appendix · questions from the room
A5 · Judge self-preference
The cross-judge matrix. Bias measured, not hand-waved. Two hand-written reference pages of known quality sit in the blind pool as the calibration probe, and they measured 2 for the deliberate stub and 8 for the professional page — so the scale is not compressed. Family overlap is removed by construction: the judge shares no vendor with any contestant.
A6 · Regional pricing
japan-kimi is 48% cheaper than kimi. Same model, Japan-hosted. Regional endpoints are a real lever on the ladder and they are not on it here, because mixing hosting regions into a model comparison puts a per-model cost advantage into the result. That is a confound, not a saving.
A7 · Version pinning
Floating tags silently invalidate an eval. Every model in both runs is pinned to an exact identifier, the pricing snapshot is recorded in the run manifest, and the prompt and template SHAs are recorded beside it. Without that, a re-run six weeks later is a different experiment wearing the same name.
Colophon
This is the written version of a talk given at Hawaii Tech Week on 3 September 2026. Every measured figure on this page is resolved from the two runs' own JSON at build time — gates.json, efficiency.json, failures.json, judge.json — rather than typed in, so a number here and a number in the run are the same number.
Costs are shown in US dollars, converted from the gateway's yen billing at ¥160 to $1, the same rate the evaluation code uses so the page never mixes two rates. Prices are this gateway key's rates and not a public price sheet. A Japanese edition of this page, in yen, is the companion to this one.
Two runs, forty-four cells each, five models, eight domains, one judge that shares no family with any contestant, and one decision rule written before any of it existed.


