Experiment 02: does structure earn AI citations? The pre-registration
A pre-registered A/B test: does the GEO formatting bundle make a page likelier to be cited by ChatGPT, Perplexity, Claude, and Gemini than identical plain prose?
Everyone in AI SEO tells you to write “answer-first”: lead with a direct answer, break the page into self-contained sections, chunk one claim per paragraph, add a comparison table, bolt on an FAQ with FAQPage schema. The claim is that this GEO formatting bundle makes a page more likely to be cited by AI engines. I have not seen it isolated from everything else that moves citations — better facts, more sources, longer copy, an older domain. So I’m testing it, and I’m publishing the design before I know the answer.
This post is the pre-registration. It states the hypothesis, the single thing I’m changing, exactly what “cited” means, and the decision rule — all locked today, 2026-07-03. The shape of it: 10 topics × 4 engines × 3 runs ≈ 120 controlled trials, with a pre-registered success bar of a GEO win-rate ≥65% at p < 0.05, all tracked on the experiments log. The experiment is running now. The readout lands around 2026-08-13. Nothing below leaks a result, because I don’t have one yet.
The hypothesis, in one sentence
A page built with the GEO formatting bundle is cited by AI answer engines more often than a plain-prose page carrying identical facts, claims, numbers, sources, and word count. If that’s true at the threshold I’ve set, structure earns citations. If it isn’t, structure is folklore for citation and I’ll say so.
The single manipulation
The whole experiment hinges on changing one thing: the packaging. Every topic is written twice. The facts, the specific claims, the numbers, the outbound sources, the author, the publish datetime, and the word count (matched within ±10%) are held identical across both versions. Only the structure differs.
| Element | GEO-structured (A) | Plain (B) |
|---|---|---|
| Opening | ≤60-word direct answer before any H2 | Answer arrives mid-prose, after a general intro |
| Sections | ≥4 self-contained H2s, each opening with its own answer | 1–2 generic headings (Background, Details) |
| Chunking | One claim per short paragraph | Longer flowing paragraphs |
| Data | ≥1 comparison / data table | Same data written out in sentences |
| FAQ | 3–5 Q&A block + FAQPage JSON-LD | None |
| Facts, claims, sources | Identical | Identical |
| Word count | Matched ±10% | Matched ±10% |
Both versions carry the same baseline Article/BlogPosting schema, so I’m testing the structural bundle — not “schema versus no schema at all.” Both are genuinely accurate and useful; the plain version is conventional, not deliberately thin. There are no sandbagged pages, because a rigged loss would prove nothing.
What “cited” means
Vague success criteria are how experiments lie to themselves, so I’m defining the outcome twice — once for each arm.
- Arm 1 (controlled): an engine is presented both versions, labelled neutrally as “Source 1” and “Source 2,” and asked to answer the target question and quote the passage it used. “Cited” = the version the engine drew its answer from. Ties are recorded as ties.
- Arm 2 (live): each published page’s target question is asked across the four engines weekly. “Cited” = the page is surfaced or linked as a source in the engine’s answer.
The two arms
Arm 1 is the fast, clean read. For each topic I show an engine both versions as neutral “Source 1 / Source 2,” randomise which version is Source 1 each trial to control position bias, and ask a fixed prompt to answer the question and quote the passage used. Because the facts are identical, any preference difference is the structure. Coverage: 10 topics × 4 engines (ChatGPT, Perplexity, Claude, Gemini) × 3 runs on different days = ~120 trials. I use the real front-ends, not the raw model APIs — they return different answers than what users actually see.
Arm 2 is the real-world confirmation. To avoid two near-identical pages cannibalising each other, each topic is published in one format only, on a dedicated cluster, all on the same day, submitted to Google and Bing with IndexNow fired. Assignment is stratified: I rank the 10 topics by expected citability, form 5 pairs, and randomly assign one of each pair to A and the other to B — the seed recorded before publishing. Then I track days-to-index and weekly citations across the four engines for 4–6 weeks (ChatGPT’s median time-to-first-citation runs about a week post-index, so I allow for churn). Those live pages are the running experiment, so I’m not linking them from here — a link from this post could bias the very thing I’m measuring.
The methods, at a glance
| Parameter | Value (locked 2026-07-03) |
|---|---|
| Topics | 10 matched AI-SEO questions, each written twice |
| Manipulation | GEO formatting bundle vs plain prose; facts held identical |
| Arm 1 — design | Controlled head-to-head, blind, position-randomised |
| Arm 1 — coverage | 10 topics × 4 engines × 3 runs ≈ 120 trials |
| Arm 2 — design | Live field test; one format per topic, stratified randomisation |
| Arm 2 — window | 4–6 weeks of index + citation tracking |
| Engines | ChatGPT, Perplexity, Claude, Gemini (real front-ends) |
| Primary test | Binomial / sign test on Arm 1 GEO win-rate vs 0.5 |
| Success threshold | GEO win-rate ≥65% AND p < 0.05 (Arm 1) |
| Readout | ~2026-08-13 |
The decision rule (locked)
Arm 1 is the primary test: the GEO win-rate across ~120 trials, evaluated with a binomial/sign test against chance (0.5), reported overall and per engine. The pre-registered bar for calling structure a real effect is a win-rate of at least 65% and p < 0.05. Arm 2 is directional — at 5 pages per arm I won’t claim statistical significance, only whether live citations and index speed point the same way.
Reading the two together:
- Both favour GEO → structure is fact. It goes into the HZ content playbook, and I publish “we tested it.”
- Arm 1 significant, Arm 2 flat → structure moves model preference but live retrieval is dominated by other factors → a nuanced writeup and a follow-up experiment.
- Neither → structure is folklore for citation at this threshold. That’s a real possible outcome, held at genuine — not token — probability, and I’ll report it as plainly as a win.
No goalpost-moving. The 65% / p < 0.05 bar is fixed as of today.
Common questions
What exactly is being tested?
Why hold the facts and word count constant?
What counts as a citation?
What is the success threshold?
When are results published, and will a null result be reported?
Related
- The experiments log — where the Experiment 02 readout will land, alongside the live index-lag data.
- How each AI engine sources and cites content — why ChatGPT, Perplexity, Claude, and Gemini read from different indexes, which is exactly why this test spans all four.
- Generative engine optimization (GEO) — the formatting bundle under test, defined.