Blog
AI SEO Jul 3, 2026 7 min read

Experiment 02: does structure earn AI citations? The pre-registration

A pre-registered A/B test: does the GEO formatting bundle make a page likelier to be cited by ChatGPT, Perplexity, Claude, and Gemini than identical plain prose?

Everyone in AI SEO tells you to write “answer-first”: lead with a direct answer, break the page into self-contained sections, chunk one claim per paragraph, add a comparison table, bolt on an FAQ with FAQPage schema. The claim is that this GEO formatting bundle makes a page more likely to be cited by AI engines. I have not seen it isolated from everything else that moves citations — better facts, more sources, longer copy, an older domain. So I’m testing it, and I’m publishing the design before I know the answer.

This post is the pre-registration. It states the hypothesis, the single thing I’m changing, exactly what “cited” means, and the decision rule — all locked today, 2026-07-03. The shape of it: 10 topics × 4 engines × 3 runs ≈ 120 controlled trials, with a pre-registered success bar of a GEO win-rate ≥65% at p < 0.05, all tracked on the experiments log. The experiment is running now. The readout lands around 2026-08-13. Nothing below leaks a result, because I don’t have one yet.

The hypothesis, in one sentence

A page built with the GEO formatting bundle is cited by AI answer engines more often than a plain-prose page carrying identical facts, claims, numbers, sources, and word count. If that’s true at the threshold I’ve set, structure earns citations. If it isn’t, structure is folklore for citation and I’ll say so.

The single manipulation

The whole experiment hinges on changing one thing: the packaging. Every topic is written twice. The facts, the specific claims, the numbers, the outbound sources, the author, the publish datetime, and the word count (matched within ±10%) are held identical across both versions. Only the structure differs.

ElementGEO-structured (A)Plain (B)
Opening≤60-word direct answer before any H2Answer arrives mid-prose, after a general intro
Sections≥4 self-contained H2s, each opening with its own answer1–2 generic headings (Background, Details)
ChunkingOne claim per short paragraphLonger flowing paragraphs
Data≥1 comparison / data tableSame data written out in sentences
FAQ3–5 Q&A block + FAQPage JSON-LDNone
Facts, claims, sourcesIdenticalIdentical
Word countMatched ±10%Matched ±10%

Both versions carry the same baseline Article/BlogPosting schema, so I’m testing the structural bundle — not “schema versus no schema at all.” Both are genuinely accurate and useful; the plain version is conventional, not deliberately thin. There are no sandbagged pages, because a rigged loss would prove nothing.

What “cited” means

Vague success criteria are how experiments lie to themselves, so I’m defining the outcome twice — once for each arm.

  • Arm 1 (controlled): an engine is presented both versions, labelled neutrally as “Source 1” and “Source 2,” and asked to answer the target question and quote the passage it used. “Cited” = the version the engine drew its answer from. Ties are recorded as ties.
  • Arm 2 (live): each published page’s target question is asked across the four engines weekly. “Cited” = the page is surfaced or linked as a source in the engine’s answer.

The two arms

Arm 1 is the fast, clean read. For each topic I show an engine both versions as neutral “Source 1 / Source 2,” randomise which version is Source 1 each trial to control position bias, and ask a fixed prompt to answer the question and quote the passage used. Because the facts are identical, any preference difference is the structure. Coverage: 10 topics × 4 engines (ChatGPT, Perplexity, Claude, Gemini) × 3 runs on different days = ~120 trials. I use the real front-ends, not the raw model APIs — they return different answers than what users actually see.

Arm 2 is the real-world confirmation. To avoid two near-identical pages cannibalising each other, each topic is published in one format only, on a dedicated cluster, all on the same day, submitted to Google and Bing with IndexNow fired. Assignment is stratified: I rank the 10 topics by expected citability, form 5 pairs, and randomly assign one of each pair to A and the other to B — the seed recorded before publishing. Then I track days-to-index and weekly citations across the four engines for 4–6 weeks (ChatGPT’s median time-to-first-citation runs about a week post-index, so I allow for churn). Those live pages are the running experiment, so I’m not linking them from here — a link from this post could bias the very thing I’m measuring.

The methods, at a glance

ParameterValue (locked 2026-07-03)
Topics10 matched AI-SEO questions, each written twice
ManipulationGEO formatting bundle vs plain prose; facts held identical
Arm 1 — designControlled head-to-head, blind, position-randomised
Arm 1 — coverage10 topics × 4 engines × 3 runs ≈ 120 trials
Arm 2 — designLive field test; one format per topic, stratified randomisation
Arm 2 — window4–6 weeks of index + citation tracking
EnginesChatGPT, Perplexity, Claude, Gemini (real front-ends)
Primary testBinomial / sign test on Arm 1 GEO win-rate vs 0.5
Success thresholdGEO win-rate ≥65% AND p < 0.05 (Arm 1)
Readout~2026-08-13

The decision rule (locked)

Arm 1 is the primary test: the GEO win-rate across ~120 trials, evaluated with a binomial/sign test against chance (0.5), reported overall and per engine. The pre-registered bar for calling structure a real effect is a win-rate of at least 65% and p < 0.05. Arm 2 is directional — at 5 pages per arm I won’t claim statistical significance, only whether live citations and index speed point the same way.

Reading the two together:

  • Both favour GEO → structure is fact. It goes into the HZ content playbook, and I publish “we tested it.”
  • Arm 1 significant, Arm 2 flat → structure moves model preference but live retrieval is dominated by other factors → a nuanced writeup and a follow-up experiment.
  • Neither → structure is folklore for citation at this threshold. That’s a real possible outcome, held at genuine — not token — probability, and I’ll report it as plainly as a win.

No goalpost-moving. The 65% / p < 0.05 bar is fixed as of today.

Common questions

What exactly is being tested?
Whether the GEO formatting bundle — an answer-first opener, four or more self-contained H2s, one claim per paragraph, a comparison table, and an FAQ with FAQPage schema — makes a page more likely to be cited by AI engines than plain prose carrying identical facts, claims, sources, and word count. The only thing that changes between the two versions is the structure.
Why hold the facts and word count constant?
So the test isolates structure. If the GEO version also had better facts, more sources, or more words, any citation advantage could be from those instead of the formatting. Matching everything but the packaging is what makes a difference attributable to the packaging.
What counts as a citation?
In the controlled arm, it's which of the two labelled sources an engine draws its answer from and quotes. In the live arm, it's whether a page is surfaced or linked as a source when its target question is asked across ChatGPT, Perplexity, Claude, and Gemini.
What is the success threshold?
It is pre-registered: a GEO win-rate of at least 65 percent across roughly 120 controlled trials, and a p-value below 0.05 on a binomial test against chance. Below that bar, the bundle does not clear the pre-set threshold for calling it a real citation effect.
When are results published, and will a null result be reported?
The readout is around August 13, 2026. Yes — a null or negative result will be reported as plainly as a positive one. That commitment is the whole point of pre-registering the design in public before the data exists.