TL;DR / Key Takeaways
- GEO's founding paper measured the field: quotations +41%, statistics +40%, citations +30–40% visibility — while keyword stuffing did roughly nothing.
- Generative engines are biased toward challengers: rank-5 sites gained +115% from GEO methods while top-ranked sites lost ~30% of answer share.
- The three gates are retrieval (crawlability), synthesis (liftable claims and numbers), and citation — and ChatGPT's API only browses the web ~74% of the time.
- The heavy half of GEO is off-site: engines assemble answers mostly from third-party pages, so coverage and mentions decide what gets said about you.
You'll meet this category under three names — GEO, answer engine optimization (AEO), and AI visibility — plus "LLM SEO" in developer circles. They're one field; search demand just hasn't converged (our July 2026 pull of Google Ads volumes: ~4,400/month for "generative engine optimization," ~2,400 for "answer engine optimization," ~1,600 for "ai visibility tools," ~880 for "llm seo"). This guide uses GEO because the research literature does.
Where the term comes from — and what the paper actually measured
"GEO: Generative Engine Optimization" (Aggarwal, Murahari et al., published November 2023, presented at KDD 2024) built GEO-bench — 10,000 real user queries across nine domains — and tested nine content interventions against generative engines, measuring how much answer real-estate each source captured. That measurement is the field's founding artifact, and its numbers are worth knowing precisely, because almost every blog citing "GEO boosts visibility 40%" drops the method breakdown that makes the finding useful:
| Intervention | Measured effect on visibility | Takeaway |
|---|---|---|
| Add quotations | ~+41% (position-adjusted word count) | Quotable sentences get lifted into answers |
| Add statistics | ~+40% | Numbers are the most portable form of authority |
| Cite sources | +30–40% | Content that shows receipts gets treated as a receipt |
| Fluency / readability edits | +15–30% | Clear prose is easier to synthesize from |
| Keyword stuffing | ≈ no improvement | The classic SEO reflex does nothing for answers |
The paper's most under-quoted finding is about who benefits: lower-ranked sites gained dramatically more — adding citations produced a 115% visibility increase for sites ranked fifth in traditional search, while top-ranked sites actually lost about 30% of answer share. Generative engines don't just re-serve the incumbent order; they re-decide it per answer. If you've been losing the ten-blue-links game to entrenched domains, GEO is the first ranking regime in years that's structurally biased toward challengers.
How generative engines actually build an answer
Every major engine follows the same three steps: retrieve candidate pages (usually via a conventional search index), synthesize an answer from them, and — sometimes — cite the sources used. Each step is a gate you can fail. If crawlers can't fetch your page or your text only exists after JavaScript runs, you fail retrieval. If your content has nothing liftable — no clear claims, numbers, or quotable sentences — you fail synthesis. And whether you get cited depends heavily on a variable most dashboards hide: from our own production monitoring, ChatGPT's API browsed the web on only ~74% of calls — the rest answered from frozen training memory, where months-old third-party descriptions of you, not your current site, decide what gets said.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
What GEO means in practice: the two halves
On your site — the half you control:
- Pass retrieval mechanically. Allow the AI crawlers in robots.txt, serve server-rendered text, ship titles, descriptions and structured data. The free ten-second check covers all of it.
- Write liftable content. Every important page should contain the three things the paper measured: concrete statistics, quotable single-sentence claims, and linked sources. A page that answers a question in one clean sentence is a page an engine can reuse.
- Publish original data. Engines cite numbers, and someone has to be the source of them. First-party measurements are the most reliably cited artifact type we've seen — and the one thing competitors can't copy.
Off your site — the half most guides skip: generative engines mostly assemble answers from third-party pages — comparisons, reviews, directories, coverage. In our monitoring, the pages engines actually cite for buyer questions are overwhelmingly not the vendors' own sites. That makes presence on retrieved third-party surfaces the heavy half of GEO: the mentions, listings and editorial coverage that exist about you on the pages engines fetch. It's the same authority-building work that link building has always gestured at — except the currency is now being quotable in someone else's answer-worthy page, not just the link's PageRank.
What doesn't work, measured
- Keyword stuffing — approximately zero effect in the founding paper's tests. Density-tuning tools are optimizing for the old physics.
- llms.txt as a ranking play — no major AI service has announced using it; publish one as a ten-minute courtesy, not a strategy.
- Buying your way into answers — OpenAI's own ad policies state ads don't influence ChatGPT's answers, and ads never reach paid subscribers; we took that apart with the primary sources.
- Chasing a single mention count — corpus mention counts are floors, not rankings; we've measured a site with a corpus count of one getting cited in three of four live answers. Any vendor selling a zero as proof of invisibility is selling a vibe.
Measuring GEO without fooling yourself
Three rules from running our own monitor: insist on web-grounded numbers (and ask every vendor whether theirs are), treat all mention counts as "at least N", and always read the row-level answers with their cited sources — the receipts are the product. The tool market for this is mapped, with measured collection costs, in our AI visibility tools guide; collecting one question across six engines costs about 24.5¢ at vendor-billed rates, so never pay for scores you can't click through.
The honest summary
GEO is real, measured, and unusually kind to underdogs — but it isn't magic and it isn't separate from the fundamentals. Make your site mechanically readable, write pages worth quoting with numbers worth citing, and build the third-party surface area engines retrieve from. Disclosure, as ever: we sell in this market — Stork Wire builds exactly that third-party coverage, and our free checker is the no-cost place to start — so weight our framing accordingly. The paper's numbers, though, aren't ours. They're the field's, and they're linked above.

