Search used to be about appearing in a list a person clicks. Increasingly it is about being the source an AI system quotes, often with no click at all. This walks through what actually changed, which techniques are real standards and which are marketing, what the research shows earns a citation, and why so much of it cannot be measured. One question runs through all of it: who says, and is it adopted?
Module 01
The shift
For twenty years the goal of a page was to rank: to sit high in a list of links so a human would click. Generative and answer engines move the target. Somebody asks a question and gets an answer built for them, sometimes with citations, often without a click. Rankings to citations. That is the whole shift, and everything else in this walkthrough follows from it.
The same question, answered two ways
What changed for the person who wrote the page
In the list, being useful and being visited were the same event. In the answer, they come apart. Your work can be read, summarised and relied upon by somebody who never arrives, never appears in your analytics, and could not name you if asked. The new question is not only how to rank. It is how to be the source the machine reaches for, and how you would even know if you were.
Those are two separate problems. The rest of this deals with the first one honestly, and admits how badly the second is solved.
Module 02
The words for it, and whether they are settled
SEO, GEO, AEO; do these mean anything fixed?
They do not. They are competing labels for one discipline that is still forming, and knowing that is worth more than memorising any of them.
Four labels, side by side
Why GEO is the one worth anchoring to
GEO is not a marketing coinage. It comes from a paper, "GEO: Generative Engine Optimization" by Aggarwal and colleagues, posted to arXiv in November 2023 and published at the KDD conference in 2024, with authors from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi. That matters less because the term is better and more because the paper contains the only part of this field with evidence behind it: an actual benchmark of which tactics worked.
Everything in Module 04 comes from that paper. When a vendor makes a claim about what earns citations, this is the thing to check it against.
So: which of the techniques you are being sold are real, and which are stories?
Module 03
What machines actually read
Myth we'll unpick: "llms.txt is the new robots.txt that replaces it"
Two files, two entirely different jobs, and only one of them is a standard.
They are not the same kind of thing
robots.txt is about permission. It tells a crawler what it may fetch, and it is written into a formal specification, IETF RFC 9309, published in 2022. llms.txt is about convenience. It offers a model a tidy Markdown map of your site so it does not have to wade through messy HTML. It grants nothing and blocks nothing. They are complementary at most; one cannot replace the other any more than a menu can replace a lock.
llms.txt was proposed by Jeremy Howard of Answer.AI in September 2024. The idea is reasonable. The claims made about it are where this goes wrong.
Place it on the spectrum
The evidence on llms.txt, plainly
●Almost nothing reads itAn analysis of 137,210 domains published in June 2026 found that of the roughly 38,000 that published a valid file, about 97 per cent received no requests for it at all in the sampled month. No bots, no humans.
●What does read it is the wrong audienceIn the same analysis, the leading readers were a training crawler and a coding tool, outranking every AI search and assistant bot. An independent 90-day test found the file hit in roughly one in a thousand AI-bot requests.
●No standards body has taken it upNeither the IETF nor the W3C has a process running on it, and its own author has said plainly that it is not an official standard.
●Google has declined, repeatedlyIts Search Relations staff have compared it to a long-abandoned signal, and a section of Google's generative-AI guidance published in 2026 tells site owners directly that machine-readable files of this kind are not needed to appear in AI search. No major AI company has documented its production answer system reading one.
●What would change the verdictOne thing, and it is specific: a major AI company publishing documentation that its live system fetches and uses llms.txt. Until then the honest label is "proposed convention, contested adoption". It costs little to publish, and it may help an agent already working on your site find its way around.
The deeper reason self-declared files get ignored
A file where a site describes itself to machines lets that site show machines a flattering version of itself that differs from what a person sees. That is cloaking, and it is exactly why search engines learned decades ago to discount signals a site asserts about itself. The lesson generalises well beyond this one file: an easily-gamed claim is the first thing a serious system stops trusting, which is why the things that do work are the ones that are hard to fake.
Which brings us to what is hard to fake, and what the research says about it.
Module 04
What earns a citation
A claim to test: the old tricks still work.
The GEO paper tested nine tactics across ten thousand queries and checked them against a live answer engine. Guess each one before you reveal it.
Helped, or hurt?
Keyword stuffing
The classic search manipulation: pack the target phrases in, as densely as the prose will bear.
It made things worseIt scored below the untouched baseline on the paper's main visibility metric, and below baseline again in the live validation. Not merely useless; actively counterproductive.
Adding statistics
Replacing vague quantities with specific figures attributed to a named source.
It helpedOne of the strongest tactics in the study, and stronger still in combination with clearer prose. Concrete, attributable numbers are quotable in a way that "many" and "most" are not.
Citations and quotations
Citing your sources, and quoting relevant ones directly rather than paraphrasing them away.
It helpedAdding quotations was the single best-performing tactic in the benchmark, with citations close behind. A page that shows its working gives the machine something to carry across.
Fluent, authoritative prose
No new facts; the same content rewritten to read better and speak with more confidence.
It helpedAmong the top performers, which is uncomfortable and worth naming: writing well moves the needle independently of whether the content improved. Combined with added statistics it beat any single tactic in the study.
These are research-benchmark results measured on the paper's own metrics, not promises about your site. The direction is trustworthy; the size of any particular gain is not a guarantee.
What that list actually amounts to
Read the winners back as a description of a page and you get: says something specific, attributes its numbers, quotes its sources, reads well, knows what it is talking about. That is not a growth tactic. That is just good, honest, well-sourced writing, which is a strange and rather cheering result for a field this full of hype. One more finding worth keeping: in the study, lower-ranked sites gained more from these tactics than sites already at the top.
Which is the encouraging part for a small, carefully written site. The advantage the incumbents hold in traditional search is not the same advantage here.
And why you can barely measure any of it
●Referrers get stripped, citations get no clickMany AI surfaces pass no referrer, and a citation the reader never clicks produces no visit at all. Ordinary client-side analytics therefore undercount AI influence systematically, not randomly.
●The new analytics channels are partialAnalytics platforms have begun adding AI-referral channels; one major platform added a native AI assistant channel in May 2026. Its published definition covers a handful of named assistants, excludes the platform's own AI answer surfaces, and leaves others landing in generic referral or direct buckets. Verify the current list before relying on it; this changed twice within a year.
●Server logs are the ground truthFor crawler activity, read the server or edge logs and look for the bots by name. A tool that runs in the browser cannot see a crawler that never runs its code. Crawler names change several times a year, so check them against each vendor's own documentation rather than a blog post.
●Vendor conversion multipliers are not findingsSeveral vendors publish figures showing AI-referred visitors converting many times better than ordinary search traffic. Those are vendor benchmarks, not peer-reviewed work. The robust, teachable fact is not a multiplier; it is that a large share of the influence is unmeasured.
Why this matters
The engines will change and the tactic lists will churn; most of what is sold as GEO today will read as quaint within two years. What lasts is the habit underneath. Ask who says, and whether it is adopted. Place each new technique somewhere on the spectrum from ratified standard to somebody's blog post before you adopt it. Build correctly, and report what you could and could not measure rather than the number that flatters you. Being the kind of source a machine reaches for turns out to look a great deal like being a good source, full stop.
Sources used: the origin of GEO and every tactic result in Module 04 are from Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization", arXiv:2311.09735, November 2023, published at ACM KDD 2024. robots.txt as a formal standard is IETF RFC 9309 (2022). The llms.txt proposal is Jeremy Howard / Answer.AI, 3 September 2024, at llmstxt.org. The adoption telemetry is an Ahrefs analysis of 137,210 domains published 15 June 2026, corroborated by an independent 90-day bot-traffic test. Google's position is from Search Relations statements and its own 2026 generative-AI guidance; the structured-data evidence from schema and citation studies of late 2025; the analytics detail from the platform's published channel definition after its May 2026 addition; the terminology framing from industry analysis of 2026. All of it carries dates and confidence labels in the reach report written for this course and its research brief, checked August 2026. Figures here are stated as research or telemetry results, not promises; re-verify crawler names and analytics coverage before use, as both change several times a year.