Two things that come after getting a site online
You can already publish a static site. This course covers what happens next.
Access control. Stopping people you have not chosen from reading a page. Most tutorials, and most AI coding assistants, produce a passcode box that does not do this. You build one, break it, then build one that works.
Machine readability. How AI systems read your site and decide whether to quote it. This half sorts the techniques with evidence behind them from the ones vendors are selling.
Both halves turn on one question: where is the decision made? A gate works when the decision to send or withhold happens on the server. A signal you publish works when the system reading it has decided to use it.
No rankings, no citation counts, no traffic numbers. Much AI influence cannot currently be measured, and the techniques with evidence behind them come from research benchmarks rather than guaranteed field results. What the course teaches instead: how to build correctly, how to check a claim against its primary source, and how to report what you could not measure.
Who this is for, and what you need first
Written for someone who already publishes static sites and wants to know what the two new problems are and how to solve them.
// you need- Comfort editing HTML and CSS, and a text editor
- A repository host account, and basic Git
- A site already deploying, or willingness to set one up first
- A browser with developer tools, which is all of them
- No paid accounts to begin. Two phases mention paid tiers, always with a free path alongside
No backend or framework experience is assumed. If you have never published a site, start with Putting a website online and come back.
The two streams are independent. Take Stream A alone if you only need access control. Take Stream B alone if you work in marketing or communications and the security half is not your problem.
// what you will be able to do- Explain where code runs and why that decides everything else
- Demonstrate why a browser-checked passcode protects nothing
- Choose between three access-control options and justify the choice
- Store secrets so a repository leak does not defeat your gate
- Tell a ratified standard from one party's proposal
- Decide which AI crawlers to allow and which to block
- Write pages that the research says get quoted
- Find AI crawler activity in your own server logs
- Report on your own site, including what you could not measure
Contents
- FoundationHow a page gets delivered
- FoundationHow to check a claim before you act on it
- Stream AAccess control
- Stream BBeing read and cited by AI systems
- CapstoneApply both streams to your own site
- OptionalAdapting Stream B for a marketing audience
- Before useWhat to re-verify before delivering this
- ReferenceSources used
How a page gets delivered
What happens between a visitor's request and your page appearing on their screen. Both streams depend on this.
// covers- The request and response cycle
- Where code runs: in the browser, on the server, or at the edge
- What a static site can and cannot do as a result
- How a crawler differs from a human visitor
- The difference between a model answering from training data and a system fetching your live page now
The point to take from this phase
Code you send to the browser runs on the visitor's machine, under their control. Code that runs on the server does not. Stream A is about that difference. Stream B is about which of the things you publish are actually read, and by what.
Watch a page arrive
Open developer tools on any site and reload with the network panel open. Read the response headers on the main document. Then open the sources panel and look at the JavaScript. Everything visible there is visible to any visitor.
How to check a claim before you act on it
A method for telling a web standard from a proposal. You will need it constantly in Stream B, where most of the published advice comes from vendors selling something.
// covers- What makes something a standard: a standards body, a published specification, and adoption in practice
- What a proposed convention is: published by one party, adopted variably or not at all
- Worked comparison, robots.txt against llms.txt. robots.txt is IETF RFC 9309, published 2022 (high). llms.txt was proposed in September 2024 and has no standards-body process (high)
- How to find the primary source behind a vendor claim, which is usually one link away
The point to take from this phase
Two questions to ask about any technique in this course, and any technique someone sells you later. Who says it works, and has anyone adopted it?
Trace one claim to its source
Take a confident claim from a vendor blog about AI search. Find the primary source it rests on. Note what the source actually says, when it was published, and whether the claim survived the trip.
Access control
In one sentence: a passcode checked in the browser protects nothing, and understanding why tells you how access control works.
Build a passcode gate, then break it
You build the gate an AI assistant writes when you ask it to make a page private, then defeat it three ways using tools already in your browser.
// covers- Building a frontend-only JavaScript passcode gate
- Reading the passcode out of the page source with developer tools
- Requesting the "protected" page directly, so the gate is never involved
- Editing the check in the console so it stops asking
Build the broken version yourself. It takes about ten minutes, and none of the rest of this stream lands without it.
Why a browser check cannot work
Three separate failures, and the general rule underneath them.
// covers- The passcode is readable. It had to be sent to the browser for the browser to compare against it. Hashing does not help: the hash arrives too, and can be tested against guesses offline
- The content was never withheld. If a script hides a section, that section is in the downloaded HTML. If it redirects instead, the target page is usually served to anyone who requests its address
- The check can be rewritten. The comparison runs on the visitor's machine, so the visitor can change it. They do not need the passcode; they can remove the question
- What the broken version is genuinely good for: keeping a page out of the way of a casual passer-by, and nothing more
The rule, and what else it applies to
Anything the browser can check, the visitor can read and change. That covers hidden prices, disabled buttons, client-side validation and feature flags, not just passcodes. Real access control decides what to send before it sends it.
Three access-control options that work
For the common case: showing a site to people you choose, without running a subscriber system. All three are real. They differ in effort, cost and how finely you can target them.
| Option | How it works | Effort and cost | Reach |
|---|---|---|---|
| Server function with a signed cookie | A serverless function compares the submitted passcode against a secret in an environment variable. On success it returns a signed, HTTP-only, Secure cookie. Content is served only when a valid cookie arrives. | Most effort. Usually free at personal-site scale on standard function allowances (medium, verify current) | Per page or per section |
| Edge access with a one-time PIN | An access layer in front of the site intercepts every request and emails a short-lived PIN to an allow-listed address. Nothing is served until the PIN is accepted. | No code. A free tier covers a small allow-list, then per-user pricing (medium, verify current) | Whole site or per path |
| Host-level site password | One password for the whole site, set in the host's dashboard and enforced at the host's edge before anything is served. | Least effort. Paid plan feature on the hosts that offer it (medium, verify current) | Whole site only |
What each protects against
The server function stops the passcode being read, the token being forged, and scripts stealing the cookie. It does not stop guessing, so choose a long passcode and consider limiting attempts. The other two stop everything in this stream, for the same reason: nothing is delivered until the check passes.
One option to know about and not use
Hosts have offered built-in account systems that give a static site real logins. At least one was announced as deprecated in early 2025, then partially walked back to "still supported", with no active development since (medium; the status is genuinely ambiguous). That is the worst state for a dependency: still recommended in old tutorials, but nothing will be fixed. Check whether a convenience feature is maintained before you build on it.
Pricing and plan names for all three options changed within the past year, and one host moved to a different pricing model entirely. Dated figures with confidence labels are in the access-gates resource written for this course. Re-verify against each vendor's current page before quoting a number to anyone. No price appears on this page or in the interactive explainer, deliberately.
Where to put secrets
The gate is only as good as where its signing secret lives.
// covers- Why secrets go in the host's environment variables, not in committed code
- Why deleting a committed secret does not remove it, because it stays in the history
- Why a repository holding gate logic should be private: the logic describes how to get past it
- What someone can do with a leaked signing secret, which is forge a valid token and enter as a legitimate visitor
Private repository and environment variables are two separate protections. Use both.
Being read and cited by AI systems
In one sentence: the target has moved from ranking in a list of links to being a source an AI system quotes, and most published advice on how to do that has no evidence behind it.
SEO, GEO and AEO: what the terms mean
Four labels, what each one means, and how settled it is. The terminology is contested, so this phase exists to stop you treating vendor jargon as fixed categories.
| Term | Means | How settled |
|---|---|---|
| SEO | Search Engine Optimisation. Ranking in ordinary search results. | Settled. Nobody argues about it (high) |
| GEO | Generative Engine Optimisation. Visibility inside a generative engine's answer. | The safest label to use, because it comes from a research paper rather than a vendor (high) |
| AEO | Answer Engine Optimisation. Being the extracted answer in featured snippets, AI overviews and assistant replies. | "Answer" is the mainstream reading (high). "Agent Engine Optimisation" is a minority usage (medium; may be rising) |
| LLMO, AIO, "AI SEO" | Overlapping descriptions of the same practice. | Vendor labels. Not settled (medium) |
Where GEO comes from
"GEO: Generative Engine Optimization" by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, arXiv:2311.09735, posted November 2023 and published at ACM KDD 2024, with authors from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi (high). Every tactic finding later in this stream comes from that paper. It is the thing to check vendor claims against.
What machines read, ranked by how settled it is
Four layers, from ratified standard down to unproven proposal. Learn to place a new technique on this list before adopting it.
| Layer | Status |
|---|---|
| robots.txt | A genuine standard, IETF RFC 9309 (2022), respected in stated policy by the major search crawlers and most major AI crawlers. It controls what a bot may fetch, and nothing else (high) |
| Semantic HTML and an XML sitemap | Foundational and universally respected. The base layer of being readable by any machine (high) |
| Schema.org data as JSON-LD | Useful infrastructure, mixed evidence on citations. Helpful on some AI surfaces; a late-2025 study found no correlation between schema coverage and citation rates in several major assistants, and some engines missed data placed only in the JSON-LD. Treat it as amplification of an already-strong page, not a citation switch (medium; evidence genuinely mixed) |
| llms.txt | A proposed convention with contested adoption. Proposed September 2024, no standards-body process, described by its own author as not an official standard, and largely unread by the systems it is aimed at on 2026 telemetry (high). It is not a replacement for robots.txt, which does a different job. Cheap to publish; not currently earning citations |
What would change the llms.txt verdict is specific: a major AI company documenting that its production answer system fetches and uses the file. That has not happened.
Which AI crawlers to allow, and which to block
The distinction that trips people up, and why blocking the wrong crawler removes you from AI answers.
// covers- Training crawlers. Blocking one is a content-rights decision. It costs you nothing in AI-answer visibility (high)
- Retrieval and search crawlers. Blocking one removes you from that system's live answers. This is self-inflicted and common (high)
- Reading logs, because robots.txt states your policy and logs show what happened
- Crawlers that ignore robots.txt, and why blocking those needs edge or firewall rules rather than a Disallow line (high)
Check the crawler names against the vendors
Crawler names change several times a year, so this is an exercise rather than a table to memorise. Take the list from the course, open each vendor's own bot documentation, and confirm or correct it. Record the date you checked.
What earns a citation
The GEO paper tested nine tactics across ten thousand queries and validated against a live answer engine. Here is what worked and what did not.
| Tactic | Result |
|---|---|
| Quoting sources directly | Best single tactic in the benchmark (medium; benchmark result) |
| Adding citations | Among the strongest (medium; benchmark result) |
| Adding statistics from named sources | Among the strongest, and stronger again combined with clearer prose (medium; benchmark result) |
| Writing more fluently, with an authoritative voice | Among the strongest, with no change to the facts on the page (medium; benchmark result) |
| Keyword stuffing | Scored below the untouched baseline. Worse than doing nothing (high, within the study) |
What to do with that
Front-load the answer. Use headings that match the questions people actually ask. Attribute every number to a named source. Quote the sources you rely on. Write clearly. The winning tactics describe good, well-sourced writing rather than a technical trick.
One more finding: lower-ranked sites gained more from these tactics than sites already at the top (medium). Treat the paper's percentages as benchmark results, not predictions about your site.
Why you cannot measure most of this
Ordinary analytics undercount AI influence systematically, not randomly. Knowing the size and shape of the gap is part of the job.
// covers- Many AI surfaces pass no referrer, so the visit lands in a generic bucket (high)
- Many citations produce no click at all, so there is no visit to record (high)
- Analytics platforms have started adding AI referral channels; coverage is partial, excludes some surfaces, and changed twice within a year (medium; specifics in flux)
- Server and edge logs are the reliable way to confirm crawler activity, because a browser-based tool cannot see a crawler that never runs its code (high)
- Vendor conversion multipliers for AI-referred traffic are vendor benchmarks, not peer-reviewed findings (medium)
The defensible statement is not a multiplier. It is that a large share of the influence is unmeasured.
Apply both streams to your own site
Four pieces of work on a site you actually run, and one piece of writing.
- Add a real access gate to a chosen section, and be able to say why you picked that option over the other two
- Make the public part of the site machine-readable, working up the layer list rather than down it
- Find AI crawler activity in your own logs. Record which crawlers, and when you looked
- Audit your secrets: nothing sensitive in the repository or its history, and the repository private if it holds gate logic
- Write a short piece on what you could and could not measure
How it is marked
The writing carries the most weight. "I could not determine whether this made any difference, and here is why" is a stronger result than an unfounded number. Reporting uncertainty accurately is the habit this course is trying to build.
Adapting Stream B for a marketing audience
Stream B works for a marketing or communications group with a shared foundations layer added underneath and an applied layer added on top.
// shared foundations layer- What a large language model is, and how it is trained
- How it stores and recalls content
- The difference between generative and agentic use
- Whatever AI policy and compliance rules the delivering organisation works under
- How AI systems and traditional search treat content differently
- Rankings against citations, and what each is worth
- Crawlability and scannability
- The trade-off between AI surfaces, which are strong for brand-perception and comparison queries, and traditional search, which is stronger for high-intent conversion
- Where models form their impressions: which sources they weight, including community platforms, video and established references
This layer works best against the organisation's own performance data rather than generic examples. Say so when someone commissions it.
What to re-verify before delivering this
Two categories of fact here go stale faster than the course does. Both are flagged wherever they appear.
- Pricing and plan names for every option in Stream A. One vendor changed pricing model within the past year
- The AI crawler names in Stream B, which change several times a year
- Which assistants your analytics platform's AI channel currently covers, which changed twice within a year
- Whether any major AI company has documented its production system reading llms.txt. That single fact would change the verdict in the layers phase
- Whether the "Agent" reading of AEO has moved from minority toward mainstream
Both are single self-contained files that run without a build step and work opened directly from disk. The security explainer is a closed simulation: no working passcodes, no signing code, nothing that functions against a live site. The GEO explainer keeps every status claim in line with the confidence labelling of the report behind it.
Sources used
Every factual claim above traces to one of these. Nothing on this page is invented, and no citation is fabricated.
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, "GEO: Generative Engine Optimization", arXiv:2311.09735, posted November 2023, published at ACM KDD 2024. Source for the origin of GEO and every tactic result in "What earns a citation". (high)
- IETF RFC 9309, Robots Exclusion Protocol, 2022. Source for robots.txt being a formal standard. (high)
- Jeremy Howard / Answer.AI, the llms.txt proposal, 3 September 2024, at llmstxt.org. Source for what llms.txt is and its non-standard status. (high)
- Ahrefs telemetry analysis of 137,210 domains, published 15 June 2026, corroborated by an independent 90-day bot-traffic test. Source for llms.txt adoption in practice. (medium to high; independent sources agree on direction, exact figures vary)
- Google Search Relations public statements and Google's 2026 generative-AI guidance. Source for Google's position that llms.txt is not used or needed. (high)
- Schema and citation studies, late 2025. Source for the mixed evidence on structured data and AI citations. (medium)
- Analytics platform channel documentation, 2026. Source for AI referral channel coverage and its gaps. (medium; specifics in flux)
- Vendor documentation, pricing pages and changelogs for the three access-control options, including a host blog post of 3 November 2025 on password protection and an edge provider's Zero Trust documentation on one-time PINs. (medium; pricing changes frequently)
- Standard web security references for client-side against server-side enforcement, token signing, and the HTTP-only and Secure cookie attributes. (high)
The two companion documents written for this course hold the full detail with dates and confidence labels on each claim: a self-teaching resource on access gates and login options, and a report on the shift from SEO to GEO and how AI systems cite sources. Both name what their author could not verify.
Not verified, and flagged as such
The exact current pricing of all three access-control options; the exact current crawler names; which assistants each analytics platform captures today; and whether the "Agent" reading of AEO is still a minority usage. Check each before delivery.
Last updated 14 August 2026
