Original analysis

The Technical AEO Audit: A Practical Guide

A repeatable check for bot access, page output, main URLs, schema, speed, sitemaps, feeds, and AI discovery.

What matters

  • Test the full page path, including CDN and firewall behavior.
  • Use several crawler classes because search, training, and user-fetch are distinct.
  • Initial HTML should contain the primary answer and crawlable links.
  • A good audit names the proof and owner for each issue. A broad score is not enough.

Audit the path, not only the page

A source file can be perfect while the public page stays hidden. DNS, TLS, redirects, server access, CDN rules, bot rules, browsers, and app paths all affect delivery. Test the live URL through the same path a bot uses.

Step 1: list the public pages

List each main page type. Include the home page, bio, guide, article, case study, tool, service, policy, sitemap, feed, robots file, and error page. Add a few old URLs and URLs with query text. For each one, set the right status, main URL, robots state, content type, and sitemap state.

These rules become tests for each release.

Step 2: inspect response behavior

For each representative URL, verify:

  • HTTPS and one main host name.
  • No redirect chain.
  • Correct 200, 301, 404, or 410 response.
  • Correct Content-Type and useful cache headers.
  • No cookie or login needed for public facts.
  • A real not-found response rather than a soft-404 homepage.
  • The same response for a normal browser and real bots.

Repeat from another network if a firewall or regional server may change the result.

Step 3: inspect initial HTML

Download the response without running JavaScript. Confirm it contains:

  • A unique title and description.
  • One descriptive H1.
  • The main answer and all key text.
  • Named author, published or updated date, and source links.
  • Standard links for the menu and site references.
  • A full main URL that points to itself.
  • Matching index directives.
  • Valid JSON-LD for facts people can see.

JavaScript can improve a tool. People and bots should not need it to find the topic or read the main answer.

Step 4: test the crawler policy matrix

Run the same public URL using documented identities for:

  • Googlebot and Bingbot.
  • OpenAI’s OAI-SearchBot, GPTBot, and ChatGPT-User.
  • Anthropic’s Claude-SearchBot, ClaudeBot, and Claude-User.
  • PerplexityBot and Perplexity-User.
  • A neutral, honest custom user agent.

A bot name alone does not prove identity. Do not make unsafe allow rules from this test. The test finds surprise checks, blocks, or different text. Live firewall rules should use provider IP checks or trusted bot groups when possible.

Step 5: validate discovery files

Check that robots.txt is plain, easy to reach, and links to the main sitemap. Each sitemap URL should be full, current, open to search, and working. Test the XML. Remove old or moved URLs fast.

An llms.txt file can be a small list for tools that choose to read it. It is not a search rule or security tool. It cannot replace site links or a sitemap.

Step 6: compare page text and schema

Test every JSON-LD block. Check each linked @id. Then compare key fields with the page:

  • Person name, roles, group ties, and profiles.
  • Article title, dates, author, and main page.
  • Breadcrumb names and destinations.
  • Organization identity and contact details.
  • Service descriptions, offers, and regions.

Missing extra schema is often safer than false schema.

Step 7: enforce performance budgets

Measure real Core Web Vitals when the site has traffic. Run lab tests before launch. Here is a useful budget for a static site:

Resource Internal target
Initial HTML under 100 KB compressed
Critical CSS under 25 KB compressed
Default JavaScript 0 KB on reading pages
Hero image under 160 KB at the served size
Fonts system stack or tightly subset, self-hosted files

These limits are budgets, not rank factors. They prevent surprise slowdowns and leave room for real devices.

Step 8: test access for people and agents

Use the site with a keyboard. Check heading order, focus, link names, form labels, errors, page regions, and a 320-pixel screen. Review the access tree. Clear HTML helps screen readers, agents, and all users.

Step 9: make a task list with proof

For each failure, save the URL, time, request name, headers, key page text or image, expected result, impact, owner, and retest date. Mark it as a block, key issue, or improvement. Skip mystery health scores that do not explain the failure.

Use the free robots policy builder for robots rules. The paid AEO Visibility Snapshot adds live tests, fact review, a prompt baseline, and a ranked task list.

Sources

  1. JavaScript SEO basics, Google Search Central
  2. Build and submit a sitemap, Google Search Central
  3. Understanding Core Web Vitals, Google Search Central
  4. Does Anthropic crawl data from the web?, Anthropic
  5. OpenAI bots, OpenAI
  6. Perplexity bots, Perplexity