Services
ORRJO Intelligence Brand, Content and Websites Demand Generation Lead Generation The Audit The Growth Engine Platform
Company
About Case Studies Pricing Careers
Also from ORRJO
ORRJO.aiBook a Strategy Call →

How to Choose an AEO Agency: The Eleven Questions That Separate the Real Ones

How do you choose an AEO agency?

Choose on evidence you can check yourself: a baseline citation test you watch them run, a prompt list built from your own customers, content with a real author's name on it, and a report you could reproduce without their software. Eleven questions get you there. The table below is the version to paste into an email.

The reason this needs a framework at all is that the usual buying signals are missing. Answer engine optimisation has been a named service for about two years. There is no long client roster to check. There is no agreed metric either, so the pitches all sound identical and the quotes do not. Two proposals for the same site can be eight times apart and both look reasonable on the page.

One disclosure before the table. This is published by an agency that sells this work, so point all eleven questions at us as well. If a question is uncomfortable for us to answer, it is doing its job.

The eleven questions

Questions one to four can be answered in a single call, and most of the separation happens there. Nine to eleven are the ones people skip and regret.

#The questionA good answer sounds likeA bad answer sounds likeWhy it matters
1Run the baseline test in front of me. What do the engines say about my category today?They share a screen, open clean logged-out sessions, run eight to ten buyer prompts, and save the raw answers with dates and the country setting."We will send a report next week." A visibility score out of 100 with no stated method.A baseline you did not watch being taken cannot be audited in month six, which is exactly when you will want to.
2Which prompts are we buying? Show me the list.Thirty to eighty questions in a buyer's own words, sorted into commercial and research intent, sourced from sales calls and customer interviews.A keyword export from a rank tracker with the word AI in the column header.The prompt is the unit of work here. A keyword list is a different product with a new label on it.
3Which engines do you test, how often, and from where?Named engines, a fixed cadence, a stated country, and logged-out sessions so the result is repeatable."All the major AI platforms." No cadence. No country.Engine answers vary by locale and account state. A test with no stated conditions cannot be repeated, so it cannot show movement.
4How much do these answers move on their own, without anyone doing any work?They separate the surfaces, because the surfaces behave differently, and they can point at published research when they say so.One screenshot of a good answer, presented as proof that a method works.Sistrix measured 74% of ChatGPT Search domains changing week to week. An agency that does not know this will sell you noise as progress.
5Who writes the content, and will their name be on it?A named writer or analyst with a role and a bio, whose byline goes on the page, and who attributes claims to sources in the prose."Our content team." Ghost bylines. A monthly word count as the commitment.Authorship and sourcing are the parts of this work an engine can actually read and reuse. Volume is not.
6Show me a page you wrote for another client, and the source list behind it.The live page, the sources, and at least one claim they refused to publish because they could not stand it up.Publishing cadence, word counts, a content calendar screenshot.The refusal is the tell. Anyone who has never dropped a claim for want of a source is not checking them.
7How will you prove a crawler can actually reach the page?Server or CDN log lines showing the named crawlers arriving, plus a robots.txt and firewall review before anything is written."We submit your pages to the AI engines."There is no submission queue. OpenAI's own crawler documentation says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
8What are you doing off-site, and what will you refuse to do?Named publications, podcasts, communities and directories, with a stated refusal to post fake recommendations in forums."We seed Reddit threads." Any answer involving accounts they control posing as customers.Bought forum recommendations get removed and the accounts get banned, so you pay for a citation with a half-life measured in weeks.
9What is in the monthly report, line by line?Prompt-level results with dates and saved answers, what changed since last month, what shipped, and what is queued next.A proprietary index number with no published method, rising every month.An index only the agency can compute will go up every month, because the agency computes it.
10Google gives you impressions and no clicks. How will you report on it?They know the generative AI performance report exists, know what it leaves out, and pair it with their own prompt testing.A promise of AI Overview click-through data, or a claim to attribute revenue to a specific citation.Google's report carries impressions only. Any click number for AI surfaces is modelled, and the model should be shown to you.
11What happens to the work if we stop paying you next month?Pages and schema live on your domain, and the prompt list plus the full test history are handed over as a readable file.Pages hosted on the agency's subdomain. The prompt list "lives in our platform".This is the question that decides whether you bought an asset or rented a dashboard.

Google publishes a generic version of this exercise in its own documentation for hiring an SEO, last updated 5 June 2026, which tells you to ask for prior work, for how success will be measured, and for a commitment to communicate every change to your site. That list is sound and it is old enough to predate all of this. The eleven above are the parts specific to answer engines, so run both.

What should an AEO agency actually deliver?

Six artefacts. If a proposal does not produce all six, it is either cheaper for a reason or it is an SEO retainer with a new cover page. The point of listing them as artefacts rather than activities is that artefacts can be compared across two quotes and activities cannot.

DeliverableThe artefact that proves it happenedThe unit to compare across quotes
Citation baselineRaw engine answers, saved and dated, for every prompt on the list, with the country and session state recorded.Number of prompts tested, and number of engines.
Prompt listA sheet of buyer questions with intent labels and the source of each one.Prompts, and whether they came from customer interviews or from a keyword tool.
Sourced contentPublished pages with named authors, visible dates, linked sources, and answer-first opening paragraphs.Pages a month, and whether writing is included or billed separately.
Structured dataArticle or BlogPosting, Organization and Service markup that validates, with author entities that resolve to a real page.Whether it is a one-off implementation or maintained as the site changes.
Off-site citationsA list of live third-party pages that mention you, with dates and how each one was earned.Placements a quarter, and whether outreach hours are inside the fee.
Monthly re-testThe same prompts run again under the same conditions, with a diff against last month.Cadence, and whether the raw answers are handed over or only summarised.

The FAQ schema trap

Watch the structured data line closely, because a lot of AEO proposals are still priced around FAQ markup. Google added a deprecation notice to its FAQ documentation in May 2025, the FAQ rich result stopped appearing in Google Search on 7 May 2026, support was pulled from Search Console reporting and the Rich Results Test from January 2026, and Google removed the documentation page entirely on 15 June 2026.

FAQPage markup is still valid schema.org and still a tidy machine-readable version of a question and its answer, which is why this page carries it. What it no longer does is produce a Google rich result. If a quote lists FAQ schema as a rich-result deliverable in 2026, the agency has not read Google's changelog since last spring, and you should ask what else it has missed.

The same caution applies to anything sold as an AI-specific file format or markup. Google's guidance on AI features, updated 10 December 2025, is blunt about it: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary", and "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." An agency charging a line item for a proprietary file is charging you for a file Google says it does not read.

What can no AEO agency promise?

Three claims should end a conversation, and there is a fourth that should at least slow one down.

Guaranteed placement in a named engine. Google's documentation on hiring an SEO says it plainly: "No one can guarantee a #1 ranking on Google." Its AI features guidance goes further, and states that meeting every requirement, best practice and policy still does not mean Google will crawl, index or serve a page. Nobody selling you AEO has a lever Google says does not exist.

A fixed number of citations. Sistrix ran 82,619 prompts and 1,548,213 snapshots across six countries including the UK, over seventeen weeks from 17 December 2025 to 8 April 2026, and published the results on 1 May 2026. In Google AI Mode, 56% of the domains cited changed every week. In ChatGPT Search, 74% were new every week, from responses that cite only three or four domains at a time. AI Overviews were the steady surface: eleven domains cited on average, eight of them consistently present, and in 53% of prompts not one source changed across the whole seventeen weeks. A citation count is a promise about a system that reshuffles most of its sources weekly on two of the three surfaces tested.

A share-of-voice guarantee. Share of voice needs a denominator, and no engine publishes one. Every share-of-voice number in this category is an agency's own sample of its own prompt list, which is a reasonable internal metric and a meaningless contractual one. Ask how the denominator is calculated. If the answer is the name of a tool, ask how the tool calculates it.

Revenue attributed to a citation. This is the one to slow down over rather than walk out over. Pew Research Center tracked the browsing of 900 US adults through March 2025, covering 68,879 Google searches of which 12,593 produced an AI summary, and published on 22 July 2025. Users clicked a source link inside the AI summary on 1% of visits to those pages. They clicked any result link on 8% of visits where a summary appeared, against 15% of visits where none did. Most of the value of a citation therefore arrives with no click and no referrer, which makes it real and makes it hard to prove. An agency that promises clean attribution is promising something the measurement layer does not support today. One that says so out loud, then shows you branded search volume, direct traffic and the self-reported source field on your enquiry form, is being straight with you.

Paid placement is its own product line, bought from the platform. OpenAI's crawler documentation lists a separate bot for it, OAI-AdsBot, described as validating the safety of web pages submitted as ads on ChatGPT. It is not something an agency arranges quietly on the organic side.

The two numbers the pitch will quote at you

Most decks in this category lean on two claims: that AI referral traffic converts better than everything else, and that it is growing at triple digits. Both trace back to Adobe Analytics. Adobe's release of 17 March 2025, drawn from more than a trillion visits, reported generative AI referral traffic to US retail sites up 1,200% in February 2025 against July 2024, alongside 8% higher engagement, 12% more pages per visit and a 23% lower bounce rate. Good data, and carefully described by Adobe as US retail, with travel and banking alongside it. There is no B2B in it.

No primary source publishes the B2B equivalent, and I would treat anyone who quotes one without a link as guessing. So when an agency puts those figures in front of a B2B buying committee, ask where they came from. The honest version of the pitch is about being present in the answer at all, which is a smaller claim and a true one.

Pair that with Cloudflare's crawl-to-refer measurements, published 1 July 2025. In the week of 19 to 26 June 2025, Anthropic's crawler made close to 71,000 HTML page requests for every referral it sent back, while Mistral sent ten times as many referrals as crawl requests. What these systems read and what they send back differ by orders of magnitude, and the gap is different for every platform. Budget for influence, and treat sessions as a bonus.

How much should AEO cost?

The published ranges, what each pricing model buys, and where those numbers come from are collected in the AEO agency pricing guide, so this section covers the more useful question, which is why two quotes for the same site come back so far apart.

Four things drive nearly all of the variance:

  • Whether content production is inside the fee. This is the biggest single swing. An audit-and-advise retainer and a retainer that ships eight sourced pages a month are different businesses at different costs, and both get called AEO.
  • The size and state of the estate. Forty pages with no author markup is a fortnight of work. Four hundred pages across three templates and a legacy blog is a quarter.
  • Whether off-site work is included. Earning third-party mentions is outreach, and outreach is hours. Quotes that leave it out are cheaper, and slower to move the answers that matter most.
  • Testing depth. Ten prompts on one engine monthly is a different purchase from eighty prompts on four engines weekly, and the line item usually reads the same on both proposals.

Normalise on those four before you compare totals. A proposal that comes in 40% cheaper while covering two of the four has not saved you anything.

For a reference point on this site, a one-off baseline audit run by a named analyst is £750, with monthly re-testing at £149 a month, and the ongoing work sits inside a brand and content retainer rather than a separate AEO line item. Whether an agency prices AEO standalone or folds it into an existing retainer tells you what it thinks the work is. Most of it is content and research with a different report on the end.

What to check before the first call

Twenty minutes of your own testing will change the conversation more than any proposal does. Open a clean browser session, log out, set the country, and ask four or five engines the questions your buyers ask. Write down who gets named, and in what order. Then ask each agency to run the same exercise live on the call.

Two things usually fall out of that. First, you find out whether you have a visibility problem or a positioning problem, because sometimes you are cited perfectly accurately and it is the description that is losing you the deal. Second, you find out whether the agency's version of the test matches yours. When it does not, ask why before you ask anything else.

Related reading on this site

The UK AEO agency comparison for who does what, which AI crawlers to allow in robots.txt for the technical precondition behind question seven, and the general agency selection questions that apply whatever you are buying.

Sources

Want the baseline test run on your category?

A named analyst puts your buyers' questions to the AI engines the way your buyers would ask them, several times each, and publishes the middle answer rather than the best one. Three working days, a fix list in sequence, and a walkthrough. £750.

See the audit