If your AEO agency cannot show you citation rate, share of voice against named competitors, and AI referral traffic from three independent sources, you cannot tell whether the retainer is working, and you probably should assume it is not. Three reports settle most of the argument for free: the Bing Webmaster Tools AI Performance report, which shows Copilot citations directly; Google Search Console, which shows AI Overview impressions; and Google Analytics 4, which added an AI Assistant default channel group in May 2026 covering referrals from chatgpt.com, perplexity.ai, and gemini.google.com. Any agency charging $2,000 to $15,000 a month should be pulling all three, and you should be able to log in and verify every number yourself.
The reason this needs saying is that AEO reporting is currently the least standardized part of digital marketing. There is no equivalent of a rank tracker everyone agrees on, which leaves enormous room for reports that look impressive and prove nothing.
Which 6 metrics actually measure AEO performance?
1. Citation rate
Of the prompts you track, what percentage produce an answer that cites or names you? This is the core number. A baseline of 5 percent moving to 25 percent over two quarters is real progress; a report that never states the denominator is hiding something.
2. Share of voice against named competitors
For your tracked prompts, how often are you named versus each specific competitor? Not “visibility increased,” but “we appear in 18 of 50 prompts, competitor A appears in 31.” Named comparison is the only version of this metric that survives scrutiny.
3. AI referral traffic and its conversion rate
GA4’s AI Assistant channel group makes this trivial to check. Track sessions, but weight conversion rate more heavily, because AI referred visitors arrive pre researched and typically convert at multiples of organic traffic. Small session counts with strong conversion are a healthy early signal.
4. Copilot citations from Bing Webmaster Tools
The most transparent native surface available, showing which of your pages Copilot actually cited and for which queries. It is free, first party, and impossible for an agency to embellish.
5. Published output count
Pages published or rewritten, schema types deployed, listings reconciled, placements earned. Not a vanity metric, the input metric. Citation gains lag production by weeks, so if output is zero, no outcome metric will move later either.
6. Prompt level movement, not aggregate scores
Proprietary “visibility scores” invented by a vendor cannot be verified or compared. Demand the prompt list and the per prompt result.
Want an independent read on your current citation position before your next agency review? Get your free AI visibility audit and compare what the engines actually say against what your monthly report claims.
What should the reporting stack look like?
Layered, with independent verification at each level. First party sources you own and can check without the agency: Bing Webmaster Tools, Google Search Console, and GA4. Third party tracking, which the agency likely licenses: Profound at roughly $2,000 to $5,000 a month for enterprise deployments, Peec AI from around EUR 89 to EUR 199 a month, Scrunch from about $250 a month, or Otterly from $29 a month. Manual spot checks, which cost nothing and catch the most: pick ten of your buyers’ real questions, run them yourself through ChatGPT, Perplexity, Gemini, and Copilot once a month, and log who gets named.
That third layer matters more than it sounds. Tools sample; buyers do not. If your agency reports a 30 percent citation rate and your own ten manual prompts return your competitor four times and you zero, one of those pictures is wrong, and it is worth an uncomfortable conversation to find out which.
What are the 90, 180, and 365 day checkpoints?
Judge by the right milestone at the right time, because engine refresh cycles are fixed and no agency controls them. Perplexity retrieves live and can reflect new content within one to two weeks. ChatGPT’s Bing backed index typically takes 2 to 6 weeks. Google AI Overviews follow organic rankings, so they move on SEO timelines.
At 90 days, expect completed technical fixes, deployed schema, reconciled listings, a defined volume of published pages, a documented prompt baseline, and early Perplexity citations. Outcome metrics may still be flat, and that is acceptable if output is strong.
At 180 days, expect citation rate movement on ChatGPT and Copilot, measurable AI referral sessions in GA4, and Copilot citations appearing in Bing Webmaster Tools. Flat citation rate at six months with real output shipped means the strategy is wrong. Flat citation rate with no output shipped means you bought reporting.
At 365 days, expect share of voice gains against named competitors, AI referral traffic converting at a rate you can put next to other channels, and a content library that keeps earning citations without new spend. If you cannot draw a line from the retainer to revenue by month twelve, renew only with a rewritten scope.
What reporting red flags mean you are being managed, not served?
Six patterns worth catching early. Proprietary scores with no methodology, where “AI Visibility Index up 34 percent” cannot be reproduced or benchmarked. Missing denominators, where citations are counted but the prompt set is never disclosed. Prompt sets that quietly change month to month, which makes trend lines meaningless and is the most common form of soft manipulation. Screenshots as evidence, since a single ChatGPT screenshot proves the engine said something once, to one user, in one session, and answers vary. Vanity impression counts borrowed from PR, the AEO cousin of advertising value equivalency, a metric serious PR measurement abandoned years ago. And traffic charts with no conversion data, because AI referrals are valuable specifically for intent quality, so an agency reporting sessions while ignoring conversion is avoiding the harder number.
None of these are proof of bad faith. All of them are proof you need first party verification, which you have for free.
How do you separate the agency’s work from everything else?
Attribution is genuinely hard here, so use structural controls rather than pretending precision. Freeze a prompt set at kickoff, in writing, and refuse mid contract changes without documenting both versions. Hold a control group where practical: leave a segment of pages or a product line untouched for two quarters and compare. Track output and outcome side by side, because if published pages are the leading indicator, you can see three months ahead whether outcomes are coming. And separate paid from earned effects, since a concurrent PPC push or a product launch moves brand queries and therefore AI mentions independently of AEO work.
The most useful single practice is deceptively simple: ask the agency to name, at the start of each quarter, the specific prompts they expect to win and by when. Predictions made in advance are the only accountability mechanism that survives a discipline without standardized metrics. Buyers assessing whether a scope even contains enough production to move these numbers should compare it against the deliverables a full AI visibility engagement covers, because measurement disputes are usually production disputes wearing a disguise.
What should the monthly review meeting actually cover?
Six items, in this order, and the order matters because it puts input before outcome and prevents an hour of dashboard theater.
Start with output shipped against output promised: pages published, schema deployed, listings corrected, placements earned, each with links. If promised counts were missed, that conversation happens before anyone looks at a chart. Second, the frozen prompt set results, prompt by prompt, showing who was named in each answer including competitors. Third, the free first party reports pulled live in the meeting: Bing Webmaster Tools AI Performance, Search Console, and the GA4 AI Assistant channel, opened on screen rather than pasted into slides. Fourth, your own manual spot checks, which you ran independently, compared against the agency’s numbers. Fifth, what changed on the engine side, since model updates and index refreshes move entire categories and an agency tracking the space should be able to explain sector wide swings. Sixth, next quarter’s named prompt predictions.
That last item is the accountability hinge. An agency willing to write down “we expect to win these six prompts by November” is confident in its process. An agency that resists making predictions is telling you something useful, and it is worth listening to.
Keep the meeting to forty five minutes. Reporting expands to fill whatever time it is given, and long reviews correlate with thin output far more often than with complex results.
FAQ
What is a good citation rate for an AEO campaign?
It depends on category competitiveness, but the trend matters more than the absolute. Moving from 5 percent to 20 or 25 percent of tracked prompts over two quarters is strong performance in most B2B and professional service categories. In crowded consumer categories, single digit gains can still be valuable. Any agency quoting a universal target percentage without asking about your category is guessing.
Which AEO reports can I check myself for free?
Three. Bing Webmaster Tools includes an AI Performance report showing Copilot citations of your pages. Google Search Console shows impressions including AI Overview surfaces. GA4’s AI Assistant default channel group, added May 2026, tracks referral sessions from ChatGPT, Perplexity, Gemini, and other assistants. Together they cover a large share of what a paid dashboard reports, at no cost and with no intermediary.
How often should an AEO agency report?
Monthly for full reporting, weekly for citation tracking on the tracked prompt set. Monthly is the right cadence for trend analysis and competitive benchmarking, since engine indexes refresh on multi week cycles and weekly outcome reporting mostly reports noise. Weekly prompt logs are still worth keeping because they catch sudden citation losses early.
Should I expect AI traffic to replace organic traffic?
No, and an agency promising that is misleading you. AI referral volume is small relative to organic for most sites, but converts far better because the visitor arrives having already had their question answered and their options narrowed. Measure AI channels on conversion rate and revenue per session, not raw sessions, or you will undervalue the channel and misread the reports.
What if my citations dropped after the agency started?
Investigate three causes before assigning blame. Engine side changes, since index refreshes and model updates shift citation patterns industry wide and a drop affecting your whole category is not agency caused. Technical regressions, such as a site migration or new JavaScript rendering that blocked crawlers. And prompt set drift, where the questions being tracked changed. Ask for the raw prompt logs from before and after, and check Bing Webmaster Tools independently.
Can I measure AEO performance without any paid tools?
Yes, adequately. Combine Bing Webmaster Tools, Google Search Console, and GA4 with a monthly manual run of twenty real buyer questions across ChatGPT, Perplexity, Gemini, and Copilot, logged in a spreadsheet with the date, the answer, and every brand named. That costs about an hour a month and produces a defensible record. Paid tools buy scale and automation, not fundamentally better truth.
The measurement problem in AEO is not that good metrics do not exist, it is that the free, verifiable ones make weak retainers obvious and few agencies volunteer them. Log into Bing Webmaster Tools, open GA4, freeze a prompt set, and run ten questions yourself this month. Whatever the next report says, you will know whether it is true. Get a baseline no one else prepared for you: claim your free AI visibility audit and walk into your next agency review with independent numbers in hand.
Tagged