Publishers and brands alike are scrambling to show up in AI search responses, and to be described in their preferred ways.
But with new AI optimization tools cropping up faster than you can say AEO, it’s hard to pinpoint which metrics are most important to track and what it even means to measure AI search visibility when there’s so little data or replicability around both prompts and results.
But when publishers and advertisers can’t figure something out, that’s an opportunity for the IAB to step in and – hopefully – provide some clarity.
On Monday, the IAB published “Measuring Visibility in the AI Era,” a set of guidelines, recommendations and important data points for how brands and publishers should be tracking their AI search analytics.
The new guidelines are not, however, a “standard,” Caroline Giegerich, the IAB’s VP of AI, told AdExchanger.
A standard requires stability, she said, and right now marketers are in a “mass transition space.” The ad industry can’t establish AI search standards until AI searches are more consistent and well understood.
The new document is also not a framework – at least, not in name. It was going to be called a framework, Giegerich said, but the IAB made a last-minute decision to change the name “so we don’t have 100 frameworks.”
Right, because 99 is a normal quantity of frameworks.
Keepin’ it neutral
As a cross-industry trade body, the IAB doesn’t recommend any particular AEO tool or vendor over another. (Please. Did you think it was going to be that easy?) Rather, it proposes a hierarchy of metrics to determine how the context and frequency at which a brand or pub’s content is showing up in AI search, which the IAB calls “The 4P’s of AI Visibility”: presence, prominence, portrayal and persuasion.
At the base of the hierarchy is “Presence,” which addresses a straightforward question: How often does one’s brand show up in AI search queries? Or, for publishers, how often is their content cited?
“Prominence” is for determining where in the AI search results one’s product or content appears and whether it’s highlighted in any particular way or is lumped into a list of similar brands or citations.
“Portrayal” focuses on not just whether the brand is present but how it shows up – both in terms of sentiment and accuracy. Accuracy can be broken down into two categories, said Giegerich: hallucinations and factual inaccuracies.
Hallucination rates, or how often an LLM generates inexplicably inaccurate information, have declined, thanks to improvements by the leading AI models.
But factual inaccuracies – which can stem from an AI drawing from outdated or misrepresentative data sets – should still be front of mind for advertisers. That misinformation “would be the thing that would keep me up at night,” said Giegerich, when asked about the area of improvement that marketers should be most focused on.
Lastly, “Persuasion” measures how effectively a recommendation drives site traffic. This can be done by tracking details like the post-citation click-through rate. And Persuasion is especially important for publishers, who are generally losing traffic as AI search engines and search integrations like Google’s AI Overviews do a better job of encapsulating a response without the user needing to click to a new site.
Big appetites
The IAB’s new AI guidance also distinguishes between “directional” and “decision-grade” measurement – i.e., data that’s more theoretical or less well-established (like Googling your brand a few dozen times and noticing whether it appeared) versus data that’s been rigorously tested and includes a larger query volume across more platforms. Only the latter is recommended for determining budget allocation.
Since AI isn’t deterministic, it can – and will – churn out different responses to the same query. To get a sense of the scale, the IAB classifies anything fewer than 50 queries as “exploratory,” a step below even “directional” measurement. Anything less than 50 prompts “cannot meaningfully characterize a category,” per the guidelines.
The not-a-framework also emphasizes the importance of testing different kinds of prompts – “best shoes,” for instance, might generate entirely different results than “best running shoes,” or “x shoe vs. y shoe.”
Brands should pay attention to how often they show up in a variety of prompts so they can see “where there’s opportunities to lean in, if that’s what they want to do,” Giegerich said. (For instance, queries for “women’s basketball shoes” could increase as the WNBA gains traction in media, and a sneaker brand may want to seize on that moment.)
It might seem logical to assume that once the IAB has established official guidance to measure visibility within AI search, the dozens of AEO and GEO platforms will collapse into a few frontrunners. But Giegerich doesn’t see that happening anytime soon.
Since the advertising world is still scrambling like mice in an unnavigable AI maze, all of the vendors are, as Giegerich put it, fighting “to get to the cheese at the end of the tunnel.”