
Why is prompt tracking the default KPI for AI search?
Promptwatch launched in April 2025 as a platform for tracking prompts. So did most of its competitors. The category logic is familiar: pick a set of queries relevant to your brand, run them daily through ChatGPT or Perplexity, and record whether your domain appears in the response.
It works. Run the same prompts over several months, and you get a real picture of average citation rates across a query set you chose. Agencies, like Seeders, use it for client reporting. Enterprise SEO teams use it for brand monitoring. It answers one legitimate question: Is our content showing up when AI answers questions about our space?
The question it answers is not the question that determines whether AI search is using your content with real users.
The simulated prompts are ones you chose. Real users are sending queries you didn't choose. The AI's response to a real query depends on that user's account history, geographic location, memory state, and real-time web browsing behavior.
None of that is in a batch test. Gijs's description of his own product ("synthetic simulation") is the most useful thing a tool founder has said to me about the GEO measurement category all year.

Which AI crawler is actually worth tracking?
Gijs maps the crawlers AI platforms send into three functional categories. Training agents build and periodically refresh the base knowledge of the underlying model. A miscellaneous infrastructure category covers scheduled crawl operations. The type worth tracking: citation agents.
A citation agent fires when a real user sends a query the AI cannot answer from memory. The model doesn't know. It goes browsing.
The crawler that lands on your server log at that moment is there because a live human query triggered a real-time web search. Not in a scheduled cycle. Not in a model refresh.
The trigger is a real user query.
Citation agents show up in CDN logs. Promptwatch reads those logs for its customers and maps the crawler signatures to agent types.
Most teams running prompt trackers have never opened their CDN logs to look for this pattern. The two data sources sit in separate places in most organizations, and nobody's connecting them.
What does 0.1% CTR mean for an AI search strategy?
The 0.1% click-through rate comes from combining two data sources: referral traffic from AI platforms visible in site analytics, and CDN log data showing how many citation agent visits those same time windows represented.
One in a thousand AI search citation events produces a click to the underlying website.
Organic search has always treated visibility and traffic as roughly equivalent: rank higher, receive more clicks. AI search breaks that equation.
You can be cited in thousands of responses and receive almost no referral traffic. I've seen this in client data, and it's the number that consistently doesn't land until someone puts a percentage on it. 0.1% makes the argument for you.
This isn't new territory either. Google's AI Overviews, featured snippets, and zero-click trends have been eroding CTR from traditional visibility for years.
AI answers take it further: the interface is conversational, links appear embedded in generated prose, and there's no equivalent of a traditional SERP where ranked links carry an implicit social proof signal.
Conductor's enterprise benchmarks show 87.4% of AI referral traffic comes from ChatGPT across their customer dataset, and that traffic converts at measurably different rates by platform.
The problem is that optimizing for citations without measuring whether those citations come from real user queries is measuring the wrong layer.
Ready to start scaling your business?

What do CDN logs tell you that prompt tracking can't?
CDN data differs from prompt-tracking data in kind, not just degree.
Prompt tracking tells you what might appear in a synthetic test designed around queries you selected. CDN citation agent visits tell you what happened when real users triggered real-time AI web browsing.
Gijs's team uses aggregated, anonymized CDN data across their customer base to identify patterns batch prompt testing can't surface: which types of content pages attract citation agents most frequently, whether page speed or word count correlates with citation agent visits, what structural patterns distinguish pages AI systems actively retrieve from pages they pass over.
That's actionable in a way synthetic testing isn't. The data is deterministic. A citation agent visited, or it didn't.
The harder route is getting the data. Not every IT manager will hand over log files. Some Promptwatch competitors have stayed in the pure-play prompt tracking lane for that reason. Gijs acknowledges the friction. His team decided the data was worth it. The 0.1% number is the argument they use.

How do you start optimizing for AI search without drowning in variables?
Prompt tracking surfaces a practical trap for teams that try to model everything: too many variables, not enough signal.
You can vary prompts by persona (simulate different marketing profiles), location (how does the response differ for a user in Amsterdam versus Atlanta?), and platform (ChatGPT versus Perplexity versus Claude).
Each axis multiplies combinations. A team that tries to cover all of them before producing an insight produces no insight. I see this in client conversations constantly: the first GEO audit becomes a scoping exercise that never ships.
Gijs's starting advice contradicts what you'd expect from a company that charges per tracked prompt:
Start from what you already know. Take the keywords driving your paid search and SEO. Map ten prompts to those topics. Run them. Look at which sources the AI cites. Find the citation gap. Build content to close it.
Ten prompts is a tractable test. The first result arrives within weeks. A thousand prompts is a data management project that delays the first insight by months. Persona layering, geo-variation, and container-based prompt grouping all matter once you have a baseline. Before the baseline, they're noise.
What is the AI search industry getting wrong about measurement?
The AI search optimization industry has converged on prompt response tracking as the default KPI. The category name for the discipline (Generative Engine Optimization) emerged because marketers needed a label.
The tools that followed mostly measure the same thing: whether an AI system mentions your brand in responses to prompts you curated.
That data is real. The structural problem is treating it as the primary signal when the measurement is probabilistic, and the prompts are self-selected.
Citation agents in CDN logs are not probabilistic. They fire on real queries from real users. They appear whether or not you're running a prompt tracker. They tell you what no batch
simulation can: whether actual users, right now, are triggering AI browsers that retrieve your pages.
The 0.1% CTR is the tell. If AI search citations barely convert to traffic, the metric that matters is whether citation agents are visiting your content to answer real queries. Not whether your brand appears in ten responses to prompts you wrote yourself.
Most AI search optimization is happening at the simulation layer. The hard measurement layer sits in server logs; most teams have never opened them for this purpose.
Ready to start scaling your business?
What does ChatGPT advertising change for AI search visibility?
At BrightonSEO in April 2026, Gijs predicted that an advertising model around ChatGPT was coming. The data on query intent would become more accessible. Prompt selection would get easier because advertisers would have visibility into which topics drive high query volume.
By the time he said it, the prediction had already arrived. ChatGPT advertising launched on February 9, 2026. The self-serve Ads Manager opened to all US businesses on May 5, 2026. OpenAI is targeting $2.5 billion in ad revenue in 2026.
For AI search optimization, this matters in one specific way: advertising data makes the prompt landscape more legible. A standard ad platform publishes performance data on topic areas, query categories, and audience reach. That narrows the problem of selecting synthetic prompts. You're no longer guessing which queries matter.
The measurement gap between synthetic testing and real-world CDN data doesn't close due to advertising. CDN citation agent logs will still tell you something that prompt trackers can't. But the information environment improves, and the industry will have fewer excuses for not knowing which prompts to track.
The domain migration that taught Gijs SEO
Promptwatch launched on promptwatch.io. The team bought promptwatch.com early, reasoning that the .com was correct for a global product.
The problem: search results for the brand name returned the .io domain at the top. For months.
After raising €1.2M from Arches Capital, the team tracked down the owner of the .io domain. The transaction, as Gijs tells it, involved traveling to Romania. He was jokingly saying that it almost felt that it felt like a cash transaction from a car boot 😉
Then came the migration: redirects from .io to .com, debugging redirect chains, waiting for re-indexing. "I wish we'd asked Seeders to solve this for us," Gijs says.
The initial mistake was not checking the SERP for their own brand name before choosing the primary domain.
The co-founder of an AI search visibility platform missed the most basic brand visibility check. He learned organic search the way most founders do: after launch, from a problem that was already live.
Resources
- Promptwatch: AI search visibility platform with CDN log analytics, prompt tracking, and content recommendations
- OpenAI: Our approach to advertising and expanding access to ChatGPT
- Silicon Canals: Arches Capital backs Promptwatch in €1.2M round
Listen to the full conversation
The full SEO Cast interview with Gijs de Groot runs in Dutch and covers his BrightonSEO talk on CDN analytics, the three-agent taxonomy in detail, how Promptwatch uses personas and geo-variation for prompt tracking, and the domain migration he is still not proud of.
The 0.1% number comes up about halfway through. Worth the 20 minutes if you want to hear a prompt tracking founder explain exactly why his own category's default metric is a simulation.















