Answer first: Answer engines often disagree on who belongs in a category shortlist and how those brands are framed. In a locked Lab sample of unbranded category prompts, pairwise shortlist overlap between major assistants stayed partial—not identical. Brand teams should measure engine disagreement as a first-class metric, not noise.
What we measured
World Brand Lab ran a methodological sample (not a census) designed to mirror how operating teams should instrument BrandAEO:
- 48 locked unbranded category prompts (discovery + best-of phrasing)
- 3 assistants sampled on the same wording the same week
- Peer set of 6 brands pre-registered before scoring
- Scores: inclusion (mention), first-mention on best-of, and trust-framed language among mentions
How much do engines agree?
On the same 48 prompts, pairwise shortlist overlap (Jaccard on mentioned peer-set brands) looked like this:
| Pair | Overlap (Jaccard) | Read |
|---|---|---|
| ChatGPT ↔ Perplexity | 0.61 | Shared core, different edges |
| ChatGPT ↔ Gemini | 0.54 | Larger framing / source drift |
| Perplexity ↔ Gemini | 0.58 | Citation habits diverge |
What this means: If you only monitor one engine, you will misread category weather. A “win” on one assistant can coexist with absence or harsh framing on another.
Where disagreement concentrates
| Prompt family | Highest disagreement driver | Typical team mistake |
|---|---|---|
| Category discovery | Category nouns / synonyms | Measuring only branded vanity prompts |
| Best-of | First-mention + trust adjectives | Treating SOV as preference |
| Comparison | Citation to docs vs reviews | Fixing homepage copy only |
A worked pattern (illustrative)
On implementation-heavy category prompts in the sample, one assistant repeatedly cited vendor docs; another leaned on roundup reviews with stale feature matrices. Shortlist overlap looked "fine" at the brand-name layer, while framing diverged: "powerful but complex" vs "best fit for mid-market."
Operating translation: do not celebrate name inclusion while the citation graph teaches two different stories. Fix the stale matrix and the docs conflict as separate tickets—then re-measure the band, not a single engine screenshot.
Operating rules for brand teams
- Report a disagreement band, not a single SOV number—min/max mention rate across engines for the same prompt family.
- Separate inclusion from preference on every scoreboard (Share of Voice Is Not Preference).
- Assign repairs by evidence class—owned specs, encyclopedic spine, third-party corroboration—not by which engine embarrassed you in a screenshot.
- Re-run the same 48 next week. One dramatic day is not a strategy.
- Version the instrument—prompt-set ID + peer set + sample week printed on every leadership slide.
Pair the readout with Brand Hub identity hygiene and BrandSight structure when gaps look architectural, not merely editorial.
What disagreement is not
- A reason to abandon measurement
- Proof that "AI is random"
- An excuse to keep rewriting prompts until one engine flatters you
- A substitute for honest peers and locked wording
Disagreement is weather. Your job is instruments and repairs—not mood.
FAQ
Is engine disagreement a bug in measurement?
Usually no. Assistants optimize under different retrieval and safety priors. Disagreement is a signal about which evidence graphs each system trusts.
How large should a prompt sample be?
Enough to cover category, comparison, and best-of families with stable wording—often dozens, not three. Lock it before you celebrate a spike.
Should we optimize for the “worst” engine?
Optimize for the buyer journeys that matter, then watch the band. Chasing a single hostile sample without a locked set creates thrash.
Bottom line
AI search is not one leaderboard. It is a set of partially overlapping shortlists. Measure the overlap. Explain the gaps. Repair the evidence. Then the Tuesday meeting has something real to own.
