The AIVI™ (AI Visibility Index) scores a brand on how it appears in organic AI answers — not on how it ranks in a search engine, and not on anything it pays for. This document is versioned because the scoring model will change as the engines change, and a benchmark whose rules move silently is not a benchmark. Every published index states which version of this document produced it.
1. Document version and status
This is AIVI™ Methodology v1.0, effective 2026-07-31. It is the controlling document for the Q2 2026 Insurance AI Visibility Benchmark and for every derived analysis published from it, including the Hartford teardown.
| Version | Effective | Governs | Change |
|---|---|---|---|
| v1.0 | 2026-07-31 | Q2 2026 Insurance Benchmark | Initial public release. Formalizes the scoring model already applied to the Q2 2026 index and adds the correction policy in Section 8. |
Two rules govern this document going forward. First, a change to any component weight, to the attribution rules, or to the engine set is a major version bump, and indices scored under different major versions are not directly comparable — where a comparison is drawn anyway, the discontinuity is stated on the page that draws it. Second, corrections to this document are additive: superseded text is recorded in the version history above rather than deleted, so a reader can always reconstruct the rules that were in force when a given index was published.
2. What AIVI measures
AIVI answers one question: when a buyer asks an AI assistant a question in this category, how present, how early, and how favorably does this brand appear in the answer? It is a measure of the answer layer specifically. It is not a search ranking, not a traffic estimate, and not a brand-health survey.
Four things are deliberately outside the index, and stating them plainly is part of the method:
- Paid placements. Sponsored and ad-click results are quarantined out of the score entirely and disclosed in a separate supplement. A brand cannot buy AIVI points.
- Traffic and conversion. AIVI measures presence in the answer, not what a reader does next. Referral behavior is a separate measurement problem with separate instrumentation.
- Sentiment. The index scores whether a mention functions as an active recommendation, which is a structural judgement about the answer, not a tonal one about the brand.
- Private or client-supplied data. Every figure in a published AIVI index is derived from public model outputs. No client relationship, analytics account, or CRM export contributes to any score. This is what makes the index publishable and independently checkable.
3. Study universe and prompt set
An index is only as defensible as the question set that produced it, so the construction is fixed before any responses are collected and locked before scoring.
| Parameter | Q2 2026 value |
|---|---|
| Brands in the study cohort | 134 |
| Stratified prompts | 300 |
| Engines | 5 — ChatGPT, Gemini, Claude, Perplexity, Copilot |
| Model responses collected | 1,500 (300 prompts × 5 engines) |
| Organic brand mentions observed | 4,328 |
| Organic citations observed | 1,742 |
Stratification. The 300 prompts are not a keyword list. They are drawn across three axes at once — query type (definitional, comparative, recommendation-seeking, troubleshooting), funnel stage (early research through active shortlisting), and product line (personal auto, home, life, commercial, specialty) — so that no single intent or product line can dominate the mention count. This matters because the engines behave differently by intent: a definitional prompt and a “which should I buy” prompt do not surface the same brands, and an index built only on the second would flatter incumbents.
Prompt phrasing follows buyer language, not marketer language. Prompts are written the way a person actually asks an assistant, which is the operating premise of generative engine optimization as a discipline — see Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024), the foundational academic treatment of how generative engines select and surface sources.
Cohort selection. The 134-brand cohort is the set of carriers that appear in the category at all: every organization named organically at least once across the prompt set, plus the licensed national and regional carriers that the prompt set was designed to give a fair chance of surfacing. A brand that never appears is still in the cohort and still scores — it scores zero on presence, which is itself the finding. Regulatory context for the carrier universe is taken from the National Association of Insurance Commissioners.
Collection window. All 1,500 responses for a given edition are collected inside a single bounded window, so an edition is a snapshot rather than a rolling average. The window is disclosed on each published index.
4. Attribution rules
Three rules turn a wall of model output into countable events. They exist to stop the same answer being counted twice in a brand’s favor.
Organic-only
Only mentions the model surfaced on its own are counted. Sponsored placements, ad units, and paid comparison slots are identified at collection time and excluded from the index. They are not discarded — they are reported in a separate paid supplement, so the two populations never mix inside a score.
First-brand-mention
Within a single model response, a brand is counted once, at its first appearance. A response that names a carrier once in a list and then again in a closing summary contributes one mention, not two. Without this rule, verbose answers and brands with long product-name tails would accumulate score purely from repetition.
Citation origin classification
Every citation URL is classified by its relationship to the brand it supports:
- 1st-party — the brand’s own domains and properties.
- 2nd-party — affiliated marketplaces, quote aggregators, and comparison sites with a commercial relationship to the brand.
- 3rd-party — independent editorial, regulatory, and reference sources.
The scored citation component counts all three, because a sourced mention reaches the reader regardless of who published the source. The stricter lenses are reported separately in Section 6, which is what prevents a brand from manufacturing authority by citing itself.
5. The five weighted components
AIVI is a weighted composite, normalized so the highest-scoring brand in the study cohort scores 100. Every other brand is expressed relative to that leader, which is why an AIVI score is a competitive position rather than an absolute quantity.
| Component | Weight | What it measures | API field |
|---|---|---|---|
| Share of Model | 35% | The brand’s share of all organic mentions in the category. The single largest component, because being named at all is the precondition for everything else. | shareOfVoicePct |
| Recommendation strength | 25% | The share of the brand’s mentions that function as an active recommendation rather than neutral inclusion in a list. This is the component that separates being listed from being chosen. | recommendationRatePct |
| Position weighting | 20% | Where in the answer the brand appears. Earlier placement carries more weight, on the same logic as search position: readers do not finish long answers. | — |
| Citation Influence | 15% | How often the brand’s mentions are backed by a linked source, counting all citation origins. Only two of the five engines emit sources at all, which caps this component structurally — see Section 7. | citations |
| Top-3 rate | 5% | The share of the brand’s mentions landing in the first three positions of a response. Weighted lightly on purpose: it is conditional on already being mentioned, so it rewards brands with small, well-placed footprints more than it should if weighted heavily. | top3RatePct |
A note on two names for one measure
The 35% component is called Share of Model in this document and in the
glossary, because that is its scored,
operational name. The published JSON emits it as shareOfVoicePct, the general-industry
term for the same measure. They are the same number. The API field names are frozen for the life
of a published edition so that anyone who built against them keeps working; the prose uses the
canonical name.
The 5% weight on Top-3 rate is the most common source of confusion about an AIVI score, and the Hartford result is the clean illustration: a carrier can hold the best top-three rate in the entire study cohort and still rank eighth, because top-three rate is conditional on a mention that Share of Model says is rare. The Hartford teardown works through that arithmetic in full.
6. Companion metrics
Two citation lenses are computed and published but deliberately kept out of the weighted score. They are reported alongside every index.
- Citation Authority — third-party citations only. A brand’s own pages confer no authority, so this cannot be inflated by self-publishing. It is the answer-layer analog of earned external links.
- Citation Independence — the strictest lens: 1st-party and 2nd-party sources are both removed, isolating genuinely independent editorial endorsement.
These sit outside the score for a specific reason. Both are strongly influenced by how much independent press a category happens to generate in a quarter, which is not a property of the brand’s AI visibility work. Folding them into the composite would import that volatility into every carrier’s score. Reporting them separately keeps the signal without destabilizing the index.
7. Known limitations
These are the constraints a reader should hold in mind when using an AIVI figure. They are published here rather than buried in a footnote because a benchmark that only lists its strengths is marketing.
Three of five engines emit no sources
In the Q2 2026 edition, all 1,742 organic citations came from Perplexity and Copilot. ChatGPT, Gemini, and Claude produced 2,549 organic mentions between them and cited zero sources in the collected responses. Perplexity documents inline numbered citations as standard behavior in its own product documentation, and OpenAI documents inline citations for ChatGPT search specifically. The practical consequence for the index is that the 15% citation component is measured across a minority of the answer surface, and a brand invisible in the three non-citing engines cannot recover that ground through citations.
Model responses are non-deterministic
The same prompt does not return the same answer twice. AIVI addresses this through prompt-set breadth rather than repeat sampling: 300 stratified prompts across 5 engines is a wide enough base that a single brand’s score is not hostage to any one response. It does mean small inter-edition movements — roughly a point or two of AIVI — should be read as noise rather than trend. Movements are reported, but only multi-point shifts are characterized as directional.
An edition is a snapshot
Engines update their retrieval stacks, indexes, and system prompts continuously and without notice. An AIVI edition describes the answer layer during its collection window and nothing outside it. Where an engine materially changes behavior between editions, that change is disclosed on the affected index rather than smoothed over.
Category scope
Scores are normalized within a study cohort. A 63.4 in insurance and a 63.4 in a future category are not the same achievement, and AIVI scores must not be compared across categories.
Independence
Brainpan.AI publishes AIVI indices independently. Being named in an index — favorably or otherwise — reflects observed public model output and nothing else. Brands cannot pay for inclusion, placement, or removal, and no client relationship has ever influenced a published figure.
8. Correction policy
Any organization named in a published AIVI index can ask for a figure to be corrected. This is a standing commitment, not a courtesy, and it applies whether or not the organization is a Brainpan.AI client.
What qualifies
A correction request should identify a specific published figure and state why it is wrong. Three categories qualify:
- Factual error — a miscount, a transcription error, a brand misattributed to the wrong corporate entity, or a figure that does not reconcile with the published open data.
- Methodological misapplication — a rule in this document applied incorrectly to the organization’s data, for example a paid placement counted as organic.
- Entity error — the wrong legal entity, subsidiary, or brand family associated with a set of mentions.
What does not qualify
Disagreement with the methodology itself is not a correction request; it is a methodology comment, and it is welcome as such, but it is answered in a version of this document rather than by changing a score. Likewise, a request to remove an unflattering but accurately measured figure will be declined and the exchange noted. The published figures are what the models actually said.
How to file
Send the request to kwalsh@brainpan.ai with the subject line AIVI correction request, naming the index edition, the specific figure, and the basis for the challenge. Supporting evidence — model output showing different behavior, corporate-structure documentation, or a reconciliation against the published JSON — speeds the review substantially.
| Stage | Target | What happens |
|---|---|---|
| Acknowledgement | 2 business days | Written confirmation of receipt, naming the figure under review. |
| Determination | 10 business days | The underlying response set is re-examined against this document and a written finding is issued: upheld, corrected, or partially corrected, with reasoning. |
| Publication | 5 business days from determination | Any corrected figure is republished across the page, the open JSON, and any derived analysis, and logged in the correction record below. |
How corrections are published
Four commitments govern what happens after a correction is upheld:
- Nothing is changed silently. A corrected figure carries a dated correction note on the index page itself. The original figure is stated alongside the corrected one.
- Corrections propagate. A change to a scored figure is applied to the index page, the machine-readable JSON endpoints, and every derived analysis that cites it — not to the headline number alone.
- Rank changes are called out. If a correction moves a brand’s position, the movement is stated explicitly rather than left for a reader to notice.
- Declined requests are recorded too. Where a challenge is reviewed and the original figure upheld, that determination is logged, so the record shows scrutiny survived and not merely that no one asked.
Correction record
No corrections have been requested or issued against any published AIVI index as of 2026-07-31. This record is maintained here and updated on determination, whether the outcome is a correction or an upheld figure.
9. References and open data
The strongest form of methodological transparency is letting a reader recompute the result. Q2 2026 is published as open data under CC BY 4.0: the top-10 leaderboard, an OpenAPI description, and a per-brand record for every named carrier. Reuse it freely, including inside generated answers, with attribution to Brainpan.AI and a link to the source page.
Standards and primary sources
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD’24). DOI 10.1145/3637528.3671900.
- Schema.org Dataset — the vocabulary used to declare each published index as a machine-readable dataset.
- Google Search Central, Intro to How Structured Data Markup Works — the structured-data conventions the published pages follow.
- The /llms.txt file — the proposal this site implements at /llms.txt.
- Perplexity, How does Perplexity work? — first-party documentation of inline numbered citations.
- OpenAI, ChatGPT search — first-party documentation of inline citation behavior in search-backed responses.
- National Association of Insurance Commissioners — regulatory reference for the US carrier universe.
- NIST AI Risk Management Framework — the risk-management vocabulary this methodology’s limitations section is written against.
Related documents
- Q2 2026 Insurance AI Visibility Benchmark — the index this version governs.
- The Hartford teardown — a worked application of the scoring model to a single carrier.
- Glossary — canonical definitions for every metric named here.
- Data handling and procurement — what data Brainpan.AI does and does not process, for legal and procurement review.
Get the full index
The Q2 2026 Insurance AI Visibility Index scores all 134 evaluated carriers on the five components described above, with the category breakdowns and the per-engine splits. $3,500, one-time.

