The Medicare AIVI™ (AI Visibility Index) family scores named Medicare plans and carriers on how they appear in AI answer surfaces — not on how they rank in a search engine, and not on anything they pay for. This document is versioned because the scoring model will change as the surfaces change, and a benchmark whose rules move silently is not a benchmark. Every published index states which version of this document produced it.
Key takeaways
- Medicare AIVI™ reports two measures separately: Share of Voice (reach, against two denominators) and QVI™ (Quality Visibility Index, per-mention quality). Unlike the Insurance AIVI™, QVI™ is deliberately volume-neutral — it does not blend reach into the score.
- QVI™ = 8 × mean Position Score + 7 × mean Recommendation Score + 25 × own-domain citation rate, scaled 0–100.
- Commercial share and total entity share are two distinct, deliberately separate denominators — 59,055 scored commercial rows and 229,372 total extraction rows, respectively.
- This is Methodology v1.0, effective 2026-10-01; a correction process (Section 8) is open to any organization named in a published index, client or not.
1. Document version and status
This is Medicare AIVI™ Methodology v1.0, effective 2026-10-01. It is the controlling document for the Q3 2026 Medicare AI Visibility Benchmark and for every derived analysis published from it, including the UnitedHealthcare analysis and the Kaiser Permanente analysis.
| Version | Effective | Governs | Change |
|---|---|---|---|
| v1.0 | 2026-10-01 | Q3 2026 Medicare Benchmark | Initial public release for the Medicare vertical. Formalizes the QVI™ scoring model, the two-denominator (Commercial / Total) reporting convention, and the correction policy in Section 8. |
Two rules govern this document going forward. First, a change to the QVI™ formula, to the attribution rules, or to the surface set is a major version bump, and indices scored under different major versions are not directly comparable. Second, corrections to this document are additive: superseded text is recorded in the version history above rather than deleted.
2. What Medicare AIVI™ measures
Medicare AIVI™ answers one question: when a Medicare beneficiary asks an AI assistant a question in this category, how present, how early, and how favorably does this plan or carrier appear in the answer? It is a measure of the answer layer specifically. It is not a search ranking, not a Star Rating, not an enrollment estimate, and not a brand-health survey.
The index deliberately reports two measures separately rather than blending them into one score:
- Share of Voice (reach). How much of the conversation belongs to a brand, reported against two denominators: Commercial share (against the 59,055 scored commercial rows) and Total entity share (against the 229,372 total extraction rows). Neither is a direct estimate of demand, attention, traffic, enrollment, or revenue.
- QVI™ (quality). How strongly each individual appearance performs, independent of volume. This is the deliberate departure from the Insurance AIVI™, which blends Share of Model into its composite at 35% weight. A Medicare QVI™ figure is therefore not directly comparable to an Insurance AIVI™ figure.
Four things are deliberately outside the index:
- Paid placements. Sponsored and ad-click results are excluded from the score entirely. A plan cannot buy QVI™ points.
- Traffic and enrollment. The index measures presence in the answer, not what a reader does next or whether they enroll.
- Star Ratings and plan benefit design. QVI™ scores how an AI answer treats a brand, not CMS Star Ratings, premiums, or benefit structure.
- Private or client-supplied data. Every figure in a published index is derived from public model outputs. No client relationship contributes to any score.
3. Study universe and prompt set
An index is only as defensible as the question set that produced it, so the construction is fixed before any responses are collected and locked before scoring.
| Parameter | Q3 2026 value |
|---|---|
| Brands tracked | 121 (115 with any recorded visibility) |
| Base query templates | 60 local + 90 national-only = 150 stratified queries |
| Geographic contexts | 11 — National plus 10 states (Florida, Texas, California, Pennsylvania, Ohio, New York, Illinois, North Carolina, Georgia, Arizona) |
| Surfaces | 7 — ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, Microsoft Copilot, Bing Copilot |
| Runs per query/surface combination | 3 |
| Prompt-surface observations (designed sample) | 14,490 (90 × 3 × 7, plus 10 states × 60 × 3 × 7) |
| Populated responses | 14,405 (99.41% of 14,490) |
| Source extraction rows | 229,372 |
| Scored commercial-brand mentions | 59,055 (25.75% of source rows) |
| Collection window | 2026-06-29 to 2026-07-03 |
Two denominators, deliberately kept apart. Use 14,490 for overall visibility/availability measures — a non-render counts as zero brand visibility on that query, for every brand. Use 14,405 for analysis of what is inside generated content: entity extraction, the Total Visibility Universe, Share of Voice, and everything built on it. Both are reported so neither figure is read as the only correct one.
Google AI Overview non-renders are a real outcome, not missing data. AI Overview rendered for 95.89% of its allocated prompts (1,985 of 2,070); the other 85 are observed non-render outcomes, carried as zero-visibility in the 14,490-denominator figures.
Response coverage. Computed against the canonical 115-brand list from sealed per-surface data: 11,638 of 14,405 populated responses (80.79%) name at least one scored commercial brand; against the full 14,490 designed sample with non-renders folded in as zero-visibility, the figure is 80.32%.
Entity classification and rollups. Name variants are rolled up under one canonical brand — for example, state Blue Cross Blue Shield affiliates (Illinois, Texas, Georgia, and unbranded “BCBS” mentions) roll up to Blue Cross Blue Shield Medicare. Brand references are separated from generic Medicare vocabulary, and consolidated tool and phone references (Medicare Plan Finder, 1-800-MEDICARE) are each treated as a single entity rather than counted as separate variants.
Reliability threshold. QVI™, mean recommendation score, own-domain citation rate, top-3 rate, and mean position score are published only for brands with at least 100 mentions — a reporting rule, not a statistical confidence guarantee. Mentions and Commercial share are complete for every brand regardless of this threshold; small-volume QVI™ scores below it can be extreme and should not be read as robust leaders.
4. Attribution rules
Row-based, not response-based
59,055 scored rows underlie every commercial-share and position figure in this report. QVI™ and position scores remain row-based; repeated mentions of one brand within a response can occupy several positions, and a brand's mention count is a count of scored rows, not of distinct responses it appeared in.
Scored-subset marking (†)
QVI™, mean recommendation score, own-domain citation rate, top-3 rate, and mean position score are computed on the subset of rows that received full per-mention scoring. For brands marked † on the leaderboard, these five columns reflect that scored subset rather than the brand's complete mention count; mentions and Commercial share remain complete for every brand regardless.
Recommendation coverage
All 23,715 rows from the four API configurations (ChatGPT, Gemini, Claude, Perplexity) have Recommendation Score zero. Nonzero scores occur only on consumer-facing surfaces (Google AI Overviews, Microsoft Copilot, Bing Copilot). QVI™ can therefore reflect surface mix and incomplete recommendation signals as much as brand performance.
5. The QVI™ formula
| Component | Scoring rule | Meaning |
|---|---|---|
| Position Score | clip(6 − ordinal rank, 1, 5) | Higher is better; five points for the first scored row. Repeated brand rows can consume positions. |
| Recommendation Score | Strong = 5; moderate = 3; mention only = 0 | Mean score is not the percentage of answers recommending a brand. |
| Citation Score | Own-domain match = 3; otherwise 0 | Matched against the captured citation corpus. |
| Top-3 flag | Ordinal rank ≤ 3 | Descriptive placement metric; not an additional QVI™ term. |
| Visibility Points | Position + Recommendation + Citation | Maximum 13 per mention; totals depend on volume. |
QVI™ formula. Quality Visibility Index = 8 × mean Position Score + 7 × mean Recommendation Score + 25 × own-domain citation rate. Citation rate enters as a fraction from 0 to 1. The theoretical maximum is 100. QVI™ is a per-mention average, not a volume count. It is deliberately volume-neutral, unlike the Insurance AIVI™ score, so a brand's rate of being answered well can be read separately from how often it is answered at all (Commercial share, Section 3).
A note on the two denominators
Commercial mention share divides a brand's mention rows by 59,055 tracked-brand rows — relative visibility within the tracked commercial set, including carriers and intermediaries. Total extracted-entity share divides by 229,372 total extracted rows — composition of the extraction output, sensitive to vocabulary and entity classification. They answer different questions and are reported side by side rather than collapsed into one.
6. Known limitations
These are the constraints a reader should hold in mind when using a Medicare AIVI™ figure. They are published here rather than buried in a footnote because a benchmark that only lists its strengths is marketing.
Capture quality
Some Microsoft Copilot records contain only prompt/location text, an interface greeting, or an unfinished response indicator. Populated is therefore a storage test, not a substantive-answer test.
Recommendation Score is not uniform across surfaces
All 23,715 API-surface rows have Recommendation Score zero; this component is not comparable as a uniformly measured cross-surface preference signal. A brand's QVI™ can be shaped by which surfaces it is captured on as much as by how it is treated within them.
Model responses are non-deterministic
Across 7,363 observed brand/query/surface combinations, 52.0% are not present in all three reported runs, and 40.6% show a rank swing of at least two positions when consistently present. Single-answer visibility claims are not a stable baseline; repeated observations support a more useful descriptive one.
Scope
This benchmark does not validate extraction accuracy beyond the classification rules stated here, resolve every plan alias, or rescore recommendations. Prominence, Product Ownership, Geography, Query Ownership, Co-Mention Network, and Cross-Surface Agreement analyses require response-level, product-, geography-, or query-tagged brand data beyond this benchmark's current scope and are not included.
An edition is a snapshot
Data were captured 2026-06-29 to 2026-07-03, crossing the Q2/Q3 boundary; the workbook retains the Q3 study label as a convention, not a claim that every observation fell inside Q3. Surfaces update their retrieval stacks and system prompts continuously and without notice, so an edition describes the answer layer during its collection window and nothing outside it.
Category scope
Scores are normalized within the Medicare study cohort. A QVI™ figure here and an Insurance AIVI™ figure must not be compared — the two indices weight volume differently by design (Section 2).
Independence
Brainpan.AI publishes AIVI™ indices independently. Being named in an index — favorably or otherwise — reflects observed public model output and nothing else. Brands cannot pay for inclusion, placement, or removal, and no client relationship has ever influenced a published figure.
7. Correction policy
Any organization named in a published Medicare AIVI™ index can ask for a figure to be corrected. This is a standing commitment, not a courtesy, and it applies whether or not the organization is a Brainpan.AI client.
What qualifies
- Factual error — a miscount, a transcription error, a brand misattributed to the wrong corporate entity, or a figure that does not reconcile with the published data.
- Methodological misapplication — a rule in this document applied incorrectly to the organization's data, for example a paid placement counted as organic.
- Entity error — the wrong legal entity, subsidiary, or brand family associated with a set of mentions.
What does not qualify
Disagreement with the methodology itself is not a correction request; it is a methodology comment, and it is welcome as such, but it is answered in a version of this document rather than by changing a score. A request to remove an unflattering but accurately measured figure will be declined and the exchange noted.
How to file
Send the request to kwalsh@brainpan.ai with the subject line AIVI correction request, naming the index edition, the specific figure, and the basis for the challenge.
| Stage | Target | What happens |
|---|---|---|
| Acknowledgement | 2 business days | Written confirmation of receipt, naming the figure under review. |
| Determination | 10 business days | The underlying response set is re-examined against this document and a written finding is issued: upheld, corrected, or partially corrected, with reasoning. |
| Publication | 5 business days from determination | Any corrected figure is republished across the page and any derived analysis, and logged in the correction record below. |
Correction record
No corrections have been requested or issued against any published Medicare AIVI™ index as of 2026-10-01. This record is maintained here and updated on determination, whether the outcome is a correction or an upheld figure.
8. References and related documents
- Q3 2026 Medicare AI Visibility Benchmark — the index this version governs.
- Medicare AIVI™ Benchmark — the series hub.
- AIVI™ Insurance Benchmark — the Insurance edition of this same measurement discipline.
- Multi-Engine AIVI™ Benchmarks — every industry the series covers, compared side by side.
- UnitedHealthcare Medicare analysis — a worked application of the scoring model to the commercial-share leader.
- Kaiser Permanente analysis — a second worked application, isolating the citation-authority term of the score.
- Glossary — canonical definitions for every metric named here.
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024) — the foundational academic treatment this prompt design follows.
- Schema.org Dataset — the vocabulary used to declare each published index as a machine-readable dataset.
Get the full index
The Q3 2026 Medicare AI Visibility Index scores all 115 visible brands on the components described above, with per-surface detail and the non-commercial layer. $3,500, one-time.
Single-organization commercial license / One-time payment — delivered within one business day.
Frequently Asked Questions
Why is QVI™ different from the Insurance AIVI™ score?
The Insurance AIVI™ is a weighted composite that blends Share of Model (reach) into the score at 35%. QVI™ is deliberately volume-neutral: it scores only position, recommendation, and citation performance, and reports reach (Share of Voice) as a separate measure. This means a brand can rank highly on QVI™ with very little mention volume — Medicare.org leads QVI™ at 48.05 while holding just 1.64% of commercial mentions. A QVI™ figure and an AIVI™ figure from the Insurance benchmark are not on the same scale and must not be compared directly.
How do you evaluate metrics accuracy for AI brand visibility?
Accuracy is enforced in two stages. Before an index is published, every scored figure must reconcile with the underlying response set. After publication, any organization named in an index can file a correction request under the policy in Section 7, which triggers a written re-examination against three defined error categories: a factual error, a methodological misapplication, or an entity error. Every determination — upheld or corrected — is logged in the correction record, so accuracy is checked against a written standard rather than asserted.

