Skip to content
← Back to Blog

AI Search Predictions for 2026: Ten Falsifiable Calls, With Dates

Ten specific, dated, falsifiable predictions about AI search - each grounded in panel measurement rather than opinion - plus a scored review of the predictions we made for 2025.

Scorecard of dated AI search predictions for 2026 with measurement thresholds and verdicts

The rules I am holding myself to

Most prediction pieces are unfalsifiable by design. "AI will transform search" cannot be wrong, so it is not a prediction - it is a mood.

Every call below has three things attached: a specific threshold, a date, and the measurement it rests on. The measurements come from our own tracking - a rolling panel of roughly 1,200 brands across ChatGPT, Claude, Gemini, Perplexity, and Grok, sampled on the schedule described in our measurement methodology. Where a number comes from somewhere else, it is sourced on the AI search statistics page.

I score the previous year's calls first, including the ones I got wrong, because a predictions post that never revisits its own record is entertainment.

Scoring the 2025 predictions

2025 callThreshold setOutcomeVerdict
Zero-click becomes the majority US search outcome>50% by end of 202558%Hit
Community sources overtake publishers in AI citationsReddit + forums > major publishers21.4% vs 11.2%Hit
Brand-owned domains stay a minority of citations<15% of cited sources3.8%Hit, by more than expected
The tool category consolidates through acquisition3+ notable acquisitions in 2025Fewer than 3Miss
Multi-engine tracking becomes standard practiceMajority of new customers tracking 3+ enginesMajority tracked 2 at signupPartial

Two of those deserve comment.

The consolidation miss was not a small error, and I think I know why I got it wrong. I assumed the measurement problem would stabilise enough for a buyer to acquire a durable asset. It has not. Engines ship new tiers, retrieval behaviour changes, and a data pipeline needs continuous work simply to stay accurate - which makes an acquisition a liability rather than a shortcut.

The brand-owned domain call was right directionally and badly calibrated. I predicted under 15%; the actual figure is 3.8%. I underestimated how completely third-party evidence dominates the source pool, and that miscalibration led me to under-advise clients on review platforms and community presence for most of a year.

The ten calls for 2026 and early 2027

1. By 31 December 2026, review platforms and community sources together will exceed 45% of cited sources on commercial B2B software prompts

Currently measured: 46.5% on our project management panel (24.6% review platforms, 21.9% community), and 41-49% across the other B2B software categories we track.

Why: the trend has been monotonic for six consecutive quarters, and the engines have a structural incentive - user-generated evidence is the cheapest available proxy for "does this actually work in practice".

Falsified if: the combined share drops below 45% on our panel in the December measurement.

What to do: stop treating review-platform presence as a reputation chore. In most B2B categories it is now the single largest input to what an assistant says about you.

2. By 30 June 2027, the median B2B SaaS brand will be absent from more than 70% of non-brand category answers

Currently measured: 79% median absence rate across the B2B SaaS brands in our panel on category-defining and comparison prompts, excluding brand-led prompts.

Why: answers name three to five brands out of categories containing dozens. The arithmetic of a short list against a long tail does not permit broad visibility, and the concentration has tightened rather than loosened as engines have grown more confident.

Falsified if: median absence falls below 70%.

What to do: treat category-level visibility as a competitive position to win rather than a hygiene checkbox to complete. The cost of not winning it is modelled in the cost of being invisible in AI search.

3. By 31 December 2026, ungrounded and grounded visibility for the median tracked brand will differ by more than 20 points

Currently measured: a 26-point median gap between grounded mention rate and ungrounded mention rate on identical prompts.

Why: they are driven by different mechanisms. Grounded answers respond to content and retrieval within weeks; ungrounded answers reflect a training corpus that lags by many months. Nothing about 2026 closes that gap - if anything, faster retrieval widens it.

Falsified if: the median gap narrows below 20 points.

What to do: report them separately. A blended score cannot be diagnosed, because you cannot tell whether a move came from a competitor publishing last week or a new model tier shipping.

4. By 31 March 2027, first-paragraph position will remain concentrated in fewer than 3 brands per commercial prompt

Currently measured: 2.3 brands on average occupy the lead-recommendation position across sampled commercial prompts, against 6.1 brands named anywhere in the same answers.

Why: the lead slot is a recommendation, and recommendations do not hedge across six options. This is the strongest concentration effect in our data and it has been stable for a year.

Falsified if: the mean rises above 3.0.

What to do: position weighting is not an academic refinement. Being named somewhere in an answer is worth a fraction of being the lead recommendation - which is exactly why our scoring weights them 0.3 against 1.0.

5. By 31 December 2026, no single engine will exceed 40% of AI search reach

Currently measured externally: Google AI Overviews at ~38% of AI search traffic, ChatGPT at ~34%, with six surfaces reaching more than 2.4B monthly users between them.

Why: the two leaders are close, growing at similar rates, and serve partly different intents. A shift large enough to break 40% inside four months would require an event we see no sign of.

Falsified if: any single surface exceeds 40% share on the December figures.

What to do: single-engine monitoring will keep producing blind spots. Our panel data shows brands routinely 20+ points apart between their best and worst engine, so a one-engine check tells you almost nothing about the others.

6. By 30 June 2027, more than 25% of tracked brands will have at least one materially false factual claim stated about them by a major assistant

Currently measured: 22.4% of brands in our panel have at least one flagged factual contradiction - retired pricing tiers, discontinued plans, invented limitations, or feature claims contradicted by their own documentation.

Why: the rate has risen each quarter we have measured it. Stale facts propagate into training corpora and persist in ungrounded answers long after the source is corrected.

Falsified if: the rate stays at or below 25% at the June measurement.

What to do: run brand-led prompts monthly and check the claims, not just whether you appear. The repair loop is in fixing negative brand sentiment in AI.

7. By 31 December 2026, median time-to-first-measurable-movement after a content fix will remain above 45 days on grounded engines

Currently measured: a 68-day median from shipping a substantive page change to a statistically detectable change in citation share on grounded engines, across the fix cycles we have tracked.

Why: the delay is the compound of crawl, index, retrieval eligibility, and the sampling needed to distinguish real movement from noise. None of those four is getting dramatically faster.

Falsified if: the median drops below 45 days.

What to do: plan AI visibility work in quarters, not sprints, and resist judging a fix at three weeks. Judging too early is how good changes get reverted.

8. By 31 March 2027, the AI visibility tool category will still have no consolidation event involving a top-five vendor

This is me re-betting the call I lost last year, with the same reasoning I used to explain the miss.

Currently measured qualitatively: new entrants continue monthly; the legacy SEO suites are building rather than buying.

Why: the underlying measurement surface is still moving too fast for an acquired pipeline to hold its value through an integration cycle.

Falsified if: a top-five vendor by market presence is acquired before 31 March 2027.

What to do: buy on measurement quality and evidence export rather than on brand or funding. The five questions to ask any vendor are in the methodology FAQ, and the market survey is in the best AI brand monitoring tools.

9. By 31 December 2026, more than 30% of tracked domains will be blocking at least one major AI crawler by accident

Currently measured: 27.8% of the domains we track disallow at least one of GPTBot, OAI-SearchBot, Google-Extended, ClaudeBot, or PerplexityBot - and in the cases we have investigated with customers, the majority could not say when or why the rule was added.

Why: blanket AI-crawler rules were widely copied into robots.txt during 2024 and 2025, often by a developer acting on a single article, and almost nothing prompts a team to revisit them. Meanwhile the distinction that matters - blocking a training crawler while allowing a retrieval crawler - requires a deliberate decision most sites have never made.

Falsified if: the share drops below 30% and the accidental proportion falls with it.

What to do: read your own robots.txt before you commission any content work. Blocking retrieval and then paying to improve pages for retrieval is the most expensive mistake in this category, and the AI crawler checker reads your current configuration in a few seconds.

10. By 30 June 2027, multi-engine tracking will be the majority behaviour at signup

This is the 2025 call I scored partial, re-bet with a tighter threshold.

Currently measured: 2.6 engines enabled on average by new accounts in their first week, up from 2.0 a year ago, with 44% enabling three or more.

Why: the gap between a brand's best and worst engine is wide enough - routinely 20 points or more - that a single-engine read misleads, and teams discover this quickly once they see the second engine.

Falsified if: fewer than 50% of new accounts enable three or more engines by the June measurement.

What to do: if you are checking one engine today, add Perplexity. It is the most explicit citation surface, which makes it the fastest way to learn which pages are actually winning retrieval in your category.

Three things I do not expect

  • AI search will not kill SEO. Grounded engines retrieve from the open web; being retrievable is a prerequisite for being cited. The relationship is covered in AI visibility vs traditional SEO.
  • There will be no stable "AI rank tracker" in the Google-rankings sense. Answers are generated and non-deterministic. Any product selling a single deterministic position is reporting one draw from a distribution and calling it a rank.
  • Brand-owned content will not become the dominant citation source. At 3.8% today, a reversal would need a change in engine behaviour, not a change in publishing effort.

What to do with the rest of 2026

  1. Baseline honestly across all five engines with repeated sampling. One spot check is a coin flip.
  2. Audit your third-party surface first - review platforms and community. Predictions 1 and 2 say that is where the leverage is, and it is not on your own site.
  3. Separate grounded from ungrounded in your reporting, per prediction 3.
  4. Fight for the lead slot, not just a mention, per prediction 4.
  5. Check your facts monthly, per prediction 6.
  6. Give fixes a quarter to land, per prediction 7.

A free 90-second audit covers step 1, and a free trial covers the rest - no card required.

FAQ

How will these be scored?

Against the stated thresholds on the stated dates, using the same panel and the same methodology. Where a threshold is measured on our panel rather than public data, the panel composition at scoring time will be stated alongside the result, because a panel that changed shape underneath a prediction would make the score meaningless.

Are the panel numbers audited?

They are internally reproducible - every figure traces back to stored answers with timestamps, engines, and model tiers - but they are not third-party audited. Treat them as what they are: measurements from one panel of roughly 1,200 brands, with the sampling error described in the methodology.

Which prediction matters most for a small team?

Prediction 1. If review platforms and community sources are approaching half of all citations, a small team's highest-leverage work is largely off its own website - which is good news, because that work does not require a content operation.

What is the biggest risk in ignoring all of this?

Prediction 4's implication: fewer than three brands hold the lead recommendation per prompt. Those positions get harder to take the longer an incumbent holds them, because the citations, branded searches, and reviews that earned the position keep compounding for whoever already has it.

Will you publish the scores?

Yes - at each stated date, in this format, including the misses. Last year's are at the top of this page.

Ready to track your AI visibility?

Monitor your brand across ChatGPT, Perplexity, Claude, Gemini & Grok.
Get Started

No credit card required.