Skip to content

How often are AI answers wrong? What the major accuracy studies measured

The best-known AI accuracy studies define "wrong" differently and date fast. What each measured, and what carries over to facts about a company.

By Published Updated 7 min read
Research and teardowns, 7 min read — Three glass lenses focusing one violet beam onto a dark surface, each producing a differently placed and differently sharp point of light.

The short answer

Often enough to check, but the headline rates are not comparable. The major studies tested news answers and citations, each with its own definition of an error, its own sample and its own dates. What carries over to a company is the method: ask real questions, define an error in advance and check every answer against a source.

Key takeaways

  • The best-known accuracy studies tested news questions and source attribution, not facts about companies, so their rates do not transfer directly.
  • Each study defined "wrong" differently: a significant issue, an incorrect citation or a misleading health summary are not the same error.
  • Accuracy findings date quickly because engines and models change, so a study's data period matters as much as its headline.
  • Google removed AI Overviews from some liver-test queries in January 2026 after a newspaper investigation into misleading health summaries.
  • Our view: measure your own error rate on your own buyer questions, with errors defined before you read the answers.

The short answer: often enough to check, never one number

AI answers are wrong often enough that any company should check what they say about it. But there is no single error rate. The studies people quote measured different tasks, in different months, with different definitions of a mistake. Quoting one rate as "how often AI is wrong" mixes all of that up.

This post is a desk synthesis of the studies most often cited in 2025 and 2026. For each, it sets out what was asked, how an error was defined, how big the sample was and when the answers were collected. Then it asks what carries over to the facts AI states about a business.

Our view: a study's definition of "wrong" and its data period matter more than its headline rate. Read those first, and the numbers make sense.

What counts as "wrong"?

Studies use at least four kinds of error, and they are not interchangeable. Before comparing any two rates, check which kind each one counts.

Factual error
The answer states something false: a wrong date, figure, name or rule.
Sourcing error
The answer's claim is not supported by the source it cites, or cites no source, or cites the wrong one.
Misleading answer
Each statement may be true, but the answer leaves out context a reader needs, such as a range that depends on age or sex.
Outdated answer
The answer was once true but no longer is, such as an old price, a former chief executive or a closed office.

Some studies count an answer as wrong if it has any one of these. Others count only serious cases, judged by experts. A rate built on "any issue" will always look higher than one built on "significant issue", even for the same answers.

Sample design matters as much. A study that asks hard questions chosen to probe weaknesses will find more errors than one that asks typical questions. Neither is wrong, but they answer different questions: how bad can it get, and how bad is it usually.

Health summaries: errors with consequences

Health is where AI errors have drawn the strongest response. TechCrunch reported on 11 January 2026 that Google had removed AI Overviews from some liver-test queries after a Guardian investigation found misleading health summaries, including normal ranges given without the context needed to read them.1 Google declined to comment on individual removals and said its clinicians had found much of the information was not inaccurate.1

The case shows the third kind of error at work. A range can be accurate for one group and misleading for another, and a short summary drops the caveats. It also shows that platforms act when errors are documented with specific examples. Our guide for clinics on what patients ask AI looks at what a healthcare provider should check.

The EBU and BBC news study

The largest public study of AI news answers came from the European Broadcasting Union and the BBC, published in October 2025. Public broadcasters asked AI assistants news questions and had journalists review the answers against professional criteria, including accuracy, sourcing and context.

In answers collected in May and June 2025, data now more than a year old, 45% of AI news answers had at least one significant issue and 31% had sourcing problems.2 The study counted significant issues judged by journalists, not every small slip.

What it tells a company: sourcing was the most common class of problem. An answer can name the right facts and still attribute them to the wrong page, or to no page. For a business, a misattributed claim is harder to correct, because you cannot see where it came from.

The Tow Center citation test

The Columbia Journalism Review's Tow Center tested something narrower: whether AI search tools could identify the source of a news excerpt. It gave eight tools 1,600 queries built from real article excerpts (test period not stated) and checked whether each named the right publisher, article and link.3

Published in March 2025, more than a year before this post, with the test period not stated, the study found the tools gave incorrect answers to more than 60% of queries.3 It also reported that tools often answered with confidence rather than declining.

The task was deliberately hard and specific, so the rate is not a general error rate. Its lasting lesson is about confidence: tools rarely said they did not know. A company should expect a wrong answer about itself to sound as sure as a right one.

StudyTaskError countedData
EBU and BBCNews questionsSignificant issue, judged by journalistsMay–Jun 2025 2
Tow CenterIdentify a news excerpt's sourceWrong publisher, article or linkNot stated; published Mar 2025 3
Health summariesMedical search queriesMisleading summary, found by reportersReported Jan 2026 1

Our summary of each study's design. The rates are not comparable across rows because the tasks and error definitions differ.

What carries over to facts about a company?

The method carries over; the rates do not. None of these studies asked about a company's prices, products, locations or leadership. Those facts are smaller and less covered than the news, which can cut both ways: fewer sources to confuse the engine, but also fewer sources to correct it.

Three patterns from the studies are likely to apply. Sourcing errors are common, so check the cited page as well as the claim. Confidence does not track accuracy, so a fluent answer needs the same check as a hesitant one. And errors about anything that changes, such as prices or staff, often come from old pages that are still online.

Buyers already behave as if this were true. In a Gartner survey of 645 B2B buyers with fieldwork in August and September 2025 (geography not stated), 69% said they turn to sales reps to validate AI-generated insights.4 A wrong AI answer about your company does not always lose the sale, but it can put your sales team in the position of correcting it. Our guide to fixing what ChatGPT gets wrong about a company covers that correction work.

How to measure your own error rate

Define errors first, then ask, then check. That order keeps the result honest.

An error audit for facts about your company

  1. List the facts

    Prices, products, locations served, leadership, founding date, policies.

  2. Define an error

    False, unsupported by its source, misleading or outdated, decided in advance.

  3. Ask real questions

    The questions buyers ask, on several engines, on separate days.

  4. Check every answer

    Against your own records and the source the answer cites.

  5. Trace and fix

    Find the page the error came from and correct it there.

Illustrative. The order we follow; adjust the fact list to your business.

Our audits use the same structure across six engines: ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI. Each prompt runs on three separate days, and errors go into a register with the source each answer cited.

Our view: report errors as a count of wrong facts, with examples, before you report any rate. A list of six specific errors, each traced to a page you can fix, is more useful than a percentage.

Sources

  1. TechCrunch, Google removes AI Overviews for certain medical queries (Jan 2026)
  2. EBU and BBC, News Integrity in AI Assistants: answers collected May–Jun 2025 (Oct 2025)(dated)
  3. Columbia Journalism Review, Tow Center: eight AI search tools, 1,600 citation queries (Mar 2025); test period not stated(dated)
  4. Gartner, survey of 645 B2B buyers, Aug–Sep 2025; geography not stated (May 2026)

Questions readers ask

  • For news, the EBU and BBC study published in October 2025 is the largest, with journalists reviewing answers against set criteria. But its answers were collected in May and June 2025, and engines change fast. No study tests facts about companies at scale, so the most reliable figure for your business is the one you measure yourself.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.