Skip to content

Which facts about your company does AI get wrong? An accuracy audit for the details buyers check

Pricing, availability, integrations, compliance, leadership: a 40-prompt audit that finds wrong AI facts, traces their source and re-checks quarterly.

By Published Updated 12 min read
Citations and accuracy, 12 min read — A grid of glowing glass tiles on a dark table, with a few violet tiles linked by thin threads of light to faint shapes in the distance.

The short answer

The facts AI gets wrong most often are the ones that change: pricing, availability, integrations, certifications and who runs the company. Audit them directly. Write a dated fact sheet, ask around 40 buyer-style questions on each engine on several days, grade every claim against the sheet, trace each error to its source, fix it there and re-check every quarter.

Key takeaways

  • In TrustRadius's survey of 1,862 technology buyers in January 2026, 63% used AI in their purchase journey, and 94% of those fact-checked what it told them.
  • Gartner found in a survey of 645 B2B buyers in August and September 2025 that 69% preferred to validate AI-generated insights with sales reps.
  • In Profound's customer data, pricing and billing had the highest inaccuracy rate of any topic, and wrong claims came from brands' own pages as well as third-party and competitor sites.
  • An accuracy audit grades claims, not mentions: correct, outdated, wrong, unverifiable or missing, per fact type and per engine.
  • Every wrong fact needs a traced source before anyone fixes it; editing your homepage will not correct a claim that comes from a review site.
  • Re-run the audit every quarter and within a few weeks of any price, product or leadership change.

Why the details buyers check matter most

A wrong fact in an AI answer costs most when it is a fact a buyer acts on. Your founding year matters little. Your price, whether you serve their country, whether you integrate with their CRM and whether you hold the certification their security team requires decide whether you make the shortlist.

Buyers do check. In TrustRadius's survey of 1,862 technology buyers in January 2026, 63% used AI somewhere in their purchase journey, and 94% of those said they fact-checked what it told them.1 The same survey found 83% of buyers shortlisted three or fewer products.1 A short list leaves little room for a vendor whose AI description is wrong.

Checking often happens in a sales call. In Gartner's survey of 645 B2B buyers in August and September 2025, 69% said they preferred to validate AI-generated insights with sales reps.2 That is good news if the rep can correct a wrong price. It is bad news if the buyer never books the call because the answer said you were too expensive, or did not operate in their market.

Our view: accuracy on a handful of decision facts is worth more than a higher mention count. A brand named in every answer with the wrong price is worse off than one named in fewer answers with the right one.

B2B buyers and AI

Buyers use AI, then check it

63%of technology buyers used AI in their purchase journeyTrustRadius, Jan 2026 1
94%of those buyers fact-checked what AI told themTrustRadius, Jan 2026 1
69%of B2B buyers prefer to validate AI insights with sales repsGartner, Aug–Sep 2025 2
Sources: TrustRadius, Beyond the hype; Gartner B2B buyer survey.

Which facts go wrong most often?

The facts that change go wrong most often, and pricing leads. Profound's August 2026 analysis of claims checked by its own fact-checking product found pricing and billing to be the topic with the highest inaccuracy rate.3 It traced wrong claims to brands' own pages as well as to third-party and competitor sites.3 Profound sells AI visibility software, and its sample is its customers, so treat the ranking as indicative rather than universal.

The reason is simple. Engines learn and retrieve from pages written at different times. When your price changes, your old pricing page, a review site's listing, a comparison article and a reseller's catalogue may all still carry the old one. The engine sees several prices and picks one.

From the buyer's side, six fact types carry most of the risk. We use them as the spine of the audit.

Fact typeWhat usually goes wrongWhat the buyer does
Pricing and billingOld prices, discontinued plans, wrong unitRules you out on budget, or feels misled later
AvailabilityCountries, languages or segments you do not serveDrops you before a call
IntegrationsMissing or invented integrations, old versionsAssumes you will not fit their stack
ComplianceCertifications you lack, or ones you hold left outSecurity review fails or never starts
Company factsOld leadership, wrong owner, merged entitiesDoubts stability or contacts the wrong person
Product lineRenamed or retired products, wrong featuresCompares you on a product you no longer sell

Our classification, drawn from the buyer questions we see in prompt sets; not a measured ranking.

Software buyers also bring more people into the decision. G2's 2026 Buyer Behavior Report, with sample period and region not stated, found finance involvement in software purchases rose from 31% to 46%.4 Finance teams read price and contract terms first, which raises the cost of a wrong number.

Rank the risks before you audit

Rank each fact type by how likely the engines are to get it wrong and by what an error costs you. The first comes from how often the fact changes and how many third-party pages repeat it. The second comes from your sales cycle.

Error likelihood vs cost to the deal

Cost of an error →

Compliance, availabilityChange rarely, but one error ends the deal. Check every quarter.
Pricing, integrationsChange often and decide shortlists. Audit first, re-check after every change.
Founding facts, historyRarely wrong, rarely decisive. Check once a year.
Leadership, product namesChange often, matter less. Re-check after announcements.

Likelihood of error →

Illustrative. Our judgement for a typical B2B company; move the cells to fit yours.

Our view: for most B2B companies, pricing and integrations go in the top-right cell. They change several times a year, third-party sites repeat them, and buyers use them to cut a long list down to three.

How often do AI answers get facts wrong?

Often enough to plan for. The largest public study of assistant accuracy is now dated: the EBU and BBC review of news answers, collected in May and June 2025, found that 45% of AI answers had at least one significant issue, and 31% had significant sourcing problems.5 That study covered news questions, not company facts, and engines have changed since.

It still makes two points that hold for brand facts. Errors cluster around sourcing, which is why tracing matters more than complaining. And the rate varies by engine, so one engine's answer tells you little about another's. Our review of accuracy studies from 2025 and 2026 compares the methods behind the main figures.

Our view: do not wait for a study of your sector. Your own accuracy rate on your own facts is cheaper to measure than to estimate, and it is the only number you can act on.

The 40-prompt accuracy audit

An accuracy audit asks the questions buyers ask and grades every factual claim in the answers against a dated fact sheet. Forty prompts across the six fact types are enough to find the errors that matter for most companies. A visibility audit counts whether you are named; this one checks whether what is said is true.

Run the audit in five steps

  1. Write the fact sheet

    The true answer to every decision fact, with an owner and a date.

  2. Draft 40 buyer prompts

    Spread across the six fact types, phrased the way buyers ask.

  3. Run them on every engine

    Six engines, several days, full answers and citations captured.

  4. Grade each claim

    Correct, outdated, wrong, unverifiable or missing.

  5. Trace and fix

    Find the source of each error and correct it where the engine reads.

Illustrative. The order we use; the prompt count scales with your product range.

Write the fact sheet first

The fact sheet is the ground truth. List every decision fact: current plans and prices with currency and billing unit, countries and languages served, integrations with versions, certifications with issue dates, leadership and ownership, and current product names. Give each fact an owner who will tell you when it changes.

Without it, graders argue about what is correct. With it, grading is quick and two people reach the same verdict.

Draft prompts the way buyers ask

Mix three kinds of prompt for each fact type. Direct questions name you: "How much does Northwind Digital cost per user?" Comparative questions name you and a rival: "Northwind Digital or Acme Search Co. for a 50-person team, and what does each cost?" Category questions do not name you: "Which analytics tools for a small sales team integrate with our CRM?" The last kind shows whether a wrong fact is quietly keeping you off lists.

A sensible split for 40 prompts is eight each for pricing and integrations, and six each for availability, compliance, company facts and product line. Weight it towards the cells you ranked highest.

Run and capture

Run each prompt on every engine your buyers use, on more than one day. Answers vary between runs, so one capture can make an engine look better or worse than it is. Record the date, engine, country, full answer text and every citation, and keep the raw capture so a disputed grade can be checked.

A row of translucent glass cards on a dark surface, most lit evenly in cool blue, two with a faint violet fracture running through them, under a narrow beam of light.
Most claims hold up. The audit exists to find the few that are cracked, and where the crack started.

Grade claims, not answers

Break each answer into factual claims about you and grade each one. Correct matches the fact sheet. Outdated was true once. Wrong was never true. Unverifiable cannot be checked against the sheet. Missing means the answer left out a fact the question asked for, such as a price, where it gave one for competitors.

Grading claims rather than whole answers matters. An answer that gets your integrations right and your price wrong is one correct claim and one error, and you need to know which fact type failed.

Trace every wrong fact to its source

Find where each error came from before anyone touches a page. When the answer cites sources, open them: the wrong fact is often there, on an old page of yours or a third-party listing. When it does not, search for the exact wrong phrase or number; old press releases, review profiles, partner directories and comparison articles are the usual suspects.

From wrong answer to fixed source

  1. Wrong claim
  2. Cited sources
  3. Origin found
  4. Fix at source
  5. Re-run prompt
Illustrative. Tracing comes before fixing; a fix in the wrong place changes nothing.

Then fix it where the engine reads. That may be your own outdated page, which you can redirect or update today. It may be a review site profile you can edit, a directory listing you can claim or an article whose author will usually correct a dated price if asked. For a knowledge-panel or encyclopaedia entry, our guide to Wikipedia entries and AI answers explains what you can and cannot change.

Pricing errors deserve their own pass. Our note on wrong AI pricing and third-party sources covers why prices are the most persistent error, and our step-by-step fix for wrong ChatGPT answers covers reporting errors to the engines themselves.

Corrections take time to appear. MaxAEO's June 2026 piece on fixing wrong ChatGPT descriptions examines how long corrections take to show up in answers once the source is fixed.6 Plan for weeks, not days, and re-run the affected prompts rather than assuming.

Score it so you can track it

Report an accuracy rate per fact type and per engine: correct claims divided by all claims made about you. Keep missing facts as a separate count, because silence and error need different fixes. Report the number of claims behind each rate, so nobody reads a rate built on three claims as a trend.

Fact typeClaims gradedCorrectMost common error
Pricing140About 6 in 10Old plan names and the previous price
Integrations120About 8 in 10A retired integration still listed
Availability90About 9 in 10One country wrongly excluded
Compliance70About 7 in 10A certification left out

Illustrative. Northwind Digital is fictional; counts and shares are invented round numbers, not benchmarks.

Weight errors by severity if you need one number for a board. A wrong price on a direct pricing question is worse than a wrong founding year in passing. Keep the weights simple and publish them with the score.

Re-check every quarter, and after every change

An accuracy audit is a baseline, not a certificate. Prices change, products are renamed, engines update their models and sources, and an error you fixed can return when an old page is re-crawled. Re-run the full set every quarter and the affected prompts after any change.

A year of accuracy checks

  1. Fact sheet, 40 prompts, baseline runs on every engine.

  2. Trace and fix errors, highest-risk fact types first.

  3. Re-run the full set; compare rates per fact type and engine.

  4. Re-run the affected prompts a few weeks after a price, product or leadership change.

  5. Refresh the prompt set to match how buyers now ask.

Illustrative. Our recommended cadence; tighten it if prices change monthly.

Watch the engines as well as your own changes. When an engine changes how it retrieves sources, or starts showing answers in a new country, accuracy can move without anything changing on your side. A quarterly run catches that; a single annual audit does not.

Keep the history. A dated log of every graded claim, the source you traced and the fix you made turns the next audit into a comparison rather than a fresh start. It also tells you which fixes worked, which is the only way to learn which sources the engines actually lean on for your category.

Tie the re-check to your own change calendar. Whoever owns pricing should tell whoever owns the audit before a price changes, so the old prices can be retired from every page you control on the same day.

What to leave out

Leave out tactics that make errors worse. Do not publish a page arguing with the engines; publish the correct fact clearly. Do not create thin pages for each wrong claim. Do not ask staff or partners to post corrections that read as reviews. These look like manipulation to people and to engines, and they rarely move the answer.

Also leave out facts you do not want repeated. If your pricing is quote-only, say so plainly and give a starting point if you can. An answer that says "pricing on request" is more useful to a buyer than an invented figure, and a stated starting price beats both.

How we run accuracy checks

Accuracy is part of every Sigzen AI audit. We run each buyer prompt on six engines, ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI, on three separate days. Our AI answer monitor uses official APIs where an engine offers one and a licensed AI-answer data provider where it does not. Every wrong or outdated statement goes into a hallucination register with the answer it came from and the source we traced it to.

The register sits beside the share-of-answer baseline, so you see both whether you are named and whether what is said is true. The audit is published on our pricing page, and our audit page lists what it delivers.

Buyers will keep checking AI answers against your site and your sales team. Make sure that when they check, the answer and the page say the same thing.

Sources

  1. TrustRadius, Beyond the hype: 1,862 technology buyers, Jan 2026 (Jul 2026); 94% is of buyers who used AI
  2. Gartner, survey of 645 B2B buyers, Aug–Sep 2025; geography not stated (May 2026)
  3. Profound, where inaccurate AI claims come from (2026)
  4. G2, Buyer Behavior Report (Jul 2026); data period and region not stated
  5. EBU and BBC, News Integrity in AI Assistants: answers collected May–Jun 2025 (Oct 2025)(dated)
  6. MaxAEO, when ChatGPT gets your company wrong (Jun 2026)

Questions readers ask

  • Around 40 is enough for most companies with one product line: eight each for pricing and integrations, six each for availability, compliance, company facts and product line. Add prompts if you sell several products or plans. Run each prompt on every engine on more than one day, because a single capture can mislead.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.