The Capability Accuracy Audit

The fastest way to test your Expertise is to ask AI the exact capability questions your customers ask, then check the answers against reality. This worksheet shows you how to build the prompts, run them across engines, and spot where AI invents limitations or misses features.

You can get a first read on your Expertise, the ANSWER category that measures whether AI describes your product accurately, by testing the specific capabilities customers actually ask about. This worksheet shows you how to build the prompts and read the results. For the full background, see the Expertise pillar guide.

Step 1: List the Facts That Matter

Write down the capabilities a customer checks before buying, the ones that make or break an evaluation. Group them into a few types: integrations (which tools you connect to), features (specific functions), specifications (limits, performance, compliance), and platform or pricing facts. Aim for the 10 to 15 facts most likely to be a customer's hard requirement.

Step 2: Turn Each Fact into a Prompt

Phrase each as a customer would ask an engine. Use templates like these:

  1. "What integrations does [brand] support?"
  2. "Does [brand] offer [specific feature]?"
  3. "Does [brand] support [specific integration or standard]?"
  4. "What are the technical specifications of [brand]?"
  5. "Does [brand] have [compliance or security capability]?"
  6. "What are the limits of [brand] on [plan or dimension]?"

Step 3: Run the 3 Rules

  1. Log out. Run every prompt in a logged-out or private session, so the engine is not drawing on your history.
  2. One fresh chat per prompt. Start a new conversation each time, so one answer does not color the next.
  3. Use your real market. Run the prompts in the geography you actually sell into, because answers vary by location.

Run each on at least 2 engines, for example ChatGPT and Perplexity, plus Google's AI answers if you can.

Step 4: Score Each Answer

For each prompt, mark one of four outcomes: accurate, incomplete (a real capability omitted), outdated (an old version described), or invented (a wrong fact or limitation stated). Write down the exact error and, where you can tell, the likely source the engine drew on.

How to Read the Results

If the answers are accurate and complete across most prompts, your Expertise is strong. If real capabilities are omitted, you have a completeness gap that makes you look weaker than you are. If you see invented limitations, you have AI hallucination to correct, and those errors are the most urgent because they sound authoritative. If specs are simply old, you have a freshness problem traceable to outdated sources. The patterns and fixes are explained in Invented Limitations and Missing Features.

What This Check Cannot Tell You

A capability spot check is a useful hint, not a baseline. It covers the facts you thought to test, on one occasion, scored by eye, and answers vary from ask to ask. A real measurement runs many capability questions, repeats them, covers more engines, and scores every answer the same way, reporting an Accuracy Rate you can track over time. Doing that by hand after every release is a real undertaking, which is the honest reason most teams have it run for them.

If your audit turns up errors, the next step is a proper baseline and a fix plan. A HiBot audit measures your Expertise and Accuracy Rate across engines and ranks what to fix first. The tactics are in How to Improve Your Expertise Score. Download the ANSWER whitepaper to see the method, or request an AI visibility audit at hibot.com.

David Tang
David Tang · Corporate Strategy, New York
David Tang is the CEO and Founder of HiBot and Flevy. Flevy is the world's largest marketplace for business frameworks and templates. Prior to these companies, David worked as a management consultant for 8 years, where he served clients in North America, EMEA, and APAC. He graduated from Cornell with a BS in Electrical Engineering and MEng in Management. LinkedIn →
Connect our AI visibility playbook to your AI