AI visibility reports often start with one overall score showing how often your brand appears in ChatGPT.
Which is ok as an overall fast metric, but it can hide a major difference in how ChatGPT researches an answer.
Semrush and Kevin Indig ran 100 prompts through GPT-5.2 using minimal and high reasoning. Only 25.6% of the cited domains appeared in both sets of answers. The prompts stayed the same, yet almost three-quarters of the sources changed.
That creates a problem for marketers using one combined score to decide whether their content is working. A brand can appear regularly in faster answers and then disappear when ChatGPT runs more searches for a comparison, validation, or purchase decision.
Here’s what that study found and what you should be doing:
What the Study Tested
The test covered 20 buyer journeys across B2B SaaS, finance, consumer tech, and health and lifestyle.
Each journey moved through five stages: problem, exploration, comparison, validation, and selection. Every prompt was run once with minimal reasoning and once with high reasoning, creating 200 responses in total.
Minimal reasoning matched ChatGPT’s faster Instant experience. High reasoning matched Thinking mode, where the model spends more time breaking down the question and researching the answer.
A quick caveat. This was a Semrush-produced study using 100 prompts and one version of GPT-5.2. The percentages describe this test, so they shouldn’t be used as a fixed benchmark for every company or industry.
The Amount of Research Changed the Result

High reasoning did a lot more work before producing the answer:
Citation rate increased from 50% to 68%.
Sources per cited response increased from 2.6 to 4.5.
Total web searches increased from 245 to 1,130.
High reasoning used 173 unique domains, compared with 127 under minimal reasoning.
The biggest increase came during comparison prompts. High reasoning ran an average of 24 smaller searches per response, compared with 5.5 under minimal reasoning.
Think about a buyer asking which CRM is best for a 50-person SaaS company. ChatGPT may check pricing, integrations, security, implementation, support, API limits, and product documentation before giving a recommendation.
A page written around the main comparison query only covers one part of that research.
The Source Types Changed as Well
Reddit’s share of citations fell from 15% under minimal reasoning to 7% under high reasoning. Other review and user-generated sites fell from 14.3% to 6%.
Government and academic sources increased from 1.9% to 8.8%. Official documentation and support pages increased from 12.4% to 17.5%.

For marketers, this is a clear reason to use varied content formats. Short answers, product pages, reviews, and community threads can support faster questions. Detailed comparisons rely more on documentation, research, technical details, and evidence.
What Should Marketers Do Differently?
Split Prompt Tracking by Research Depth
Start by dividing the prompts you already track into two groups.
Basic questions:
What is the product?
Does it include a certain feature?
How much does it cost?
What does a term mean?
Detailed questions:
Which product is best for a specific company?
How do several products compare?
Does the product meet a security or compliance requirement?
Is it worth the price?
What are the risks of choosing it?
Report the results separately. Track the brand mention rate, citation rate, and main source types for each group.
This will show whether your brand is strong for basic product questions while competitors are being used for the detailed questions closer to a decision.
Build Complete Buyer Journeys
A list of disconnected prompts gives you a limited view of what happens as someone moves towards a purchase.
For each important product or service, create one prompt for each stage:
Problem: How does the buyer know they have the problem?
Exploration: What types of solutions could help?
Comparison: Which options meet their requirements?
Validation: What evidence would confirm the choice?
Selection: How do they buy or implement it?
Then track whether your brand appears at each stage of the journey. When it doesn’t, record which brands and sources ChatGPT uses instead.
In the study, a brand was present from the problem stage through to selection in 4 of the 20 buyer journeys under high reasoning. This didn’t happen at all under minimal reasoning.
This shows why a single prompt result is limited.
Audit the Smaller Questions ChatGPT Needs to Answer
Take your ten most important comparison and selection prompts. For each one, list the checks ChatGPT might need to make before it can recommend a product.
For example, a software product could include:
Pricing and contract terms.
Integrations.
Security and compliance.
Product limits.
Setup and migration.
Support.
Best-fit customers.
Evidence behind performance claims.
Now check whether each answer exists on a page that can be found and understood easily.
Give the question a clear heading. Answer it directly. Add the relevant detail, source, date, or proof underneath. Keep related pages connected with internal links so the product information is easy to follow.
This is where content gaps usually exist. The main product page is there, but the evidence needed for a detailed citation is spread across sales copy, PDFs, old support posts, and pages with poor content.
Change the Report
Keep the overall score as a top-level summary. Add the detail the content team needs underneath it.
Track:
Basic prompt mention rate.
Detailed prompt mention rate.
Website citations.
Brand mentions without a website citation.
Website citations without a brand mention.
Performance by buyer-journey stage.
Source types used for each prompt.
Appearance rate across three repeated runs.
The commentary should explain what changed and what needs to happen next.
For example:
Brand mentions increased for product and pricing questions. Comparison prompts still rely on competitor documentation and third-party reviews. We need stronger security, migration, and product-limit pages.
That gives the team a clear direction to work from. A report saying visibility increased by 8% gives them very little to work with - but might be satisfactory to an exec or client.
Don’t Rely on Scoring Systems
The main point is to stop taking platform AI visibility scores as the primary performance metric.
They can show needle movement, but they don’t explain why your brand appears, where it drops out, or how results change across buyer questions.
Instead, you need to know which questions your brand appears for, how consistently it appears, and what sources ChatGPT relies on as the buyer moves closer to a decision.
That gives teams something they can act on. Rather than chasing a higher score, they can focus on the stages, questions, and supporting evidence where the brand is still being missed.
