calendar Last updated: 16 September 2026
Data integration

How Accurate Is “AI Visibility” Tracking for ChatGPT?

Try It Now

Connect your data in 1 min. Start free, no credit card

Got insights from this post? Give it a boost by sharing with others!

AI-assistant referral traffic to windsor.ai has grown by roughly 780%
over the last six weeks compared to the six weeks before that, and it’s
still climbing: about 1 in 7 visitors to windsor.ai now arrive via a link
clicked inside ChatGPT or Claude.

When someone asks ChatGPT “what’s the best tool to connect Google Search
Console to ChatGPT,” how often does it name Windsor? And how much can any
tool that claims to measure that be trusted?

The problem with API-based “AI visibility” tracking

Most tools that report on brand visibility in AI answers send a prompt
to the model’s API and read the response. It’s the only practical way to
run thousands of prompts a day. The API and the consumer app are
different products, though, with different system prompts and defaults,
and in ChatGPT’s case the app has no web-search toggle while the API
does. So an API-only tracker describes the API, and there’s no guarantee
that matches what a person gets in the app.

We didn’t want to take that on faith, so over the past several weeks we
ran three rounds of testing on ChatGPT to find out how far apart the two
are.

Three rounds, growing more precise each time

Models and settings by round

Round Side Model Web search Effort Prompts Answers Window
1 ChatGPT API gpt-5-mini On Medium 30 Over 6,200, both platforms combined Aug 17 to early September
1 Claude API claude-sonnet-5 On Medium 30 Aug 18 to early September
1 ChatGPT app (manual spot-check) App default, not recorded App-controlled, no toggle n/a Same prompts, spot-checked Not counted Same period
2 ChatGPT API gpt-5-mini On Medium 25 1,306 Aug 27 to Sep 7, 8 run days
2 ChatGPT app App default, not recorded App-controlled, no toggle n/a 25 368 Aug 27 to Sep 7, 8 run days
3 ChatGPT API gpt-5-mini On Default (not set) 10 585 Sep 7 to Sep 15, 7 run days
3 ChatGPT app App default, not recorded App-controlled, no toggle n/a 10 90 Sep 7 to Sep 15, 6 run days

“Effort” is the API’s reasoning-effort setting. On the app side, a
temporary chat uses whatever model the app assigns to the account and
decides on its own whether to search the web, so neither was set by us,
and the model the app used was not captured.

Round 1: API-only tracking across Claude and ChatGPT

What we measured: whether windsor.ai and its usual competitors showed up
in API answers, on both Claude and ChatGPT, across 30 tracked prompts per
platform and over 6,200 responses, with a manual spot-check of the same
prompts in the real, logged-in ChatGPT app.

  • Claude’s answers named windsor.ai consistently, in line with other
    signals we track.
  • ChatGPT’s API answers didn’t: neither our brand nor the classic
    competitors in the space ranked in the top 25 entities named.
  • Manually testing the same prompts in the real, logged-in ChatGPT app
    produced answers that looked nothing like the API output.

One example: asking “how do I connect Google Search Console to ChatGPT”
surfaced tools in the app (HYDP, GSC Wizard) that never appeared once in
the API data.

Round 2: Open prompts, API against the real app

What we measured: the same open prompts (“how do I connect [platform] to
ChatGPT”) run in parallel through the API (gpt-5-mini, web search on) and
through a real, logged-in ChatGPT account in fresh temporary chats.
Prompts of this shape usually made the app ask a clarifying question and
wait, while the API answered in full regardless, so an API-only
measurement can’t reflect what a person sees.

  • 25 prompts, August 27 to September 7: 8 run days on both sides.
  • Cadence: 1 to 3 app runs per prompt per day (usually 2) in batches at
    different hours; a median of 14 API runs per prompt per day, spread
    across 08:00 to 22:00 UTC.
  • 1,306 API answers against 368 app answers.
  • Median answer length: 5,585 characters on the API against 982 in the
    app.
  • Tools named per answer: 5.3 on the API against 1.1 in the app.
  • Answers naming no tool at all: 7% of API answers, 57% of app
    answers.
  • Zapier named in 57% of API answers and 6% of app answers.
  • Windsor.ai named in 6% of API answers and 5% of app answers.
  • Same-day pairs: 103 (prompt, day) pairs had runs on both sides; in 51
    of them, half, the two tool lists had nothing in common.

Round 3: Direct prompts, with position tracked

What we measured: a single, direct prompt shape (“best tools to connect
[platform] to ChatGPT”) that gets a straight list from both the API and
the app, so the two can be compared answer for answer. Every named tool’s
position in the answer was recorded alongside whether it was named, so
prominence and consistency could be measured as well as presence.

  • 10 prompts, September 7 to September 15: 7 API run days, 6 app run
    days.
  • Cadence: 1 or 2 app runs per prompt per day in batches at different
    hours; a median of 10 API runs per prompt per day, spread across 08:00
    to 22:00 UTC.
  • 585 API answers against 90 app answers.
  • Median answer length: 4,883 characters on the API against 3,009 in
    the app.
  • Tools named per answer: 8.3 on the API against 4.4 in the app.
  • Answers naming no tool at all: 10% of API answers, 2% of app
    answers.
  • Zapier named in 86% of API answers and 70% of app answers.
  • Windsor.ai named in 5% of API answers and 2% of app answers.
  • Same-day pairs: 53 (prompt, day) pairs had runs on both sides; in 3
    of them, 6%, the two tool lists had nothing in common.

Round 2 and round 3 side by side

Metric Round 2 API Round 2 app Round 3 API Round 3 app
Answers 1,306 368 585 90
Median answer length (characters) 5,585 982 4,883 3,009
Tools named per answer (mean) 5.3 1.1 8.3 4.4
Answers naming no tool 7% 57% 10% 2%
Zapier named in 57% 6% 86% 70%
Windsor.ai named in 6% 5% 5% 2%
Same-day pairs with nothing in common 51 of 103 (49%) 3 of 53 (6%)

Round 2 and round 3 used different prompt wording, so the two halves of
the table describe two separate experiments measured the same way.

What round 3 found

The API names more tools per answer, leans hard toward
automation-platform names the app barely uses, and names even the shared
tools more often than the app does.

  • Pipedream: 59% of API answers, 0% of app answers.
  • Chatfuel: 16% of API answers, 0% of app answers.
  • Zapier: 86% of API answers, 70% of app answers.
  • Make: 84% of API answers, 59% of app answers.

Bar chart of the share of ChatGPT answers naming each tool, API against the real app: Zapier 86% vs 70%, Make 84% vs 59%, n8n 67% vs 26%, Pipedream 59% vs 0%, ManyChat 29% vs 4%, Supermetrics 18% vs 17%, Chatfuel 16% vs 0%, Apify 13% vs 11%

Tool Named by API Named by real app
Zapier 86% 70%
Make 83.6% 58.9%
n8n 66.8% 25.6%
Pipedream 59% 0%
ManyChat 29.1% 4.4%
Supermetrics 17.9% 16.7%
Chatfuel 16.2% 0%
Apify 13% 11.1%

Windsor.ai itself was named in 5% of API answers and 2% of app answers
for these same 10 prompts, low on both sides, and lower still in what a
real user would see.

Neither side is consistent with itself

Mention rates assume a tracker’s own reading of “does ChatGPT recommend
this tool” is stable, but it isn’t on either side. For each of the 10
prompts, we compared every repeat answer against every other repeat answer
to that same prompt and measured how much of the named-tool list
overlapped.

  • API answers to the same prompt agreed with each other 29% of the time
    on average, ranging from 17% to 44% across prompts.
  • Real app answers agreed with each other 36% of the time on average,
    ranging from 10% to 56%.

The app’s higher average comes with a wider spread, so it doesn’t make
the app the more stable side.

Bar chart of the average overlap between repeat answers to the same prompt, API against the real ChatGPT app, for each of the 10 prompts

Mention rate hides where a tool appears

A tracker that only checks whether a tool was named misses where in the
answer it landed. Using Zapier as the tool named often enough on both
sides to check this cleanly:

  • API (585 answers): named in 86%, opens the answer 27% of the time
    it’s named, average position 3.5.
  • Real app (90 answers): named in 70%, opens the answer 51% of the time
    it’s named, average position 2.1.

The API names Zapier more often, but the app, when it does name Zapier,
puts it first more than twice as often.

Bar chart of the share of Zapier mentions that open the answer, API against the real ChatGPT app, for each of the 10 prompts

Which side leads also flips from prompt to prompt:

Prompt API opens with Zapier App opens with Zapier
“best tools to connect Google Ads to ChatGPT” 75% 17%
“best tools to connect LinkedIn to ChatGPT” 62% 89%

A tracker counting mentions only would record both prompts as “Zapier
present on both sides” and miss that the two answers put it in a
different place.

Position also moves for a single prompt on a single side. In the API’s
answers to the Instagram Public Data prompt, Zapier’s position ranges
from 2nd to 15th across the 52 times it’s named; on the GA4 prompt, from
1st to 10th. We split that movement into two pieces, averaged across all
10 prompts:

  • Two API calls made hours apart on the same day differ by 1.7
    positions on average.
  • Two calls made on different days differ by 0.5.

Most of the movement happens call to call within a single day, rather
than as a slow drift across the nine-day window.

Why the gap changed between round 2 and round 3

One result from round 3 appeared to contradict round 2. In round 2 the
app’s answers were terse and mostly empty; in round 3 they were nearly as
long and tool-dense as the API’s. The setup didn’t change between the two
rounds: same account, same script, the same kind of fresh temporary chat,
back to back.

The prompt wording is the one variable that differed:

Round Prompt wording App answers naming no tool Median app answer length
2 “how do I connect [platform] to ChatGPT” (open) 57% 982 characters
3 “best tools to connect [platform] to ChatGPT” (direct) 2% 3,009 characters

The open prompt gives the app room to ask a clarifying question instead
of answering; the direct prompt doesn’t.

Where this leaves us

Bar chart of chatgpt.com and claude.ai referred visitors to windsor.ai versus the earliest period: baseline, then +780% over the last six weeks, then +1,650% over the last two weeks

AI-assistant referrals to windsor.ai are growing fast enough that whether
ChatGPT recommends us in answers like these has a direct effect on our
own traffic. We measure it to see whether the changes we make to our
content get us in front of more real users, and for that the number has
to describe what those users see.

We only ran the app comparison on ChatGPT. Claude’s API named windsor.ai
consistently in round 1 and Claude sends us referral traffic, but we
haven’t checked Claude’s app against its API, so nothing here says how
accurate Claude tracking is.

API-based tracking is still the only way to run thousands of prompts a
day, and we’ll keep using it for that. For ChatGPT, though, an API number
and an app number are two different measurements, and a reported
“visibility” rate should say which one it is.

Tired of juggling fragmented data? Get started with Windsor.ai today to create a single source of truth

Let us help you automate data integration and AI-driven insights, so you can focus on what matters, growth strategy.
g logo
fb logo
big query data
youtube logo
power logo
looker logo