Every number we publish, and how to check it
We sell measurement, so ours has to be inspectable. This is the current state of our own AI visibility, the results we got on sites we do not control, and an explicit list of what our numbers cannot tell you. Wins and losses at the same size.
3 of 4 engines identify and cite us by name
The branded question — “What is CitePack?” — asked of each engine three times, first-party rather than through a middleman. Every individual answer stored.
Perplexity
Identifies us correctly and cites citepack.com in all three samples.
Google (Gemini)
The first engine to resolve us, and steady since.
ChatGPT
Moved from a name collision to citing us, roughly three weeks after the identity work shipped.
Claude
Still reaches a 2007 academic paper of the same name. An identity problem, not a content one.
And the number that matters more: asked for the best tool in our category without naming us, we appear 0 of 12 times. Recognized, not recommended. It is the harder rung, it is earned mostly off your own domain, and we are not going to claim it before we have it.
Measured 2026-07-28 · 36 probes · m=3 per query family. Small samples, reported as counts and never as rates or trends. An engine we could not probe is recorded as not measured, never as a zero.
Seven sites, measured twice, two weeks apart
A measurement you can only ever pass is not a measurement. Before trusting our own numbers we ran the instrument across seven live third-party sites and published whatever came back — 28 engine-level readings in total.
One reading improved. Two got worse. The rest held.
The same question, two minutes apart
Answer engines are non-deterministic. Asking an identical question twice on the same day, up to 60% of the cited sources come back different. This is the single most important fact about measuring AI visibility, and it is why we sample rather than check.
- vendor-a.com
- bigreview.com/best-tools
- vendor-b.io
- reddit.com/r/…
- vendor-a.com
- bigreview.com/best-tools
- anotherblog.dev/guide
- vendor-c.com
A tool that shows you a confident score from one question is not wrong occasionally. It is wrong in a way it cannot detect. The full method →
What our numbers do not show
- We cannot promise an engine will cite you.Nobody controls what a model retrieves. We measure the current state, ship the fixes that make you resolvable and quotable, and re-measure afterwards — reporting the result either way, including when nothing moved.
- We have not proven that our fixes cause citations.Our own recognition improved after our identity work shipped, and engines re-index on their own schedule, so the honest description is a sequence and not a proof of cause. Saying otherwise would be the exact overclaim this page exists to avoid.
- Small samples are counts, never trends.Three samples per engine per question tell you the state today. They do not support a percentage, a trajectory, or a forecast, and we do not render one.
- Unmeasured is reported as unmeasured.Never quietly counted as a no. On a report you paid for, a fabricated negative is worse than no answer.
- Everything here is stored and recomputable.Each probe — engine, question, sample, answer — is written to an append-only log, because an engine’s index cannot be rewound. Nothing on this page is an estimate, a projection, or a borrowed vendor statistic.
Now run it on your own site
The same four engines, the same three samples, the same rules about what counts. About 30 seconds, free, no signup.