Ask for a dated AI citation baseline before you sign anything. That means a fixed list of real buyer questions, run more than once across the platforms that matter to your business, with the cited source recorded for each one.
If what an agency shows you instead is a single screenshot or a self-scored “visibility score” with no prompt list attached, that is not a baseline. It is a sales aid dressed up as data.
Key Takeaways
- A real AI citation baseline is a document you can check yourself later, not a number an agency states out loud on a call.
- 3 outcomes exist for any buyer question, not 2: cited as evidence, cited as a named option, or misrepresented and omitted. Most reports only track 1 of the 3.
- The same baseline, re-run 6 months later, is not measuring your progress against a fixed target. The platform itself has moved in that window too.
- Ask for the request in writing before the proposal is final, then run 2 or 3 of the exact prompts yourself.
- A provider asking for a long-term signature before showing any baseline is asking you to trust a result you have never seen.
What a Real AI-Citation Baseline Actually Proves
A baseline is a snapshot: these specific questions, against these specific platforms, on this specific date, with what got cited written down. Nothing more.
Its entire value sits in being checkable later. Without a dated starting point, no report can honestly claim your visibility “improved.”
Most agency reports collapse a buyer question into 2 outcomes: cited, or not cited. A real baseline tracks 3.
- Cited as evidence. Your page supports a general claim the AI is making, without your brand being named as an option.
- Cited as a named option. Your brand itself is one of the choices the AI is actually recommending.
- Misrepresented or omitted. The AI answers the question, names competitors, and either skips your brand or describes it inaccurately.
A brand that scores well on evidence citations but never appears as a named option is not winning the questions that actually drive business.
That gap is a real citation-rate and share-of-voice problem, not a vanity number. A competitor benchmark inside the same baseline is what makes it visible.
We flag this split for every client before a contract starts. A single combined “mentions” number hides exactly the gap a buyer needs to see.
The 7 Fields a Real AI Citation Baseline Must Show
Ask for these fields specifically. A proposal missing more than 1 or 2 of them is not a real AI citation baseline yet, no matter what it is labeled.
| Field | A Real Baseline Shows | A Fake One Skips |
|---|---|---|
| Prompt set | The exact buyer questions used, in full | “A range of relevant queries” |
| Prompt source | Whether the questions came from real buyer language | Generic, agency-invented phrasing |
| Platforms tested | Each one named (ChatGPT, Perplexity, AI Overviews, Gemini, Claude) | “AI platforms,” left unnamed |
| Run count | Each prompt run more than once, with dates stated | A single screenshot |
| Citation record | Evidence, option, or omitted, marked per prompt | 1 combined score |
| Competitor set | Who got cited instead of you | Your numbers shown alone |
| Raw data access | Prompt-by-prompt results you can inspect | A dashboard summary only |
Why the prompt set matters more than the score
A visibility percentage built from questions nobody in your market actually asks is measuring the wrong thing precisely. It can read well and still tell you nothing about your real buyers.
Why raw data beats a summary chart
A dashboard chart is a claim. A prompt-by-prompt list you can open and check yourself is evidence.
Google’s Search Console now has a dedicated reporting view for AI-generated features. Ask whether a proposed baseline also pulls from that first-party source, not only from a third-party tool.
Why the Same AI Citation Baseline Looks Different a Month Later
This is the part almost nobody explains. It changes how you should read any 2 baselines placed side by side.
An AI platform is not a fixed ruler. Between your first baseline and a re-run, the model itself can change.
Its index can refresh, and the sources it trusts can shift. Comparing 2 baselines across time measures your progress and a moving target at once, not your progress alone.
A second pattern sits underneath this one.
- Broad, high-competition questions (“best SEO agency”) tend to see their cited sources shift often between runs.
- Narrow, specific buyer questions (“SEO agency for a 12-person law firm”) hold steadier across the same runs.
A baseline built mostly from broad head terms will look noisy no matter how good your content is. One anchored in specific buyer language holds steadier and tells you more.
What this means for you: ask any agency how it separates real movement in your citations from ordinary drift in the platform itself. An answer that treats every run as a clean, apples-to-apples comparison is missing a real limitation of this entire category.
The Exact Request to Send Before You Sign
Do not ask “how do you measure AI visibility.” That invites a marketing answer, not a document.
Send this instead, whether it is an email, a call, or a formal RFP for a GEO agency, before the contract is final.
“Before we sign, send us a dated AI-citation baseline: 15 to 25 real buyer questions specific to our business, run across [name your priority platforms], each run more than once, with the cited source marked as evidence, a named option, or omitted for every question. We want the raw prompt-by-prompt results, not a summary score.”
1. Send it before the final call, not during it
Building this properly takes real time. Asking for it live on a sales call almost guarantees a rough guess dressed up as a result.
2. Name your own priority platforms
If your buyers mostly use ChatGPT and Google, say so directly. A generic “test some AI platforms” request lets a provider pick whichever one makes the numbers look best.
3. Ask who actually wrote the prompt set
The strongest answer is a joint list. The agency proposes questions from its own research, and you add or correct them from language you already know your buyers use.
A prompt set the agency alone controls is one it can quietly tune to flatter itself. Who actually defines that prompt set is the single most revealing question in a full agency vetting pass, not just this one.
Real Answer Versus Deflection
| You Ask | A Real Answer | A Deflection |
|---|---|---|
| “Can I see the full prompt list?” | Sends it, or builds it with you before signing | “Our methodology is proprietary” |
| “Which platforms did you test?” | Names them individually | “We check the major AI platforms” |
| “How many times was each prompt run?” | Gives a number, explains why runs can vary | Never mentions repeat runs |
| “Can I run 2 or 3 of these myself, right now?” | Walks you through it live | Discourages it or delays |
| “What did competitors get cited for?” | Shows a comparison, unprompted | Shows only your numbers |
| “Is this included, or billed extra?” | States it plainly, in writing | Stays vague until after you sign |
Treat 2 or more deflection answers as a real signal, not a small thing to smooth over. A provider that shows its work before being paid is telling you how it behaves once it is.
Verify It Yourself Before You Trust Their Report
You do not have to take an AI citation baseline on faith. 2 checks take under 10 minutes combined and need no paid tool.
Run a few of the exact prompts yourself
Open ChatGPT or Perplexity directly and type 2 or 3 of the proposed buyer questions.
Compare what you see against what the agency’s document claims for those same prompts, on roughly the same date. A close match builds real trust, and a wide gap deserves a direct question before you sign.
Check whether your own site can even be seen
A baseline showing near-zero citations sometimes reflects a blocked site, not weak content.
Confirm your robots.txt allows GPTBot, OAI-SearchBot, PerplexityBot, and ClaudeBot. Pull your own server logs for hits from those crawler names.
A competent provider checks crawler access before running a single prompt, since a block can quietly explain a bad result that has nothing to do with your content.
The AI Citation Baseline Authenticity Checklist
Score whatever you are handed against this list. It grades the document itself, narrower on purpose than a full agency vetting pass.
- Dated, with a specific run window stated, not just “recent”
- Prompt set shown in full, not summarized
- Each prompt tied to a real buying question, not a generic topic
- Platforms named individually, at least 3 for most businesses
- Each prompt run more than once, with the variation shown
- Citations marked evidence, named option, or omitted, not combined
- At least 1 competitor’s results shown beside yours
- Raw, per-prompt data available on request
- No guaranteed-outcome language attached anywhere
- A stated plan for when it gets re-run, so it becomes a trend, not a one-off
A document clearing 8 or more of these is genuinely usable. Below 6, treat it as a sales aid built to resemble measurement.
What It Should Cost, and When You Should See Movement
Running a genuinely useful prompt set across several platforms, more than once per prompt, takes real time and real tooling cost. A rock-bottom price on an AI citation baseline audit often signals a thin sample size: a single, unreliable prompt run standing in for real measurement work.
What matters is not whether it is free. It is whether the cost and scope are stated plainly before you commit to anything bigger.
We price a real GEO audit to reflect what a genuine, repeated-sampling baseline actually takes, and we state that cost upfront, not after you sign.
Expect early citation movement to show over months, not days. A single re-run proves nothing on its own, given the non-determinism covered above.
Treat a provider promising fast, visible citation gains the same way you would treat a guaranteed AI citation or a guaranteed ranking: as a claim nobody honestly controls.
If They Will Not Show You One
2 different refusals to show an AI citation baseline mean 2 different things.
- “We build it after you sign, as the first deliverable.” This can be legitimate for a business-specific prompt set that genuinely needs your input. Ask for a sample from a comparable past client, identifying details removed, so you can see the format first.
- “We do not share our methodology.” A method you cannot see is a method you cannot verify. Treat this specifically as a warning sign on its own, separate from the answer above.
Either way, do not accept a long-term signature as the cost of finding out. A required lock-in before any real evidence exists is its own item worth flagging on a full vetting pass.
How GVM Technologies Builds This Before Anyone Signs
We build a dated AI citation baseline before we bill a single hour.
It uses real buyer questions run across the platforms that matter for that business, with every citation marked evidence, option, or omitted.
We show it to the prospect first, so nobody has to take our starting point on faith. What a real first 30 days with a GEO agency does with that same baseline is a fair follow-up question, since a baseline that never gets re-run after signing is only half finished.
Our published case studies carry named clients and dated results for the same reason. A claim you cannot check is not proof, whatever it is attached to.
Book a Free Strategy Call and we will walk you through a sample baseline for your industry before you commit to anything. If a vague “AI visibility” promise burned your business before, this is likely the exact step that was missing.
FAQs
1. What exactly counts as a real AI-citation baseline?
A dated document listing the specific buyer questions tested, the platforms each one ran against, and how many times each prompt was run. It also shows which source got cited for every question, marked as evidence, a named option, or omitted, with raw results available on request.
2. How many prompts does a legitimate AI citation baseline need?
15 to 25 real buyer questions is a reasonable floor for most small and mid-size businesses. The exact count matters less than whether the questions reflect language your actual buyers use.
3. Should I pay for the baseline before signing a contract?
Sometimes, and that alone is not a red flag, since building one properly across several platforms takes real time. What matters is that the cost and scope are stated plainly, in writing, before you agree to anything larger.
4. What is the difference between an AI citation and an AI recommendation?
A citation can support a general claim as evidence, without your brand ever being named as an option. A recommendation is when your brand is one of the choices the AI actually names.
5. Can an AI citation baseline guarantee my citations will improve?
No single provider controls a non-deterministic AI system closely enough to guarantee that, since the same question can surface a different cited source on separate runs. A baseline proves where you started, not where you end up.
6. What if an agency won’t share its baseline methodology at all?
Treat an outright refusal differently from a provider that simply wants to build your specific baseline after signing. A provider unwilling to let you verify how it measures your own results has not earned that trust.
Conclusion
The 1-line test before you sign anything: ask to see the raw AI citation baseline, not the score.
A document you can check yourself, dated and built from real buyer questions, is measurement. A number with no document behind it is a sales aid.
Book a Free Strategy Call and see a real, dated baseline for your business before you commit to anything, or read the full GEO and SEO agency due-diligence checklist for the other 17 questions worth asking before you sign.
