How to Track AI Visibility Over Time Without Enterprise Software
You have been quoted five-figure annual fees for AI visibility monitoring platforms, and the business case is not obvious yet. The good news is that the core measurement discipline isn’t proprietary. What the platforms sell is convenience and scale — the method itself you can run with a spreadsheet, a fixed protocol and about two hours a month.
What are you actually trying to track?
Before choosing tooling, be precise about the output. A defensible AI visibility tracking programme produces four numbers per period, per assistant:
- Citation rate — prompts where one of your URLs appears as a source, divided by prompts tested.
- Mention rate — prompts where your brand name appears in the answer text, divided by prompts tested.
- Competitive share — your appearances as a proportion of all brand appearances across the same prompts.
- Accuracy count — of the answers mentioning you, how many describe you correctly.
Everything else is detail. If a manual process yields these four reliably and comparably over time, it’s doing the job an enterprise tool does.
Step one: build and freeze the protocol
Consistency matters more than sophistication. Trend data is only valid if the method is identical between rounds, so write the protocol down before the first round and treat changes to it as version releases.
Your protocol document should specify:
- The exact prompt list, verbatim, numbered and versioned.
- Which assistants you test, and in what order.
- Session conditions: logged out or in a fresh profile, no chat history, no memory or personalisation enabled, no custom instructions, location noted.
- How many times each prompt is run per round, and how you resolve disagreement between runs.
- What counts as a mention (exact name only, or common variants and abbreviations too).
- What counts as a citation (any of your domains, and whether subdomains and PDFs count).
The personalisation point deserves emphasis. If you test while logged into an account that has discussed your company before, you’re measuring the assistant’s memory of your conversation, not its view of the market. Always test in a clean session.
Step two: keep a simple, well-shaped spreadsheet
Structure the data as one row per prompt, per assistant, per run. Resist the urge to summarise as you record — store the raw observations and calculate the rates in a separate sheet.
Columns worth having:
- Round date and protocol version
- Prompt ID and prompt type (category, evaluative, comparison)
- Assistant and run number
- Brand mentioned (yes/no)
- Your URL cited (yes/no) and which URL
- Other brands named, as a comma-separated list
- All cited domains, as a comma-separated list
- Accuracy assessment (accurate / partial / wrong / not mentioned)
- A short note field for anything unusual
The “other brands” and “all cited domains” columns do the heaviest analytical lifting over time. They tell you who is occupying the answer space and which third-party sources the assistants keep returning to. That list is effectively your target list for corroboration work.
Step three: run rounds on a fixed cadence
Monthly suits most organisations. Weekly is usually over-sampling given how slowly the underlying signals move, and it burns the discipline out. Quarterly is too sparse to catch a problem forming.
Run the whole prompt set in one sitting where possible, so all observations share the same conditions. Batch by assistant rather than by prompt — switching tools repeatedly increases the chance of a protocol slip. Timestamp the round and record any model version information the interface exposes.
Three runs per prompt is a reasonable minimum for handling variance. If that’s too much volume, prioritise: run your top fifteen highest-value prompts three times and the long tail once, and note the difference in your protocol so you never compare the two segments as if they were equivalent.
Step four: reduce the cost with light automation
You don’t need a platform to remove the tedium. Several accessible approaches, in rough order of effort:
- Split the work. Assign assistants to different team members with the same protocol sheet. Ninety minutes each rather than four hours for one person.
- Use official APIs where they exist. Several assistants expose programmatic access, and a short script can run the prompt set and dump raw responses to a file for scoring. This gives you cleaner reproducibility than manual testing, though API behaviour and the web product’s retrieval can differ — don’t mix the two sources in one trend line.
- Automate the scoring, not the judgement. Simple string matching handles mention and citation detection reliably. Keep accuracy assessment human — it’s the part that requires knowing your business.
- Log server-side crawler activity. Your own access logs show which AI agents are fetching which pages. It’s not a substitute for answer testing, but it is free, continuous, and tells you whether the access layer is working.
Step five: read the data like an analyst
Three habits separate useful tracking from a monthly ritual.
Never react to one round. AI answers vary. Require two consecutive rounds moving in the same direction before you call it a trend, and hold at least three rounds before drawing conclusions about a change you made.
Annotate the timeline. Every publication, site change, robots.txt edit, press placement and partner announcement goes on the same chart as the metrics. Without annotations you’ve correlation with nothing to attach it to.
Segment before concluding. A flat blended number frequently hides a rise in one assistant and a fall in another, or gains on informational prompts alongside losses on evaluative ones. Always look at the segments before writing the summary.
When does a paid platform become worth it?
Be honest about the threshold rather than defaulting either way. A manual process stops scaling when you have several hundred prompts, multiple brands or markets, more than a handful of assistants, or a genuine need for daily rather than monthly resolution. It also stops working when nobody owns it — an unowned manual process silently degrades, and inconsistent data is worse than no data.
Below those thresholds, the manual method gives you something the platforms cannot: everyone involved understands exactly how the numbers are produced, which makes them far harder to dismiss in a leadership discussion.
If you want the loop run for you without an enterprise contract, our ARIA citation tracker holds the protocol constant across rounds. If tracking reveals a structural problem in your pages rather than a discovery one, the AI citability scorer is the next step, and our AI visibility page covers the broader programme.
Frequently asked questions
How often should I run AI visibility checks?
Monthly works for most organisations. The underlying signals move slowly, so weekly testing mostly measures variance, while quarterly is too sparse to catch problems as they form.
Do I need to run each prompt more than once?
Yes. AI answers are non-deterministic, so a single run mixes real signal with randomness. Three runs per prompt is a practical minimum, with a documented rule for resolving disagreement between them.
Can I just use my analytics to track AI visibility?
Only partially. Referral traffic from assistants shows that some citations converted to visits, but it cannot show the prompts where you were absent — and absence is the thing you most need to measure.
Does testing while logged in affect the results?
Significantly. Chat history, memory features and custom instructions all shape answers. Always test in a clean, logged-out or fresh session, and record the conditions in your protocol.
Get the next one in your inbox.
Practical insights — no fluff, straight to your inbox.
Or follow us on LinkedIn:
Follow StrategyPeeps






