CrowlyCrowly
GEO

How to build an AI-visibility monitoring dashboard without a paid tool — the complete manual method

You can monitor your AI visibility manually with a spreadsheet and 2 hours a month. Here's the method, the template, and the limits of manual monitoring — and when it stops being enough.

Crowly5 min read
Computer screens displaying monitoring data

Not every company is ready to subscribe to an AI-visibility monitoring tool right away. It could be a budget constraint, it could be skepticism that needs to be validated with data before justifying the investment, it could simply be the company's stage of digital-marketing maturity. Whatever the reason, there's a way to start monitoring at no cost — with clear limitations, but with real value as a starting point.

This post describes the complete manual-monitoring method, including how to structure the spreadsheet, which queries to use, how often to run it, what to document and, crucially, when the manual method stops being enough.

The monitoring template: the spreadsheet structure

Create a spreadsheet with the following tabs:

Tab 1 — Monitored queries. A list of the queries you'll test, with a "category" column (awareness / consideration / decision), a "priority" column (high / medium), and an "added on" column. Start with 10 queries and expand as the routine settles in.

Tab 2 — Weekly log. Columns: Date | Query | Platform (ChatGPT / Gemini / Perplexity) | Did your company show up? (Y/N) | Position (1st, 2nd, 3rd mention) | Competitors that showed up | Notes (answer context, sentiment, changes noticed).

Tab 3 — Calculated score. Each week, calculate the appearance percentage: the number of queries you showed up on ÷ total queries tested × 100. Break it down by platform. That's your weekly score — the historical series is the most valuable data.

Tab 4 — Competitor benchmark. Same structure as Tab 2, but focused on 2 or 3 main competitors. Test the same queries for them and document the same fields.

The test protocol: how to run it to avoid bias

Always in incognito mode. Use an incognito tab or a browser profile with no history. AIs personalize answers based on history — you want the default answer for the generic user.

Always with no account logged in. On ChatGPT especially, the answer with an account logged in can differ from the answer with no login.

Use the same query wording. Variation in wording can change the answer significantly. Keep the queries exactly the same across every test for comparability.

Document the full answer when there's variation. Not just "showed up or not" — when the answer changes in an interesting way, record the specific sentence where your company was mentioned (or not) and the context.

The ideal frequency and the effort per round

Recommended frequency: every two weeks. Weekly is better for detecting variations, but 2-3 hours a week is a lot for small teams. Biweekly balances granularity and effort.

Time per round: with 10 queries tested on 3 platforms, that's 30 tests per round. Each test takes between 1 and 3 minutes (including documentation). Total: 45 to 90 minutes per biweekly round.

Splitting the work: at companies with more than one marketing person, this routine works best with a fixed owner — the consistency of who runs it reduces interpretation variance in the records.

What to look for beyond presence and absence

The most valuable data from manual monitoring isn't the score itself — it's the patterns that emerge over time:

Position variation. Do you show up first today and third the next week? Or is the position consistent? Position instability can indicate your presence is marginal — the model is uncertain about whether to include you.

A change in narrative. Has the way the AI describes your company changed? Has it started using different terms, framing you differently? Those changes usually reflect changes in the indexed content about you.

Who entered or left. Competitors that appear and disappear on the same queries you compete on reveal strategic moves worth investigating: what did they publish that made the model include them?

The limits of manual monitoring — and when to move to a tool

Manual monitoring has structural limitations that become evident as the strategy matures:

Scale. With 10 queries and 3 platforms, you have 30 data points per round. A mature strategy may need 50 to 100 queries to adequately cover the universe of relevant queries — making manual impractical.

Time-of-day consistency. AI answers can vary by time of day — tests run at different times introduce uncontrolled variation.

No automated history. Maintaining the historical series manually is labor-intensive and prone to lapses. A tool maintains it automatically.

No alerting. Manual monitoring doesn't warn you when something changes — you only find out on the next round. A tool can alert you to sharp drops.

When manual monitoring starts taking more than 4 hours a month, or when content decisions need finer granularity than manual offers, moving to an automated tool has a clear ROI.

How Crowly solves what manual can't

Crowly automates the entire process described in this post — with the advantages of scale (hundreds of queries monitored), consistency (tests run to the same standard every week), automated history, and change alerting. The free initial diagnostic offers the same data as the first round of manual monitoring — without the work.

Start with the free diagnostic to get your baseline immediately. Then decide whether continuous monitoring makes sense. Free diagnostic →

Sources:

Keep reading