Prompts are essentially search queries and conversational questions that deserve grounding during RAG.
So when someone sells you on โcreatingโ a list of prompts to track AI visibility, remember: thereโs nothing particularly creative about it. More importantly, youโre being asked to outsource your understanding of your own search demand to a tool.
If you have to make up your prompts, how do you know they reflect what users actually search for rather than synthetic noise generated by the tracking tool?
Your first-party search demand should inform your prompts, not the other way around.
Itโs a bit like a snake biting its own tail: prompts are generated from assumptions about what people might search for, then used to measure visibility against those same assumptions.
Without grounding them in actual search behaviour, the whole exercise risks becoming somewhat mythological โ rather like Janus Bifrons, looking in two directions at once.
In my latest post, I walk through a pipeline I built to identify synthetic queries an LLM sends to Google to validate the retrieval (i.e ; grounding) and how to map them across your actual search queries in Google Search Console.
๐ This piece goes deep into mapping prompt tracking searches to your real search queries.
But if you want to access a broader workflow to classify your web traffic, I provided an example on how to separate prompt tracking and scraping activity from queries consistent with ChatGPT grounding.
Your Prompts pump Fake Queries and Impressions into your Search Console
The problem starts when prompt trackers send templated queries to an LLM, and the model branches them out into Google to validate its answer.
Those synthetic searches can generate new query impressions in Google Search Console, while the resulting retrieval and citation activity can also trigger HTTP requests that get recorded in your CDN logs.
In other words, the very prompts you use to measure AI visibility can create new first-party data points that look like genuine user behaviour โ feeding fake queries and impressions back into the datasets you might later use to understand search demand.
Prompt Tracking is A/B Testing With No User Context
A prompt tracker may look like a consistent visibility measurement, but what happens under the hood is anything but controlled.
The model being used, how it interprets the prompt, whether and where it searches, the userโs account and context, and even an intrinsic dose of randomness can all change the response.
The same prompt can therefore produce different outputs depending on when, where, and how it is run.
Prompt tracking means effectively running A/B testing LLM systems with no access or valid records of the user context.
Compare Prompt and Query Embeddings with Cosine Similarity and find those that Deserve Grounding
And here is the workflow idea Iโve been working on lately.
Itโs six steps from a raw prompt tracker export to a validated list of prompts anchored in real search demand
- BigQuery โ data extraction and data cleaning
- Dataform โ transformation
- Python โ embeddings generation and cosine similarity calculation
Pull tracked prompts from Botify and Peec AI, labeling each by tracker provider so you will ultimately know where your grounded prompts came from.
A reusable JavaScript UDF strips accents, poor punctuation and non-ASCII characters from every raw GSC query in one pass.
An incremental table applies the UDF, drops sub-3 words queries and competitor noise, retaining the last 30-day searches.
A simple SQL query pulls only clicks = 0 rows from the mart. Reduce the date window if the export gets too large.
Both sets are converted to vectors with all-MiniLM-L6-v2 โ fast enough to run at scale on CPU inside Colab.
Every prompt is compared against every zero-click query pulled in step 4; each prompt keeps its single best match.
Keep matches with similarity โฅ 0.80, zero clicks and non-negative impressions
The Problem With Raw GSC Query Data
Before any matching can happen, the query data itself needs cleaning.
Google Search Console exports are messy at scale โ accented characters, mixed casing, stray punctuation, and competitor terms all dilute the semantic signal youโre trying to match against.
Running embeddings on dirty text produces noisy similarity scores, so the preprocessing step is the foundation the whole pipeline depends on.
Step One: A Reusable Text-Preprocessing UDF
Just like I did when mapping out AI Mode queries, the first building block is a JavaScript User-Defined Function (UDF) in BigQuery. Rather than repeating cleaning logic in every downstream query, this gets called once, as a reusable function, from any SQL query in the project:
if (query === null || query === undefined) return ""; var accentsMap = { 'รง':'c','รฆ':'ae','ล':'oe','รก':'a','รฉ':'e','รญ':'i','รณ':'o','รบ':'u', 'ร ':'a','รจ':'e','รฌ':'i','รฒ':'o','รน':'u','รค':'a','รซ':'e','รฏ':'i', 'รถ':'o','รผ':'u','รฟ':'y','รข':'a','รช':'e','รฎ':'i','รด':'o','รป':'u', 'รฅ':'a','รธ':'o','รฑ':'n', 'ร':'C','ร':'AE','ล':'OE','ร':'N' }; // 1. Transliterate known accented characters via the map var text = query.split('').map(function(c) { return accentsMap[c] || c; }).join(''); // 2. NFKD decompose + strip any remaining combining marks text = text.normalize('NFKD').replace(/[\u0300-\u036f]/g, ''); // 3. Lowercase text = text.toLowerCase(); // 4. Strip any remaining non-ASCII (emoji, CJK, etc.) text = text.replace(/[^\x00-\x7F]/g, ''); // 5. Remove numbers text = text.replace(/[0-9]+/g, ''); // 6. Remove punctuation (keep letters/spaces only) text = text.replace(/[^\w\s]/g, ' '); // 7. \w includes underscore โ strip it too text = text.replace(/_/g, ' '); // 8. Collapse whitespace text = text.replace(/\s+/g, ' ').trim(); return text; Why this matters:
- Handle accents first, then clean up: This makes common accents like รฉ โ e work, while also catching less common ones.
- Lowercase before removing special characters: This keeps the text consistent for matching.
- Remove numbers and punctuation: They add little value when comparing the meaning of queries, so we simply drop them. Do note that a question mark is useful to track down, but it can be resource expensive. This is a condition that must be levelled when you clean up the prompt dataset.
This UDF becomes then a single callable function that the mart layer below calls on every row.
Step Two: Building the Cleaned Query Mart
With the UDF in place, the next layer is a Dataform incremental table that applies the cleaning function, deduplicates, filters out noise, and lands a clean data mart:
config { type: "incremental", schema: "searchconsole", name: "pre_processed_gsc_query", uniqueKey: ["query", "data_date"]}WITH raw_queries AS ( SELECT data_date, query, SUM(impressions) AS impressions, SUM(clicks) AS clicks FROM `your_project.searchconsole.searchdata_url_impression` WHERE data_date BETWEEN DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY) AND CURRENT_DATE() AND search_type = 'WEB' AND is_anonymized_query = FALSE AND query IS NOT NULL GROUP BY query, data_date),cleaned AS ( SELECT `your_project.searchconsole.advanced_text_preprocessing`(query) AS query, data_date, impressions, clicks FROM raw_queries),deduped AS ( SELECT data_date, query, SUM(impressions) AS impressions, SUM(clicks) AS clicks FROM cleaned WHERE query IS NOT NULL AND LENGTH(query) >= 2 AND ARRAY_LENGTH(SPLIT(query, ' ')) >= 3 GROUP BY query, data_date)SELECT data_date, query, clicks, impressionsFROM dedupedWHERE NOT REGEXP_CONTAINS(query, r'noisy|competitors') What happens at each stage
raw_queriesโ pulls 30 days of web-search data and removes anonymised queries, which have no usable query text.cleanedโ applies the UDF from Step One to standardise each query.dedupedโ combines variants that become the same query after cleaning and removes queries with fewer than 3 words.WHERE NOT REGEXP_CONTAINSโ removes competitor queries that could skew content-gap analysis.
โ ๏ธ The last two points must be dictated by the type of dataset youโre dealing with; the choice is arbitrary.
And the result is one clean pre_processed_gsc_query table that downstream queries can use directly.
Step Three: Isolating the Zero-Click Query Pool
With the mart in place, pulling the actual candidate pool for gap analysis is trivial:
SELECT query, SUM(clicks) AS clicks, SUM(impressions) AS impressionsFROM `your_project.searchconsole.pre_processed_gsc_query`WHERE clicks = 0 AND data_date BETWEEN DATE_SUB(CURRENT_DATE(), INTERVAL 14 DAY) AND CURRENT_DATE()GROUP BY query Filtering by zero-click queries is the whole point of the mart, with the option to reduce the time frame when you have a very large GSC dataset.
Step Four: Semantic Matching in Google Colab
The BigQuery side hands off a clean, structured dataset.
In Google Colab, a Python script may:
- Pull the query mart into Colab โ I used Polars to retrieve and process >= 10 GB data in the environment.
- Sample the queries if the dataset is too large or the embeddings generation will kill your RAM.
3. Clean the prompts dataset, where you have a column with the prompt tracker provider (in case you are using several). You need to lowercase and remove punctuation.
4. Create embeddings for prompts and queries using all-MiniLM-L6-v2. Among all the models I tested to generate embeddings and calculate cosine similarity, Iโve always found all-Mini the most robust, lightweight and therefore convenient of all
5. Match each prompt to its closest query using cosine similarity.
6. Keep strong matches with similarity โฅ 0.80 and zero clicks.
And export to Excel with all matches and filtered opportunities.
โ ๏ธ Disclaimer โ All data in this output is fictional and provided for demonstration purposes only.
However, the output stat worth highlighting is what percentage of the tracking prompts is anchored to a real search query.
Thatโs a much more defensible metric to bring to stakeholders than raw prompt volume, because it ties AI-visibility tracking back to demonstrable search demand.
Prompt Tracking should start with your Own Queries
Prompt tracking tools are useful, but on their own they canโt tell you whether a tracked prompt reflects something people are actually typing into Google, or whether itโs a synthetic variation invented by the tracking tool itself.
Running prompts through this pipeline doesnโt solve the โegg-and-chickenโ riddle, but at least it gives you an update on your prompt tracking strategy.
The share of prompts matched to real, zero-click search demand is not only a signal of success with your tracking setup but can be viewed as an indicator of LLM brand recall.
The more grounded prompts you capture, the stronger your brand recall in LLMs โ and the more confidence you can have in your tracking setup.
Conversely, a low proportion of grounded prompts can indicate weaker brand recall and a potential gap between real user demand and what your tracker is surfacing.
Thatโs where you need to review your tracking prompts against your underlying search demand.
Itโs a snake biting its own tail: the prompts are the tail; your first-party search queries are the head. Your tracking should start with what users actually search for, not the other way around.