How to Build a Prompt Set That Actually Measures AI Visibility
A usable prompt set has 30–50 prompts weighted toward the decision stage, contains zero branded prompts in the core index, is stable enough to trend over time, and covers the questions buyers actually ask rather than the keywords you already rank for. Most prompt sets fail on the first and second of those.
The same prompt can return two answers ten minutes apart. Tools then diverge on sampling, geography, login state, provider path, and what they count as a mention. Compare trends, not one-day percentages.
Adding and deleting prompts every sprint makes week-eight incomparable to week-one.
The trend line is the product.
A rotating set is a series of one-off anecdotes.
Freeze the core.
When you must change it, version the file with a date (prompt-set-v2026-08-21.csv) and note which IDs changed.
Too small — can n=10 tell signal from noise?
The same prompt, same engine, minutes apart, can name a different vendor.
That is how the models work.
With ten prompts you cannot tell a real loss from a redraw.
Thirty to fifty core prompts is the range that stays cheap enough to re-run and large enough to trend.
Two hundred is mostly cost.
Confirm you are actually invisible before you industrialize the set — the 7-layer diagnostic is the pre-check.
What is the three-stage distribution model?
LLMConquer tags prompts tofu, mofu, and bofu.
Those map to Awareness, Consideration, and Decision — the same topic names used in onboarding.
Stage
Tag
Prompt shape
Example
What it tells you
Suggested share (commercial)
Awareness
tofu
Definitions and category framing
"What is the difference between an AI mention and an AI citation?"
Does the model explain the space using language you can live with?
25%
Consideration
mofu
How-to, criteria, comparison frames
"What should I look for when comparing AI visibility platforms?"
Are you in the shortlist logic?
30%
Decision
bofu
Best-for, vs, alternatives, plan-fit
"What is the cheapest way to monitor ChatGPT and Perplexity weekly?"
Do you get the recommendation?
45%
Onboarding in the product defaults to 40 / 40 / 20 (Awareness / Consideration / Decision) because a first workspace needs category coverage, not only bottom-funnel queries.
You can change those weights.
For a commercial core index I still recommend Decision-weighted, because that is the leading indicator a revenue team can act on.
Brand names belong on Decision prompts only when you are deliberately testing a branded index.
In onboarding, generated prompts follow the same rule: no brand name unless the stage is BoFu.
Where should you source prompts?
Six methods, ranked by yield.
1. Sales call recordings
Highest yield.
Buyers already asked the question in their words.
Transcribe the last 20 discovery calls.
Extract questions, not your answers.
One intent per line.
2. Competitor gap
Prompts where a competitor is named and you are not.
This is the fastest way to stop measuring vanity.
In LLMConquer this is the competitor mode in AI Research (Pro).
3. Support tickets and on-site search
People who already know you still ask comparison and "does it do X" questions.
Those often belong in Consideration, not Awareness.
4. Existing keyword data, rewritten
Useful as a backlog, harmful as a paste.
Rewrite rules:
Turn the fragment into a grammatical question
Add the job to be done ("for an agency", "for five engines")
Remove your brand from the core index
Drop anything you would never say out loud
5. AI-generated variations of prompts you already trust
Once you have 15 good seeds, generate paraphrases and then semantically dedupe.
String-unique is not unique.
This is the similar mode in AI Research.
6. Category surface
Ask the engine what it volunteers about the industry with no vendor list in the prompt.
Capture the questions it implies.
This is the niche mode in AI Research.
What are the hygiene rules?
One intent per prompt.
No brand name in the core index. Keep a separate branded index.
Natural phrasing, not keyword syntax.
Locale and language set explicitly (en-US, en-GB, fr-FR). Do not assume.
Deduplicate by meaning, not by string.
Freeze the core set. Version any change with a date.
Prompt hygiene — reject the line if any box fails
Would a buyer say this out loud?
Does it ask exactly one thing?
Is the brand name absent (core index) or intentional (branded index)?
Is the locale written down next to the prompt ID?
Is there a near-paraphrase already in the set?
Will I still want this line in the set 90 days from now?
How many prompts do you actually need?
Variance is the constraint.
One answer is a draw from a distribution, not a census.
n=10 cannot separate a 8-point real drop from a bad afternoon.
n=30–50, re-run daily or weekly, with a 7-day rolling mention rate, is enough for most teams to see direction.
n=200 multiplies engine cost and refresh time.
You learn more from a frozen 40 than from an ever-growing 200 you cannot afford to re-run.
Starter plans in this category often cap prompts (LLMConquer Starter includes 25).
If you are under a cap, cut Awareness first, not Decision.
How often should you re-run, and what is a real change?
Daily is enough for a 30–50 set on the engines you sell against.
Hourly sampling changes the number without changing the decision.
A real change is a shift that survives a rolling window and shows up in relative share, not only in one engine's one-day mention %.
Absolute percentages disagree across tools.
That is expected.
Compare trend and rank order.
Do not fire a vendor — or a teammate — over a single-day dip.
What is generative engine optimization and who is it for?
Awareness
tofu
How do AI answer engines choose which brands to mention?
Awareness
tofu
What is the difference between an AI mention and an AI citation?
Awareness
tofu
Why would a brand track ChatGPT visibility instead of only Google rankings?
Awareness
tofu
Which AI engines matter for B2B software buyers in 2026?
Awareness
tofu
What does share of voice mean in AI search results?
Awareness
tofu
How do large language models decide which vendors to recommend?
Awareness
tofu
What sources do AI engines usually cite for software comparisons?
Awareness
tofu
Is appearing in ChatGPT different from appearing in Google AI Overviews?
Awareness
tofu
What problems does AI visibility tracking actually solve for marketing teams?
Consideration
mofu
How do teams measure whether ChatGPT recommends their category of software?
Consideration
mofu
What should I look for when comparing AI visibility platforms?
Consideration
mofu
How many prompts do I need to track before AI visibility data is usable?
Consideration
mofu
What is a reasonable weekly refresh frequency for AI prompt monitoring?
Consideration
mofu
How do I tell a real drop in AI mentions from normal model variance?
Consideration
mofu
Should I track branded prompts and category prompts in the same index?
Consideration
mofu
How do agencies report AI visibility across more than one client brand?
Consideration
mofu
What engines should a mid-market SaaS team monitor first?
Consideration
mofu
How do citation gap reports differ from mention-rate dashboards?
Consideration
mofu
What does a good AI visibility prompt set look like for B2B software?
Consideration
mofu
Can I measure AI visibility without paying for an enterprise platform?
Consideration
mofu
How should I weight awareness versus decision prompts in a tracker?
Decision
bofu
What is the best AI visibility tool for a 50-prompt daily tracker?
Decision
bofu
Which AI visibility platform is most cost-effective for five engines?
Decision
bofu
What is the best alternative to Profound for mid-market teams?
Decision
bofu
Otterly vs Peec AI: which is better for citation tracking?
Decision
bofu
Which AI visibility tool should an agency use for multiple brands?
Decision
bofu
What is the cheapest way to monitor ChatGPT and Perplexity mentions weekly?
Decision
bofu
Which platform is better if I already pay for Semrush?
Decision
bofu
Which platform is better if I already pay for Ahrefs Brand Radar?
Decision
bofu
What AI visibility tool has an API suitable for a custom dashboard?
Decision
bofu
Who should buy an AI visibility platform versus running prompts manually?
Decision
bofu
What is the best AI visibility tool for a team that needs competitor gap analysis?
Decision
bofu
Which AI visibility platforms include sentiment on brand mentions?
Decision
bofu
What should I ask a vendor about sampling location and login state before I buy?
Decision
bofu
Which tool is better for tracking Google AI Overviews alongside ChatGPT?
Decision
bofu
How do I choose between a suite add-on and a dedicated AI visibility product?
Decision
bofu
What is the best starting plan if I need 25 prompts and four engines?
Decision
bofu
Which AI visibility tool is best if I need exportable citation data?
Decision
bofu
What should a first AI visibility stack include besides a prompt tracker?
Replace the Decision rows with your category ("HRIS", "warehouse WMS", "payment orchestration") before you treat this as your index.
Keep the shape.
If you have not confirmed a real gap yet, start with the 7-layer diagnostic.
Frequently asked questions
Should I track competitor brand names?
Yes, but not inside the core unbranded index. Keep a separate competitive set for "[competitor] alternative" and "[competitor] vs [other]" prompts. Mixing them into the core index makes share of voice look like a brand-awareness score.
How do I handle multiple markets?
Freeze one prompt set per language and locale. Do not silently reuse an English set against a French or US-vs-UK engine path. Geography changes answers. Label the set with country and language in the filename and in the tool.
Do prompts need search volume?
Volume is useful for prioritization, not for validity. A low-volume question from a sales call can be a better Decision prompt than a high-volume keyword that nobody asks an AI engine. Use volume as a tie-breaker, not as the source list.
What about long-tail prompts?
Long-tail is where recommendation happens. Keep them if they are one intent, naturally phrased, and stable. Drop near-duplicates. A 40-prompt set with 12 unique intents and 28 paraphrases is still n=12.
Should the core index include my brand name?
No. Brand-name prompts belong in a second index so you can still watch "does ChatGPT know us" without polluting "does ChatGPT recommend us."
How often should I change the set?
Almost never. Version the core set with a date when you add or retire a prompt. Changing ten prompts every month destroys the only number that matters — the trend.
Twelve AI visibility platforms compared on a normalized cost: monthly price divided by included prompts × engines × refreshes per week. Prices checked 21 Aug 2026.
Work access, rendering, entity, corpus, competitive, prompt-fit, and freshness in order. Most ChatGPT invisibility is a configuration problem, not a content problem.