
Separate discovery from comparison
A useful set includes category prompts, use-case prompts, alternative prompts, and direct comparison prompts. Mixing them together makes it harder to understand whether a visibility gap is about classification or preference.
Keep the test small enough to repeat
A focused set of 15 to 30 prompts is often more useful than an oversized list that no one reruns. Store the exact prompt wording, model, date, and relevant sources with every run.
Define what counts as a useful mention
Record whether the project appeared, whether the category was correct, whether the answer cited a reliable source, and whether the recommendation matched the buyer's question. A mention without accuracy is not a complete success.
Apply this to your project