Skip to content
BroadcastwellExchange
Sign in
Menu

Discussion

What would make you trust an AI visibility number?

Sairam SivakumarBroadcastwell team, Founder

My starting list: published questions, named engines, stated dates, the valid base and interval, readable answer records and clear exclusions.

I also want to know who reviewed the answers and what the measurement cannot establish. A precise number can still answer the wrong question.

What is missing? If a report met this list, what would still stop you from using it for a decision? One check and the reason for it is a complete answer.

Source: broadcastwell.com/methodology

The author used AI assistance.

Replies

  1. Reply by Yamac Nova Kurul

    Yamac Nova Kurul

    Adding an engineering perspective to your list: Session Isolation and State Reproducibility.

    Beyond the stated date, engine, and base interval, I would need to trust that the measurement was executed in a completely clean, zero-shot environment. If an API call or a scraper session retains any cached context or lingering tokens from previous queries, the LLM's output and citations will be skewed by hidden prompt-chaining bias.

    In my recent work architecting an AI Copilot backend and building API workflows, ensuring data integrity meant treating every request as an idempotent operation. For an AI visibility number to be truly trustworthy and reproducible, the methodology must explicitly confirm that every single run (whether it's run 1 or run 3) was executed in a strictly isolated, fresh state. If the testing environment isn't clean, the number isn't reliable.

    1. Reply by Sairam Sivakumar

      Sairam SivakumarBroadcastwell team, Founder

      Yamac, I would add a visible session-state section to the receipt: new or reused conversation; prior turns, if any; signed-in or signed-out state; memory or personalization settings where they can be inspected; the engine experience; and any controls that were unavailable or unknown. No account identifiers, cookies or tokens belong in a public receipt.

      I would also separate a repeatable collection procedure from a promise of identical answers. My proposed check is to repeat one exact question under a documented fresh-session procedure, retain each answer separately, and record the differences instead of treating consistency as a pass/fail label.

      For a first worked example here, which two checks would best demonstrate that prior conversational context was excluded? A hypothetical example, or a public answer you have permission to share, would be enough to start.

      The author used AI assistance.

  2. Reply by Priya Nair

    Priya NairBroadcastwell team, Management Consultant, Delivery Lead

    The first thing I ask is what the number actually measures. I want the exact question set, the engine or surface, the capture date, the eligible-answer denominator, the scoring rule, and whether the questions were run more than once. A percentage without that context can look precise while hiding a very narrow test.

    If I had to pick one missing detail that makes a visibility number unusable, it is the sample definition: which questions were asked and which answers counted. Repeat runs are a close second because these systems are not deterministic. I also want results separated by engine rather than blended too early. A useful number should let me reconstruct the experiment well enough to understand what changed, what stayed fixed, and whether I would expect a similar result if we ran the same test again.

    The author used AI assistance.