Skip to content
BroadcastwellExchange
Sign in
Menu

Contribution record

This is the member's own record of contributions and responses, to share if they choose. Taking part in the Exchange does not affect applications. It is not a credential or a hiring decision.

Reply in "Same question, different engines: what stays fixed?"

Author
clane
Posted
Revision
Original version
Thread
Same question, different engines: what stays fixed? (Discussion)

The contribution

One additional control I would add is run order. In a hypothetical comparison, querying every prompt on engine A first and engine B hours later can mix an engine difference with a time difference. I would interleave engines within a narrow capture window and record the order and timestamps, with a preset repeat count.

I would not claim to hold hidden retrieval indexes, experiments, model updates or undisclosed personalization fixed. Mark those as uncontrolled or unknown, retain every run, and describe the result as a comparison of the recorded product experiences. API and consumer-app outputs should also remain separate unless the study explicitly compares those surfaces.

AI-assisted response; proposed protocol, not a completed measurement.

The team reply

For a cross-engine comparison, I hold the input and the test conditions as fixed as the products allow: exact prompt text, language and locale, category definition, fresh-session state, authentication or personalization state, capture window, repeat count, and the rule we use to score a name. I also record whether an engine was using web retrieval or a different search mode, because that can change the task materially.

The control people forget most often is conversation state. A "same question" test is not the same test if one engine receives a clean prompt and another has prior chat context, memory, or personalization shaping the answer. I would rather accept that engines have different architectures than accidentally introduce avoidable differences in the setup. The goal is not to make the products identical. It is to make our comparison protocol explicit enough that someone else can understand what was held constant and what could not be.

A contribution record shows what was posted publicly on the Broadcastwell Exchange, by whom and when. It is not a certificate or an endorsement.