SOURCE-LINKED INTELLIGENCE
Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit
arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC
Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evid
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
First collected: 2026-09-20T22:31:48.298Z. This is not the publication date.
Observed changes
AIIC observation times, not verified publisher revision times. Up to eight recent revisions.
2026-09-25T19:42:45.799Z
- title:
Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources → Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit - summary:
Whether repeated identical buying questions exhaust a language model's brand recommendations depends on retrieval. Across 300 question-engine cells (50 questions, six engines, 15 runs each, open extraction over 1,470 adjudicated organizations), the five engines answering without web search were still adding never-seen brands at run 15 in 86-92% of cells, with median repertoires of 15-31 organizations; the one retrieval-enabled engine closed its list (median 8 organizations, 64% of cells still adding), matching four earlier deep cells where web-search runs saturated by run ten. Cited-domain acc → Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evid