The Dice Roll Method Sets Iteration Counts for Repeated-Query LLM Audits, and Finds Fixed Tiers Do Not Transfer
arXiv 2609.04047 formalizes a protocol for the increasingly common practice of auditing LLM outputs with repeated identical prompts, which currently has no standard for iteration counts, stability metrics or reliability thresholds. Reanalyzing five brand-recommendation auditing studies covering roughly 190,000 observations, 270+ brands, 6 languages and iteration counts of 5 to 40, the D-study yields three tiers: exploratory at n=5 (G=0.58), confirmatory at n=10 (G=0.74) and rigorous at n=15 (G=0.81). A preregistered external validation on three independent corpora reproduced the reliability prediction in 37 of 39 cells with no failures, but the fixed tiers themselves did not transfer, supporting a pilot-then-solve approach rather than copying an iteration count from someone else's study.
↳ Follow the thread