Research
Benchmarkless Safety Scoring: Comparing LLM Safety When No Labeled Benchmark Exists
Formalizes the setting of comparing candidate models for safety before any labeled benchmark exists for the relevant language, sector, or regulatory regime. Specifies the contract under which a scenario-based audit can serve as deployment evidence, with validity depending on fixed scenario pack, rubric, auditor, judge, sampling, and rerun budget. Directly applicable to enterprise deployments in regulated industries without domain-specific safety benchmarks.
Source
↳ Follow the thread