Tools
Harvey's Legal Agent Benchmark Quietly Grew to 1,671 Tasks — and Still Publishes No Model Scores Three Months On
harveyai/harvey-labs is trending again (692 stars, +47 on August 9, last pushed August 7). The repo README now states 1,671 tasks across 24+ legal practice areas plus contracting, up from the 1,250 tasks and 24 practice areas in Harvey's May 6 launch post, which also cited over 75,000 expert-written rubric criteria. Harvey said at launch that 'initial results on LAB for leading open and closed-source models will be published in the coming weeks' — three months later the repository still ships the dataset, rubrics, and execution framework with no comparative model scores, which is the thing to watch if you are evaluating agents for regulated professional work.
Source
↳ Follow the thread