Model rankings can narrow a candidate set, but an enterprise deploys a workflow on its own data, permissions, review capacity, and failure surface. The review object keeps the reported score separate from the evidence that the score transfers to that operating environment.
benchmark_score
The reported score, benchmark version, task split, and evaluation conditions.