1/ We use LLM judges to scale up costly human evaluation.
1/ We use LLM judges to scale up costly human evaluation. But to trust an LLM judge, you need… human evaluation. Our new preprint tackles this circularity: "Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability"