AI
Blind Your LLM Judge Before You Trust the Score
My eval harness leaked model names into the judge prompt for six weeks. What two recent papers say about judge bias and calibration, and the…
Rayyan |
August 16, 2026 |
7 min
Read More