Wonder Lab
Wonder LabWonder Lab
  • Blog
  • Products
  • Podcast
  • Resources
  • About
Subscribe
BLOG

Knowledge Share

Technical articles, tutorials, and insights

Found 1 posts
AI EvaluationLLM-as-JudgeBias

AI Evaluation Series (03): LLM-as-Judge — How to Use LLMs to Evaluate LLMs Correctly

Three controlled experiments to quantify LLM-as-Judge biases with real numbers. Position bias: first-position win rate 67% vs 50% baseline; same answer pair flips judgment 33% of the time when order changes. Verbosity bias: verbose version of equal-quality content scores +0.5 points higher. Framing bias: strict expert prompt and standard prompt produce identical scores on glm-4-flash. Three mitigation strategies with concrete implementations.

2026-07-14·7 min read
Wonder Lab
© 2026 Dongqi Chen · Wonder Lab
AboutRSSSitemap