Technical articles, tutorials, and insights
Three controlled experiments to quantify LLM-as-Judge biases with real numbers. Position bias: first-position win rate 67% vs 50% baseline; same answer pair flips judgment 33% of the time when order changes. Verbosity bias: verbose version of equal-quality content scores +0.5 points higher. Framing bias: strict expert prompt and standard prompt produce identical scores on glm-4-flash. Three mitigation strategies with concrete implementations.