Wonder Lab
Wonder LabWonder Lab
  • Blog
  • Products
  • Podcast
  • Resources
  • About
Subscribe
BLOG

Knowledge Share

Technical articles, tutorials, and insights

Found 2 posts
AI EvaluationDeepEvalRAGAS

AI Evaluation Series (06): DeepEval in Practice — Enterprise Agent Evaluation Suite

Running DeepEval on the same customer support Agent: AnswerRelevancy (avg 0.767, 60% pass), Faithfulness (avg 0.600, 40% pass), ToolCorrectness (20% pass — Agent's failure to call tools is the root cause). Key focus: the paradigm difference between DeepEval and RAGAS. RAGAS is a batch analysis tool; DeepEval is a CI quality gate tool. They solve different problems. Also covers how to connect glm-4-flash as the DeepEval Judge LLM.

2026-07-16·5 min read
AI EvaluationAgent EvaluationTool Call Accuracy

AI Evaluation Series (05): Agent Evaluation — Tool Call Accuracy and Trajectory Quality

Systematic evaluation of a customer support Agent (3 tools, 15 test cases). Results: tool name accuracy 73%, parameter accuracy 100%, step efficiency 0.73x. All 4 failures are 'should have called a tool but didn't' — not 'called the wrong tool.' The bottleneck is tool triggering logic, not tool execution. Trajectory quality scored 3.73/5 by LLM-as-Judge, with multi-step and edge cases performing slightly worse.

2026-07-15·6 min read
Wonder Lab
© 2026 Dongqi Chen · Wonder Lab
AboutRSSSitemap