Knowledge Share
Technical articles, tutorials, and insights
AI Evaluation Series (06): DeepEval in Practice — Enterprise Agent Evaluation Suite
Running DeepEval on the same customer support Agent: AnswerRelevancy (avg 0.767, 60% pass), Faithfulness (avg 0.600, 40% pass), ToolCorrectness (20% pass — Agent's failure to call tools is the root cause). Key focus: the paradigm difference between DeepEval and RAGAS. RAGAS is a batch analysis tool; DeepEval is a CI quality gate tool. They solve different problems. Also covers how to connect glm-4-flash as the DeepEval Judge LLM.
AI Evaluation Series (05): Agent Evaluation — Tool Call Accuracy and Trajectory Quality
Systematic evaluation of a customer support Agent (3 tools, 15 test cases). Results: tool name accuracy 73%, parameter accuracy 100%, step efficiency 0.73x. All 4 failures are 'should have called a tool but didn't' — not 'called the wrong tool.' The bottleneck is tool triggering logic, not tool execution. Trajectory quality scored 3.73/5 by LLM-as-Judge, with multi-step and edge cases performing slightly worse.