Zhaokun's Blog
HOME
RESEARCH
PROJECTS
TALKS
ARCHIVES
ABOUT
English
简体中文
Deutsch
HOME
RESEARCH
PROJECTS
TALKS
ARCHIVES
ABOUT
45
Tags
4
Categories
17
Posts
Evaluation
2025
3
Your AI Agent Passed the Demo. Now Try to Break It.
Why LLM Evaluation Gets Weird the Moment You Add Tools
RAG Works Great—Until Your Documents Disagree
2024
2
What Actually Breaks When You Quantize an LLM?
Did the Model Reason—or Has It Seen the Answer Before?
1
DE
文