Evaluation moved from format to execution

The most important trend in the research: evaluation moved from format correctness (can the model emit a valid call?) to execution correctness (does the multi-step task complete reliably under state, policy and dynamic users?). Gorilla/APIBench tested syntax; AgentBench, ToolSandbox and τ-bench test state, policy and outcomes.