---
type: "Concept"
title: "Evaluation moved from format to execution"
description: "Benchmarks shifted from argument syntax to state, policy and outcome reliability."
tags: [theme, evaluation, benchmark]
generated: { by: human:crpage, at: 2026-07-09T09:44:00Z }
status: stable
---

The most important trend in the research: evaluation moved from **format correctness** (can the model emit a valid call?) to **execution correctness** (does the multi-step task complete reliably under state, policy and dynamic users?). [Gorilla/APIBench](../../research/gorilla.md) tested syntax; [AgentBench](../../research/agentbench.md), [ToolSandbox](../../research/toolsandbox.md) and [τ-bench](../../research/tau-bench.md) test state, policy and outcomes.
