---
type: "Research"
title: "ToolSandbox (2024)"
description: "Stateful, conversational evaluation with a user simulator."
resource: "https://arxiv.org/abs/2408.04682"
tags: [benchmark, agents, state, research]
generated: { by: human:crpage, at: 2026-07-09T09:44:00Z }
status: stable
sources: [{ id: primary, resource: "https://arxiv.org/abs/2408.04682" }]
---

A **stateful** conversational benchmark with a user simulator and milestone-based trajectories — showing that state dependence and insufficient-information handling remain hard. Simulated tools only.

Relates to: [Execution and orchestration](../stack/execution-and-orchestration.md).
