---
type: "Research"
title: "AgentBench (2023)"
description: "Multi-environment agent evaluation across 8 interactive environments."
resource: "https://arxiv.org/abs/2308.03688"
tags: [benchmark, agents, research]
generated: { by: human:crpage, at: 2026-07-09T09:44:00Z }
status: stable
sources: [{ id: primary, resource: "https://arxiv.org/abs/2308.03688" }]
---

Evaluated agents across 8 interactive environments, revealing a large gap between open and proprietary agents — agentic competence, not API syntax.
