Research & benchmarks
The tool-use and federated-learning research lineage.
OKF Authoring Research
- Open data for the Explore OKF pilot
selects a bounded Coventry everyday-services journey and authoritative UK
sources for testing readable graph labels, useful semantic linking,
CPSV-AP/SKOS mappings and multi-resolution geography before reviewing
okf-uk-living.
AI And Tool-Use Research
- FedAvg — McMahan et al. (2017) — Communication-efficient decentralised training by model averaging.
- Google keyboard FL — Hard et al. (2018) — On-device federated training for next-word prediction at production scale.
- Kairouz et al. — Advances and open problems (2019) — The field's open problems: heterogeneity, privacy, fairness, evaluation.
- REALM (2020) — Retrieval-augmented pre-training; knowledge can be retrieved, not only stored.
- MELLODDY (2019–2022) — Ten pharma companies trained a shared model without sharing raw data.
- MRKL (2022) — Neuro-symbolic modular routing to expert/knowledge modules.
- ReAct (2022) — Interleaving reasoning traces with actions improves task success.
- Toolformer (2023) — Self-supervised learning of when and how to call tools.
- API-Bank (2023) — Runnable evaluation exposing planning, retrieval and calling gaps.
- Gorilla & APIBench (2023) — Fine-tuned LLM + document retriever; adapts to doc change, cuts hallucination.
- ToolLLM / ToolBench (2023) — 16,000+ APIs; ToolLLaMA with a neural API retriever; DFSDT.
- AgentBench (2023) — Multi-environment agent evaluation across 8 interactive environments.
- Berkeley Function Calling Leaderboard (BFCL) — Function-call accuracy: relevance, parallel and sequential calls.
- APIGen (2024) — Pipeline for verifiable function-calling data; 3,673 executable APIs.
- τ-bench (2024) — Tool-Agent-User evaluation with domain policies.
- ToolSandbox (2024) — Stateful, conversational evaluation with a user simulator.
- ToolACE (2024) — Self-evolution synthesis; 26,507 APIs; dual verification.