Benchmarks
AI Agent Task Success & Reliability
These figures are evidence about evaluation, not a universal production target. The tau-bench result shows that realistic multi-turn, tool-using service tasks remain materially harder than answer-generation demos, especially when a task must succeed repeatedly. Nubank's deployment result shows the upside of evaluation-driven iteration in one use case, but its percentage-point lift is a before-and-after result, not a cross-company benchmark.
What is AI Agent Task Success? Read the definition| Figure | Segment | Source | Year |
|---|---|---|---|
| +29 percentage points | Nubank card-delivery deployment: self-service rate versus prior AI-agent variant | Nubank, Building Customer Support AI Agents at 100M-User Scale | 2026 |
| +37 percentage points | Nubank card-delivery deployment: AI transactional NPS versus prior AI-agent variant | Nubank, Building Customer Support AI Agents at 100M-User Scale | 2026 |
| < 50% | State-of-the-art function-calling model task success on realistic retail and airline service tasks | ICLR 2025, tau-bench | 2025 |
| < 25% pass^8 | Retail tasks passed consistently across eight trials by the evaluated leading model | ICLR 2025, tau-bench | 2025 |
Benchmarks describe other teams, not yours. Read them as a sanity check on direction, not a target to hit — the right number depends on your channel mix, customers, and how you define the metric.
Continue exploring