Five years running AI: what creates value and what holds it back
Applied AI Insights 2026 gathers what we've learned running AI agents across conversations, data, voice and real business processes. Operating data, not intention surveys.
- interactions and messages
- 4.8M+
- conversations evaluated
- 10,247
- evaluation runs
- 12,000+
- models and configurations
- 35+
Five lessons that explain most of the value, and the cost
- 01
Execute, don't just answer
2.7×
operational savings from AI that executes versus AI that only answers (32% vs 12%)
Value appears when the agent changes the state of the business.
- 02
Context before instructions
+29 pp
of correct resolution contributed by context, state and tools
Instructions contributed 9 pp.
- 03
Advanced reasoning, only when it decides
14–19 pp
improvement in complex decisions, versus ≈2 pp in simple tasks
Thinking more pays off when it changes a decision.
- 04
Several models, not one
−46%
average cost per task when each task goes to the right model
Quality dropped by 2 pp.
- 05
Cost lives outside the model
69%
of total operating cost was not inference
Integration, data, voice and control weigh more.
What's inside
-
Business impact
A pilot goes live in 14 days; the return arrives between the third and ninth month.
-
Autonomy and reliability
80% of critical failures came from data, tools or rules.
-
Data and context
The more context a model receives, the worse it uses it: 3 to 5 chunks is the sweet spot.
-
Architecture and control
The orchestration layer and Loop Engineering raised resolution from 72% to 91%.
-
Models and economics
On simple tasks, cost varied up to 22× while quality stayed relatively close.
-
Voice and channels
In voice, response speed shaped the experience as much as the content.
Operating data, not intention surveys.
View the report