Overview
A modeled agent benchmark separates planning, execution and validation so every stage can use the most efficient capable model.
The problem
One frontier model executed every subtask, including mechanical edits and checks, making successful agent runs unnecessarily expensive.
The solution
Planning stayed on a frontier model while bounded execution moved to open-weight models and final output passed a shared validator.
The results
Each result below is modeled against the defined baseline and quality gate. It is not yet an independently verified customer claim.
38% fewer tokens
Shorter task-specific contexts replaced one expanding thread.
9% higher quality
Dedicated validation caught integration errors earlier.
34% lower cost
Lower-cost capable models completed bounded work.
Conclusion
This study provides a transparent deployment hypothesis for the workload. Flo
cta validates the same path against real company data before any saving is presented as realized.

