Overview
A controlled workload study for a high-volume software platform measured how mixed-model routing changes token use, quality and total inference cost.
The problem
Every request was sent to the same frontier model, including repetitive support tasks that did not need its full reasoning capacity. Token spend grew with volume while quality was not measured consistently.
The solution
Flocta classified each task, compressed reusable context and routed bounded work to efficient open-weight models. Frontier models remained available for ambiguous cases, with every output checked against one shared quality gate.
The results
Each result below is modeled against the defined baseline and quality gate. It is not yet an independently verified customer claim.
35% fewer tokens per accepted response
Prompt compression and reusable context removed repeated input without changing the user-facing task.
8% higher measured quality
Task-specific routing outperformed the one-model baseline on the defined evaluation set.
35% lower modeled inference cost
The mixed-model path reduced cost while preserving escalation to a frontier model when required.
Conclusion
This study provides a transparent deployment hypothesis for the workload. Flo
cta validates the same path against real company data before any saving is presented as realized.

