What is context compression?
Context compression is the practice of sending a language model only the facts, constraints, instructions and evidence that the current task actually needs, instead of the whole conversation, document set or company knowledge base. Because models are billed by token, the context is usually the larger half of the bill, and most of it is material the task never uses.
How it works
- 01The task is read first, and what it requires is determined before anything is assembled.
- 02The available context is filtered against that requirement: relevant passages, the constraints that bind the answer, the evidence it has to cite.
- 03Everything that does not change the answer , unrelated history, duplicates, material already superseded , is left out.
What context compression is not
Compression here does not mean summarising the context into fewer words, and it is not a smaller context window. Summarising loses the exact wording a task may need to quote, and it introduces a second model call to save the first one. Selecting is not summarising: the passages that go in are unchanged, there are simply fewer of them. It is also not the same as retrieval-augmented generation, although the two are often confused , retrieval decides what to fetch, compression decides what of it survives into the prompt.
How Flocta uses it
Flocta filters company context against the current task and sends only the required facts, constraints, instructions and evidence. The model receives less noise while the meaning needed for a high-quality answer stays intact.
Questions people ask
Why does context cost so much?
Because input tokens are billed too, and in a workload with company knowledge attached the input is usually larger than the output. Sending a whole knowledge base with every request pays for the same material over and over.
Is context compression the same as a summary?
No. A summary rewrites the material and loses the exact wording; compression selects from it and leaves what it keeps untouched. Rewriting also costs a model call, which works against the saving it is supposed to produce.
Flocta runs all three of these underneath the assistant your team already uses.
Request a demo