Back to Flocta

What is semantic caching?

Semantic caching is the reuse of a language model's previous answer when a new request means the same thing as an earlier one, rather than only when the two are character-for-character identical. A conventional cache misses as soon as a word changes; a semantic cache compares meaning, so a question asked a second time in different words can still be answered without paying a model to produce the answer again.

How it works

  1. 01The incoming request is turned into a representation of its meaning, together with the inputs, the context it depends on and the version of the underlying data.
  2. 02That fingerprint is compared against the fingerprints of work the system has already done and verified.
  3. 03On a match the stored answer is returned. On any meaningful change , different inputs, changed context, newer data , the request goes to a model as normal.

What semantic caching is not

Semantic caching is not the same as a prompt cache, and it is not the provider-side caching some model APIs bill at a lower rate. Provider caching reuses parts of the prompt inside one model's own infrastructure and expires in minutes. Semantic caching sits above the provider, lasts as long as the underlying facts do, and can return an answer without calling any model at all. It is also not an answer that gets staler the longer it is kept: a cache entry that cannot prove its inputs are still current has to be treated as a miss, otherwise the saving is paid for with a wrong answer.

How Flocta uses it

Flocta fingerprints each task and reuses a verified result when the task, its inputs, the relevant context and the data version all still match. Any meaningful change triggers a fresh run. It is one of the three mechanisms behind the platform's up to 35 percent lower token cost at the same output quality.

Questions people ask

How is semantic caching different from normal caching?

A normal cache keys on the exact request, so a single changed word is a miss. A semantic cache keys on what the request means, so the same question phrased differently can still hit. The hard part is not the matching but the invalidation: an entry whose inputs or source data have moved on must be treated as a miss.

Does semantic caching make answers worse?

Only if it is allowed to return an answer whose basis has changed. Done correctly it returns an answer the system has already produced and verified for the same task on the same data, which is the same answer the model would produce again , at no cost.

Flocta runs all three of these underneath the assistant your team already uses.

Request a demo

More terms