Blog · · 11 min read
AI agent caching without stale answers

AI agent caching can reduce model latency and cost, but a careless cache can bypass the product's current state. It may return an answer created for another tenant, under an older permission set, with a previous agent release, or from a record that changed seconds ago. Semantic matching adds another failure mode: a nearby result can sound plausible while answering a different question. Cache policy therefore belongs alongside authorization and data-freshness controls. The implementation below separates four cache layers, uses authority-aware keys, excludes side effects, and validates hits before users see them.
Separate the four cache layers
There is no single cache layer in an agent system. Each mechanism below reuses a different artifact and needs its own key, lifetime, and acceptance rules.
| Layer | Reused value | Match rule | Main risk |
|---|---|---|---|
| Provider prompt cache | An identical prompt prefix | Exact prefix match | Assuming it reuses a prior answer |
| Exact-response cache | A completed model response | Exact normalized request or custom key | Omitting identity or release inputs from the key |
| Semantic-response cache | A completed response | Vector similarity plus policy checks | Reusing a plausible but wrong answer |
| Tool-result cache | A deterministic read result | Explicit function arguments and source versions | Replaying stale reads or pretending a write ran |
Provider prompt caching does not normally return an old answer. It reuses work performed on an identical input prefix. OpenAI requires exact prefix matches and separates caches by organization. Anthropic also uses exact prompt-prefix matching, and changes to cached blocks or tool definitions invalidate the prefix. Stable instructions, tool definitions, and repeated reference material fit this layer.
An exact-response cache returns a completed response. Cloudflare AI Gateway documents whole-request keys, custom keys, bypass controls, and TTLs. This is where application key design matters. A custom key based only on the user's text can collapse requests that differ by tenant, authorization, model, locale, or product state.
A semantic cache accepts similar wording rather than requiring an exact request. Redis describes this as vector-similarity retrieval and recommends tuning thresholds, TTLs, and invalidation for the volatility of the data. The broader match creates more hits, but it is probabilistic. Similar wording does not establish equivalent intent, authority, or source state.
Tool-result caching needs separate rules. A read such as get_plan_limits(tenant_id, plan_version) may be cacheable when its inputs and source version are explicit. A write such as close_ticket(ticket_id) is never a cached execution. Returning an earlier success payload would claim that a side effect happened in the current run even though no tool ran.
Classify each operation before caching it
Put cache policy in the tool or operation registry. Individual prompts must not decide whether a result can be reused.
Use four classes:
stable_prefix: model instructions, schemas, and reference blocks that can use provider prompt caching.pure_versioned_read: deterministic reads whose result is tied to explicit source versions.bounded_response: generated answers that may be reused only inside a narrow authority and freshness scope.non_cacheable: side effects, security decisions, approvals, volatile values, and any operation without enough provenance to validate a hit.
Set the default to non_cacheable. Before enabling reuse, the tool owner must identify the source of truth, key inputs, freshness rule, invalidation events, sensitivity class, and hit validator.
- Cache static policy text by content hash, not by a mutable document name.
- Cache a product read by tenant, resource ID, selected fields, and record version.
- Bypass cached responses for authorization questions because current authority is the answer.
- Bypass for inventory, balances, quotas, incidents, and other fast-changing facts unless the product can prove a bounded freshness requirement.
- Never cache an approval decision for a different action, amount, target, user, or policy version.
- Never represent a previous tool receipt as evidence that a current write executed.
Google documents implicit and explicit context caching, while Amazon Bedrock reports cache reads and writes separately. Both optimize model input. Neither provider feature decides whether the application may reuse a response across identities or product states.
Build an authority-aware cache key
A safe key includes every input that can change whether a result is correct or visible. Request text is only one such input.
{
"cache_schema": "agent-cache-v3",
"layer": "semantic_response",
"tenant_id": "tenant_42",
"authority_fingerprint": "role:support-agent|scope:ticket:read",
"user_segment": "internal-support",
"locale": "en-GB",
"agent_release": "support-agent-2026-08-12.2",
"model": "provider/model-version",
"tool_contract_set": "sha256:tool-set-version",
"policy_epoch": "support-policy-19",
"source_versions": {
"ticket:T-1842": "v44",
"account:A-91": "v12"
},
"normalized_intent": "summarize-ticket-for-handoff",
"sensitivity": "internal"
}
Hash the canonical serialization for storage, then keep the manifest with the entry. Operators need the manifest to explain a hit; the hash alone provides only an index.
Include current authority
Users in the same tenant can have different access to a resource. One user can also lose a role while a long-running agent waits for approval. Build the authority fingerprint from the current policy decision, including the relevant role, resource scope, entitlements, and policy version. Ignore role names supplied in model input.
A permission change should advance a policy epoch or alter the authority fingerprint so old entries become unreachable. Keep sensitive generated responses within a user-specific boundary unless the authorization system can prove that everyone in a broader segment may see the same result.
Bind entries to the agent release
Prompt text is only one behavioral input. Tool schemas and descriptions, model versions, safety policy, retrieval configuration, and output parsing can all change the answer. Include the immutable agent release that bundles those components. A changed tool contract should make responses from the old release ineligible.
Record source versions
A TTL measures an entry's age; it cannot reveal whether the underlying record changed. Include database versions, ETags, content hashes, or another change token for every mutable source that affects the answer.
Store a source-version map when an answer combines several sources. The hit remains eligible only while every recorded version still matches. If a source has no usable version marker, assign a short TTL and document the weaker guarantee. An unexpired clock does not prove freshness.
Invalidate when product state changes
Use TTL as a backstop. Product events should invalidate entries when their dependencies change.
For a support summary built from a ticket and account record, subscribe to events such as:
- ticket message added;
- ticket status changed;
- account plan changed;
- user role or team membership changed;
- policy published;
- agent release activated or rolled back;
- source record deleted or access revoked.
Invalidate by dependency instead of scanning prompt text. Maintain a reverse index from each source identity or policy epoch to its cache-entry IDs. When exact fan-out is too expensive for a high-volume source, increment a namespace version. Existing entries can expire later, but new reads can no longer reach them.
Reserve stale-while-revalidate for low-risk display paths with a stated maximum stale age, and tell the UI that the content is stale. Do not use it for actions, compliance decisions, current permissions, or facts that could change the agent's next action.
When a source is deleted or removed after a privacy request, purge both the cached payload and its derived embedding. Keeping the embedding after deleting the text leaves a derived representation outside the intended lifecycle.
Validate semantic hits before reuse
A semantic lookup should return candidates, not an accepted answer. Its vector score measures proximity in embedding space. It cannot establish that the old response satisfies the new request.
One LangChain issue describes a request for boredom advice receiving the joke cached for another question. A separate LangChain issue shows per-query hashes missing paraphrases, while a static key returns the first answer for unrelated questions. These author reports concern specific configurations, but they expose the tradeoff between an overly specific key and an overly broad one.
Filter candidates in this order:
retrieve semantic candidates inside exact authority namespace
-> reject mismatched locale, release, model, policy, or sensitivity
-> compare every mutable source version
-> enforce TTL and maximum stale age
-> compare structured intent and required entities
-> run a task-specific answer suitability check
-> emit a hit with provenance, or regenerate
"Summarize ticket T-1842" and "close ticket T-1842" share an entity but request different operations. Keep the operation, resource type, resource ID, requested fields, and risk class separate from free text.
A second model may judge fuzzy equivalence for low-risk content, but it cannot replace authority, source-version, or policy checks. Record the judge model and prompt version, measure false-hit rates, and return a miss when the judge is unavailable. A GPTCache proposal adds a second validation gate after vector similarity because related wording can still produce a factually inappropriate hit.
Work through a tenant-safe support example
Consider a support agent that summarizes ticket T-1842 for an internal handoff. The first request runs under tenant 42, user U-7, policy epoch 19, ticket version 44, and account version 12. The runtime stores those inputs with the generated summary.
A second support user asks for the same summary. The lookup proceeds as follows:
- Resolve the second user's current tenant and resource authority.
- Enter only the tenant and authority namespace that the user may access.
- Match the normalized intent and ticket ID.
- Read current ticket and account version tokens.
- Compare policy, agent release, model, tools, locale, and sensitivity.
- Return the cached summary only when every hard boundary matches.
- Record which entry was used and which checks passed.
If someone adds a private note and advances the ticket to version 45, the source-version check forces a miss. A perfect semantic score and 20 minutes of remaining TTL do not matter. The agent regenerates from version 45.
If user U-7 loses access while the entry remains in Redis, the new authority fingerprint cannot address it. The permission-change event also advances the policy epoch, which makes old keys unreachable for other affected users. Authorization still comes from current product policy, not from cache contents.
Handle failures without weakening the boundary
Define how each operation behaves when the cache is unavailable or cannot validate an entry.
- If the cache is unavailable, run the uncached path when the underlying operation is safe and within budget.
- If authority resolution fails, do not query a broader namespace. Fail closed.
- If a source version cannot be checked, treat the candidate as a miss unless the operation has an approved bounded-staleness rule.
- If invalidation delivery is delayed, namespace versions and short backstop TTLs should limit exposure.
- If a semantic validator fails, regenerate instead of accepting the nearest candidate.
- If a cached payload fails schema validation, quarantine it and emit a miss.
- If provenance is incomplete, do not serve the entry to a user.
- If an entry appears under the wrong tenant or sensitivity label, disable that cache layer and investigate it as a security incident.
Limit which services can write to each namespace. Validate payload schemas, cap object size, and encrypt sensitive values according to the product's data policy. Model-generated key material must never override tenant or authority fields. User text may contribute to normalized intent, but trusted application code supplies identity, policy, release, and source versions.
Verify AI agent caching before enabling hits
Begin in shadow mode. Run lookups and record candidates while continuing to generate a fresh response for the user. Compare each candidate with that response and with current source state.
Measure these results separately for each layer and operation:
- eligible requests and cache-hit candidates;
- accepted and rejected hits by reason;
- exact and semantic false-hit rate from labeled checks;
- stale-hit attempts blocked by source-version checks;
- cross-authority candidate attempts, which should be zero after namespace filtering;
- p50 and p95 latency saved;
- model input and output units avoided;
- regeneration rate after invalidation;
- entries without complete provenance;
- incident and user-correction rate for served hits.
Exercise these negative paths:
- Change a record without waiting for TTL expiry.
- Revoke a user's role between cache write and read.
- Switch tenant while preserving identical request text.
- Deploy a new tool schema under the same human-facing agent name.
- Ask a semantically similar question with a different operation or entity.
- Delete a source and verify that payload and embedding are purged.
- Make the invalidation consumer lag.
- Attempt to cache a write tool's success result.
- Corrupt the provenance manifest.
- Disable the cache and verify the uncached path still works.
Keep semantic reuse disabled until false-hit review meets a task-specific threshold and every authority or source-version mismatch produces a miss. Any observed cross-tenant answer blocks release, regardless of the measured savings.
Start with one versioned read
Choose one high-volume, read-only operation backed by a stable version or ETag. Register it as pure_versioned_read, build the authority-aware key, connect change events to invalidation, and run shadow lookups for a week. Inspect every accepted candidate and rejected stale hit. After the negative tests pass, enable exact hits for a small tenant cohort. Add semantic reuse later, once equivalence errors can be measured separately from freshness and authorization failures. The resulting AI agent caching policy will have a defined boundary and evidence for every hit.
References
- OpenAI: Prompt caching supports exact prefix matching, tool-prefix behavior, organization separation, and cache retention details.
- Anthropic: Prompt caching supports exact cached-prefix behavior, cache breakpoints, TTL options, and invalidation when cached content or tool definitions change.
- Google Gemini API: Context caching supports the distinction between implicit and explicit model-context caching.
- Amazon Bedrock: Prompt caching supports cache-read and cache-write accounting for repeated model input.
- Cloudflare AI Gateway: Caching supports exact request caching, custom keys, bypass controls, TTLs, and hit status.
- Redis: What is semantic caching? supports vector-similarity caching, threshold tuning, volatility-aware TTLs, invalidation, and wrong-answer risk.
- LangChain issue #25161 provides an author-reported example of a changed question receiving an earlier semantically cached response.
- LangChain issue #29517 provides an author-reported example of the tradeoff between per-query hashes and an over-broad static key.
- GPTCache issue #682 provides an author proposal for a second answer-suitability gate after vector similarity.