Docs · Traces & cache
Traces and the answer cache
Updated September 25, 2026 · by Nishanthan Janarthanarajah (Nishy), founder of Agent Studio
Reading a trace
- Status: ok, blocked, or error, plus “cache” when no model was called.
- Timeline: shield → input guardrails → orchestrator steps (indented: tool calls and delegations) → output guardrails → cache decision. Bars show where the time went.
- What happened: findings such as “No tools or sub-agents were used” when the agent had tools but answered from memory, with the fix to try.
- Payloads: click any row for the full JSON, including tool inputs and results (truncated for size).
{ "stage": "step", "index": 2, "toolCalls": [{ "name": "delegate_to_order_tracker", "args": "{"task":"Look up A1234"}" }], "finishReason": "tool-calls", "inputTokens": 812, "outputTokens": 41, "ms": 640 }Answer cache
| Setting | Meaning |
|---|---|
| Enable | Turn the cache on for this agent's deployed API. Off by default. |
| Match rephrased questions | Also reuse answers when a new question shares most meaningful words with a cached one. |
| Cache answers that used tools | Off by default because tool answers are live data (orders, calendars). Turn on for stable lookups. |
| Keep answers for | Time to live, from 1 hour to 30 days. Deploying or changing settings clears everything. |
Cache hits appear in traces and in the API response as cached: true, with zero tokens. The playground never uses the cache so you always see current behaviour. See the API reference.
Frequently asked questions
What does a trace contain?+
One event per stage: the injection shield verdict with score and signals, each guardrail decision with its reason, every orchestrator step with the tools it called and tokens used, each sub-agent delegation with task and result, every MCP or webhook tool call with input, output and timing, the cache decision, and any error with the stage it happened in.
How long are traces kept?+
30 days. Older runs are pruned automatically.
How does the answer cache decide when to reuse an answer?+
Only single-turn questions to the deployed API are considered. The question is normalised and hashed; an exact match returns the cached answer at once. With fuzzy matching on, a rephrased question that shares most of its meaningful words is also matched. Blocked answers are never cached, and answers that used tools are only cached if you opt in.
Does the cache make answers stale?+
It can, which is why it is off by default, has a TTL you choose, and is cleared automatically on every deploy or settings change. Turn it on for agents that answer stable questions such as FAQs and product information.
What happens to long conversations?+
When the history you send is above roughly 1,000 tokens, the older turns are summarised by a fast model and only the last four turns go to the orchestrator verbatim. The trace shows a compact event with tokens before and after. Facts, names, numbers and decisions are kept in the summary; if a detail from early in a conversation was lost, the summary is where to look.
Is this the same as prompt caching?+
No. Prompt caching reuses a model's internal state for a shared prefix; Groq does not expose it. The answer cache reuses the final answer, which saves all tokens for a repeated question.