How Asana Kept Browser-Agent Cache Reuse Intact
A 144-run study found that append-only history, larger budgets and batched screenshot pruning preserved prompt-cache prefixes far better than editing context at every step.
Edited by Tyronne Panaino
Asana published an engineering study on October 8 describing changes that it says made a StackAI browser agent cheaper and faster without reducing its answer quality on the tested task. The central lesson is not a new model feature. It is a context-management change: keep the growing history stable long enough for the model provider's prompt cache to reuse it. Asana says the resulting changes shipped to StackAI, making the work relevant to teams operating long browser sessions rather than only to benchmark readers.
Why the original cache kept breaking
The Asana engineering account explains that prompt caching can reuse only the longest unchanged prefix of a request. Its browser agent already cached fixed tools and system instructions, but the session history kept changing. The agent removed the previous screenshot at every step and trimmed older text as it approached a history budget. Those edits changed earlier positions in the request, so later calls repeatedly lost the reusable prefix.
This is a subtle failure mode. A cache marker does not help much when the application rewrites the material before it. Browser agents are especially exposed because every action can add page text, screenshots and tool results. An implementation can appear to support caching while still paying close to the uncached input price on most calls.
What Asana changed
The team tested a more append-only history. It placed the cache marker on the latest tool result, expanded the text-history budget from 120,000 to 480,000 characters and stopped deleting a screenshot after every step. Under the selected 20-to-1 pruning policy, the agent could accumulate as many as 20 screenshots before cutting the set back to the newest one. That created a longer run of calls with the same prefix.
Asana used GPT-6 Astra through Codex to inspect the code, instrument requests, run the experiments and analyze traces, while people set the goal and reviewed the conclusions. Its study covered six caching and history policies at two budgets across GPT-6.1 Sol and three unnamed frontier models. The company reports 144 main runs plus a 12-run follow-up, with every answer scored against a separately prepared reference.
The accompanying OpenAI case study describes a common task across configurations: collect six fields for each of 32 books from a public demonstration catalog. Asana says the optimized GPT-6.1 Sol workflow cost 76 times less and ran five times faster than the original production setup on an unnamed comparison model. It also reports that 89% of the optimized input was served from cache.
What those numbers do and do not establish
These are first-party results from Asana and OpenAI, not an independent benchmark. The model names behind three comparison labels were withheld. Each main condition had only three runs, and Asana says varying call counts mean the study supports broad patterns rather than small differences between conditions. OpenAI also notes that some chart markers and variability measures were reconstructed from source images because the underlying run values and standard deviations were unavailable.
The large cost ratio therefore should not be generalized to every browser agent. It combines a weak original workflow, a changed history policy and a different model. The more portable finding is mechanical: editing old context can invalidate prefix caching, and batching unavoidable changes can preserve reuse for more calls.
Asana's follow-up also found that retaining every screenshot could outperform the selected pruning setup on this bounded task. That does not remove the need for limits. The company warns that long or drifting agents can still approach a context window, and a broken cache can suddenly make every call pay full input cost. Step, token and cost caps remain part of the operating design.
A practical evaluation checklist
Teams testing the same pattern should log cache-read tokens for every call, not only the total bill. They should compare stable and frequently rewritten histories on the same task, keep the answer-quality reference separate from the agent, and retain complete traces so failures can be audited. They should also test what happens when the session exceeds its normal history size.
That makes the study most useful as an experiment design, not as a universal performance promise. It shows how context structure, pruning policy and model behavior interact in one shipped browser-agent workflow, while leaving replication across sites, task lengths and providers for later work.
Status
Learning. The implementation, study design and deployment status are documented by Asana and OpenAI. Confidence is medium because the sources are participants in the same case study, several comparison models are anonymous and the run-level data were not independently reproduced in the fetched evidence.
Sources
- Asana — How we cut a browser agent's cost by keeping its cache intact
- OpenAI — Asana browser-agent case study
Update note: Last reviewed October 11, 2026. We will revise this post if Asana publishes reproducible run-level data or independent replications.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.