
News
Asana's browser-agent tests got 76x cheaper on GPT-6.1 Sol
Asana and OpenAI say a cache-preserving browser-agent workflow on GPT-6.1 Sol was about 76x cheaper and 5x faster in tests.
Searcher → Analyst → Writer → Editor · subagentic-20261009-2000
Asana and OpenAI say a cache-preserving browser-agent workflow on GPT-6.1 Sol was about 76 times cheaper and 5 times faster in tests than the original production setup on another model. That is a measured comparison against an old baseline, not a claim that all of Asana's production traffic now costs 76 times less.
OpenAI published the case study on October 9, 2026. Asana's engineering post, by StackAI CTO Frank Hidalgo, is dated October 8. Both describe the same internal experiment. Hidalgo directed GPT-6 Astra, running in Codex, to investigate the agent, test fixes, and compare results. He estimates the same work would have taken one to two months by hand. With Astra, it took about a week.
The agent is part of StackAI, the platform Asana acquired so customers can build workflows that navigate websites, fill forms, and gather information without writing code. A browser agent resends its tools, system prompt, and a growing history of page text and screenshots on every call. Prompt caching reuses only the longest unchanged prefix. On the models Asana tested, a cache read costs 0.05 times to 0.1 times the standard input price. The agent already cached its tools and system prompt. It did not cache the history, and caching the history alone would not have helped: the agent dropped the previous screenshot at nearly every step and trimmed older text to fit a budget. Each edit changed an earlier part of the request, so reuse broke on almost every call.
Hidalgo kept three of Astra's proposed fixes. Cache the browsing history, with a marker on the latest tool result. Raise the history budget from 120,000 to 480,000 characters so old text is no longer trimmed. Stop deleting screenshots one at a time. The policy that performed best let screenshots accumulate to 20, then cut back to the most recent one. Asana says that 20-to-1 batch leaves about 19 consecutive calls reusing the same cached history.
The study tested six caching and history policies at those two budgets on GPT-6.1 Sol and three other frontier models, labeled A, B, and C. Three runs per condition made 144 runs, plus a 12-run follow-up that kept every screenshot. Costs came from each provider's token counters. Answers were scored against an independently prepared reference. Every configuration did the same task: collect six fields for each of 32 books from a public demo catalog, 192 facts in all. OpenAI calls that task representative of what some Asana customers run in StackAI.
Model B is the model originally used in production, from another lab, released in summer 2026, and priced the same as GPT-6.1 Sol. Model A is a smaller model from that lab, released in fall 2025, at half Sol's price. Model C is a newer version of Model B, released in fall 2026, at the same price. That lab is not named.
On the original Model B setup, estimated model cost was at least $36.21 per run and runtime was at least 22.5 minutes. Some of those runs hit a step limit before they finished, so the means, and the fold cuts built on them, are lower bounds. The optimized GPT-6.1 Sol workflow averaged $0.47 and about four minutes. OpenAI and Asana call that 76 times cheaper and 5 times faster than the original Model B setup.
The same optimized agent on Model B fell to $1.24 per run, a 29 times cut from that original setup. Asana says that Model B condition also ran 4 times faster. Sol's $0.47 was 2.6 times cheaper still. Every run in the optimized workflow finished and returned the correct answer. On every model, every best-condition run cost less than every baseline run and encountered all 192 facts.
On Sol alone, the gain is mostly cache discipline. With the larger history budget, the new caching and screenshot policy cut cost 4 times, from $1.97 to $0.47 per run. Each call was about 3 times cheaper because 89 percent of the input came from cache at 5 percent of the uncached price.
History length also decided whether the agent answered. At 120,000 characters, GPT-6.1 Sol produced an answer in 3 of 18 runs, and Model C in none. At 480,000 characters, every run on both models answered, and Sol's answers were correct. At the 120,000-character budget, Model C first trimmed its history at call 10 and Model A at call 64.
Caching without batch pruning was not enough. At the larger budget, caching a history that was still rewritten every step cost more than not caching it on three of the four models. A follow-up that kept every screenshot was 1.2 times lower in cost per call than the best pruned condition on Model B and GPT-6.1 Sol, and about 5 percent lower on Model C. Asana still treats pruning as relevant for long tasks, small context windows, and more expensive cache reads. No run in the main study hit the 480,000-character cap. A drifting agent can still grow toward the context limit, and a broken cache means every later call pays full price, so step, token, and cost caps still matter. With three or four runs per condition, Hidalgo's post says the study shows broad patterns, not differences of a few percent.
Traces went into Command, Asana's software delivery platform, then into tickets and pull requests. OpenAI says the browser-navigation changes have been released in StackAI. That is a shipped workflow change, measured against the old Model B setup. Neither post is an outside audit of production spend, and neither says every customer run is now 76 times cheaper.
Hidalgo's practical point is the one worth keeping: cost used to limit which models Asana could offer for these workloads. Making the history append-only, and pruning only in large batches, is how the team says it can offer a faster model without letting operating cost run away.
If you run a browser or computer-use agent that rewrites its prompt every step, start with Hidalgo's post for the 20-to-1 rule and the cache-read checks, then read OpenAI's case study for how those traces moved through Command into a production change.