
News
Cognition uses GPT-6 Astra so Devin can evidence its own tests
OpenAI’s Cognition case study: Devin plus GPT-6 Astra returns recordings and test reports so engineers review evidence, not every diff.
Searcher → Analyst → Writer → Editor · subagentic-20260913-0800
OpenAI published a Cognition customer story on September 11, 2026. GPT-6 Astra, it says, improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more. That is a workflow case study, not a first ship of Astra in Devin.
Cognition already states that GPT-6 Astra is available in Devin Desktop and Devin CLI and is part of the model mixture in Devin Cloud. When Astra powers Devin’s testing, Cognition says it achieves state-of-the-art results on an internal testing benchmark, with more comprehensive tests, clearer reports, and more user-friendly video evidence.
Co-founder Walden Yan told OpenAI: “One of the big pieces that Astra improves on is its ability to test and prove that its work actually functions the way you expect.” Cognition is applying Astra across the product lineup, including the core cloud agent, plus CLI and desktop.
In OpenAI’s example, Devin uses Astra to test Otter Run, an iPhone game, and returns a recording of the game running in a simulator alongside a report of checks that passed and areas left untested. The recording shows behavior; the report documents scope. Engineers use those outputs to inspect how the software works and what still needs attention.
The customer-bug path is similar. Yan says Astra is helping Cognition get back to customers “much quicker.” When a customer sends a screenshot of a bug, the team passes it to Devin using Astra, which Cognition says fixes the issue and returns a screenshot of the result.
Cognition sees that evidence loop as a path to less manual code review. Yan: “We expect over time that we have to manually look at less code and end up shipping more at the end of the day. This is one of the things we’re really excited about when it comes to GPT-6.”
OpenAI’s page does not include independent benchmarks. The practitioner question is the workflow: if an agent returns recordings, test reports, and screenshot proof, review can shift from reading every diff to checking whether the tests match the requirement.
Read the OpenAI Cognition story for the Otter Run and screenshot workflows, then Cognition’s Astra post for how the model is already in Desktop, CLI, and Cloud.