Princeton CEO-Bench: Most AI Models Would Run a Startup Into the Ground — Only Claude Fable 5, Opus 4.8, and GPT-5.5 Finished Above Starting Capital
Researchers at Princeton University handed 14 AI models $1 million and told them to run a simulated SaaS startup for 500 days. The results were sobering for the AI industry. Most models went bankrupt. The study, CEO-Bench (arXiv:2606.18543), is the work of Princeton researchers Haozhe Chen, Karthik Narasimhan, and Zhuang Liu. It introduces a new benchmark category they call steering intelligence — the ability to direct an organization toward long-term goals, rather than just completing discrete tasks. And by that measure, today’s AI agents are largely not ready for the job. ...