OpenAI released the GPT-5.6 family in three flavors — Sol, its most powerful model; Luna, built for speed; and Terra, tuned to balance the two for everyday work. A new "ultra" mode inside Sol lets the system delegate work to submodels for the hardest tasks.
Alongside the models, OpenAI launched ChatGPT Work, an agent built on Codex that pulls context across a team's connected apps and files to produce finished spreadsheets, decks, and documents, rolling out first to Mac and Windows desktop apps.
CEO Sam Altman told CNBC that Sol is 54% more token-efficient on agentic coding tasks, framing it as a direct answer to enterprise cost scrutiny. Early reactions from builders were split — some found GPT-5.6 more reliable for everyday tasks, while others said Anthropic's Fable 5 still edges it on raw creative intelligence.
SpaceXAI's Grok 4.5 is the company's first major release since going public and acquiring the AI coding editor Cursor. Elon Musk described it on X as "an Opus-class model, but faster, more token-efficient and lower cost," positioning it against Anthropic's flagship reasoning model rather than chasing benchmark supremacy outright.
The model runs at roughly 80 tokens per second and claims close to double the token efficiency of comparable leading models, priced at $2 per million input tokens and $6 per million output tokens through the SpaceXAI API — well below most frontier competitors.
It's now the default model in Grok Build, available across all Cursor plans, and live via the SpaceXAI developer console — explicitly targeting the coding-agent market Anthropic, OpenAI, and Cursor-style tools have dominated.
Ollama has raised a $65 million Series B led by Theory Ventures, with Benchmark, 8VC, Y Combinator, and others joining, bringing its total funding to $88 million. The 14-person company says it's now used by 8.9 million developers monthly and sits inside 85% of the Fortune 500.
Founders Jeff Morgan and Michael Chiang previously built Kitematic, which Docker acquired in 2015 — work that became Docker Desktop. Ollama applies the same "hide the messy setup" playbook to running open-weight models locally or in its own cloud, billing by GPU time rather than per token.
Benchmark's Peter Fenton called open-weight models a shift, not a war: firms with heavy inference bills now have a real reason to run open models locally and lean on closed frontier labs like Anthropic only when needed.