Anthropic Ships Claude Opus 4.8 in Just 41 Days, Pushing AI Agents Toward Real Work
Anthropic is picking up the pace. On May 28, just 41 days after releasing Claude Opus 4.7, the company dropped Opus 4.8 at the same price. The update targets a persistent complaint from developers: earlier versions felt incremental. This time, the focus is on judgment and honesty, not just benchmark scores.
Early testers report sharper decision-making. One noted the model “asks the right questions, catches its own mistakes, pushes back when a plan isn’t sound.” On the Super-Agent benchmark, Opus 4.8 completed every case end-to-end, matching GPT-5.5 at similar cost. It also scores 69.2% on agentic coding measures, up from 64.3% for 4.7, and hits 83.4% on agentic compute use—topping GPT-5.5 and Gemini 3.1 Pro in several categories.
Honesty is the standout feature. Anthropic trained the model to flag uncertainties roughly four times more often than 4.7. Testers at Bridgewater Associates saw it proactively identify issues with inputs and outputs, catching flaws other models left for users to find. That matters in high-stakes fields like law and finance, where false confidence can derail projects.
The release also brings new tools. Claude Code now supports dynamic workflows that plan large tasks, deploy hundreds of parallel subagents, and verify outputs before returning results. Enterprise teams can handle codebase migrations spanning hundreds of thousands of lines from start to finish. Users on claude.ai get an effort control slider to adjust how deeply the model thinks on any query, balancing thoroughness against token costs.
Fast mode runs at 2.5 times standard speed and costs three times less than before. Regular pricing holds at $5 per million input tokens and $25 per million output tokens. The model is available through the Claude API, Amazon Bedrock, and Google Vertex AI.
Alignment assessments show progress too. Opus 4.8 scores higher on prosocial traits like supporting user autonomy, with deception rates below those of 4.7 and close to Anthropic’s best-aligned model, Mythos Preview. The company continues to hold back its more powerful Mythos-class models, which require additional safeguards before wider release.
Customer feedback reflects practical gains. Databricks reports that Opus 4.8 unlocks deeper multistep reasoning in its Genie agent while cutting token costs on multimodal documents by 61% versus 4.7. Legal tech teams at CoCounsel cite better consistency and citation precision. One tester called it a “major quality-of-life update” that carries context and style more effectively across extended sessions.
Anthropic struck a measured tone, describing the upgrade as “modest but tangible.” Executives signaled future models will deliver similar capabilities at lower prices. For now, the combination of refined judgment, parallel agent orchestration, and user-controlled effort gives teams new levers for deploying AI at larger scales.
Source: Webpronews
dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.