i can tell who’s still using opus 5 for their linkedin cringe posts. please spare my eyeballs, opus 5.5 is here
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
frontier modelby Anthropic
Drawn from 35 eligible posts by 22 tracked authors.
Collection is incomplete. Breaks in the line are unobserved intervals, not zero mentions.
| Interval start (UTC) | Mentions | Coverage |
|---|---|---|
| 2026-09-04 00:00 UTC | 2 | Partial |
| 2026-09-05 00:00 UTC | 3 | Partial |
| 2026-09-06 00:00 UTC | 0 | Partial |
| 2026-09-07 00:00 UTC | 0 | Complete |
| 2026-09-08 00:00 UTC | 0 | Partial |
| 2026-09-09 00:00 UTC | 0 | Complete |
| 2026-09-10 00:00 UTC | 2 | Complete |
| 2026-09-11 00:00 UTC | 1 | Complete |
| 2026-09-12 00:00 UTC | 0 | Complete |
| 2026-09-13 00:00 UTC | 1 | Complete |
| 2026-09-14 00:00 UTC | 1 | Complete |
| 2026-09-15 00:00 UTC | 0 | Complete |
| 2026-09-16 00:00 UTC | 0 | Complete |
| 2026-09-17 00:00 UTC | 2 | Partial |
| 2026-09-18 00:00 UTC | 2 | Complete |
| 2026-09-19 00:00 UTC | 0 | Complete |
| 2026-09-20 00:00 UTC | 0 | Complete |
| 2026-09-21 00:00 UTC | 3 | Partial |
| 2026-09-22 00:00 UTC | 8 | Complete |
| 2026-09-23 00:00 UTC | 3 | Partial |
| 2026-09-24 00:00 UTC | 1 | Complete |
| 2026-09-25 00:00 UTC | 3 | Partial |
9 unobserved intervals omitted. We did not collect during them, so their counts are unknown, not zero.
Mixed reception.
A breakdown of posts, not a product rating.
Named in these same posts. This does not imply a comparison or recommendation.
Public posts from the accounts Tech Twitter monitors, in the selected window. Findings need three supporting authors and a published source. Announcements can cite one known-affiliated account. This is a sample of the conversation, not a survey or a measure of adoption.
i can tell who’s still using opus 5 for their linkedin cringe posts. please spare my eyeballs, opus 5.5 is here
Holy shit, Claude Opus 5.5 is now up to 40% cheaper than Opus 5. Same tasks. Same quality. Way less cost. Breakdown: Input tokens: $4/M (down 20%) Output tokens: $20/M (down 20%) Cache reads: $0.20/M (down 60%) Most of a coding session's cost comes from cache reads, so this alone can cut your bill by more than half. AI is slowly getting cheaper.
There's definitely a weird, intangible vibe that I kind of miss from Fable. Opus is great to work with. Really smart, thorough, probably even better than Fable in most ways. But I still find myself missing "something". Hard to explain exactly what "something" is here...
TIME TO WAKE-UP AND VIBE WITH MY BOY OPUS 5.5 UNTIL LUNCH (listen to music while opus 5.5 does work i told it to do 100x better than its little useless brother opus 5)
Peek into the first day for Opus 5.5 on OpenRouter: - 22% of Anthropic tokens went to Opus 5.5 (best model launch by share in the last year) - nearly double the user count as day 1 for Opus 5 - doing better today than yesterday Claude cooked, will do a first week recap soon
SITUATION EXPLAINED: GPT-6 Sol Max is two spots behind Astra at a fifth of the cost. • Fourth in Code Arena at $8 per million tokens, blended input and output • On par with Opus 5 Max while costing less than half as much • On web dev it sits behind Opus 5 Max, Fable 5.1 Max, and Astra • Third on games @schisofrenia: "The Pareto is just getting constantly transformed every week."
opus 5.5 is finally a usable opus model again feels wayyyyyy better than opus 5 we are so back
SITUATION EXPLAINED: OpenAI cut GPT-6 prices in half. Sol now beats Opus 5 at 9% of the cost per task. • Sol is now $2 and $10 per million tokens, down from $4 and $20. Luna is $0.10 and $0.50, down from $0.20 and $1.20 • On AutomationBench, Sol at xhigh beats Opus 5 at max effort at 9% of its cost per task • Coding deception drops from 10.4% on GPT-5.6 Sol to 1.3% on GPT-6 Sol, with Astra at 0.5% • But every comparison in the post is against Opus 5, not the Opus 5.5 that shipped the same afternoon • OpenAI is also pitching better writing like Anthropic: more clarity, less jargon, fewer odd turns of phrase, shorter answers @theojaffee: "It's like a breath of fresh air to read LLM outputs and they don't sound like this grating, smug nonsense slop. Instead, they just sound like normal writing, finally."
Switched our internal @KhanAcademyEng PR review bot over from Opus 5 to Opus 5.5: Cost decreased by 50% Tool calls decreased by 34% Execution time decreased by 62% No decrease in quality.
GPT-6 Sol is now available to all Perplexity users. On our Wide-And-Deep-Research (WANDR) evals, it outperforms Opus 5 at one-fifth the price. Sol will become the “Light” Effort orchestrator for Computer users, while Astra remains the orchestrator for "High" effort. Congrats to @OpenAI for consistently launching pareto-optimal models!
SITUATION EXPLAINED: Opus 5.5 beats GPT-6 Astra on most benchmarks, and costs 40% less than Opus 5. • 66.4% on Terminal-Bench 4.0 against Astra's 57.9%, and 1846 on GDPval against Astra's 1542 • $4 input and $20 output, down from $5 and $25. Cache reads drop from $0.50 to $0.20, and output is 30% faster • Tested by METR and Frontier Design before release. Strongest alignment score Anthropic has posted, with 85% fewer attempts to cross containment boundaries than Opus 5 • It also often appears to suspect it's being evaluated, which Anthropic says complicates predicting how it behaves once deployed • Anthropic says benchmark margins are now a weaker guide to real-world differences, and the Opus-Fable gap is narrower than the scores suggest @theojaffee: "My favorite part of the model was that it writes in a refreshingly non-Claude-slop way. You can tell that it's Claude, but it's not so grating."
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Opus 5.5 is the model for people who do actual work: build products, write code. - It's good from low effort to max. Higher effort goes deeper, not wider. Lower effort leaves space for you to fill in, which I love. - It uses fewer tokens than the last Opus and doesn't run in circles like Opus 5 did. - It's easier to steer than Fable in some cases. - It can run for a long time, but in-the-loop work is where it feels special. I think this is the model that brings the next wave of people into AI, the way 4.5 did.
Claude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. I’ve been using Opus 5.5 as a daily driver and love its clear communication skills and ability to write in my style. We’re defaulting to effort medium across products, which is comparable to Fable 5.1 on intelligence but faster. Your rate limits will go 25% further on Opus 5.5 compared to Opus 5. Give Opus 5.5 an ambitious task and let us know what you think!
BREAKING: Anthropic just dropped Opus 5.5—and it’s pulling some of our recent Codex converts back to Claude. We’ve been testing it at @every across coding, design, writing, and knowledge work. It sometimes beats Fable 5.1 in our testing and is up to 40% cheaper than Opus 5. It's a strong contender for new daily driver model: It’s excellent at end-to-end builds that match your taste. I’ve started reaching for Opus over Fable on big, end-to-end coding projects. And @kieranklaassen is replacing Fable 5.1 with Opus for his day-to-day product work. He calls it his new favorite model. It produces legible prose, but still trails Astra on writing tasks. Opus 5.5 scores a 68.42 on reading ease—the highest on any model we've tested. But in my writing benchmark tests it consistently buries the main point in intro paragraphs, and revisions. You'll be able to understand what this model is saying (yay!) but for day-to-day writing it's still behind. The economics are striking. Anthropic says Opus 5.5 will cost $5 per million input tokens and $20 per million output tokens. That’s the same input price and 20% less for output than Opus 5’s $5/$25. Altogether it should save roughly 40% on costs than Opus 5. It still has a “do the most” problem. @hammermt tested it on our standard knowledge work benchmarks, and it's results were great when thye came back. On at least one test it ran past the 10 minute time limit before delivering. Net Result: If you build apps and interfaces, try it. It’s become my go-to for ambitious coding projects. I’m still roughly 80/20 Codex versus Claude in day-to-day use, and I still prefer Sol and Astra for editing. But I’m spending far more of my tokens with Claude than I was a week ago. State of Play: Codex is still the better harness for me, but Anthropic is steadily gaining ground. They have a history of making their smaller models perform better than their bigger ones (Sonnet 3.7 for example) and they seem to have done the same with Opus 5.5. Full vibe check will be on @every soon!
It happened. Grok 4.7 dropped Better intelligence than Opus 5. Half the price Fully baked into my favorite AI agent harness at the moment: Grok Bot If you haven't tried using cloud cursor agents inside Grok Bot, now is by far the best time to do it Choose a project you want to work on, connect your github, ask a grok bot to do work on it It will spin up Cursor cloud agents and write code in the cloud. Lightning fast and incredibly smart I recommend using a project management tools like Linear or Notion to make a bunch of tasks first, then have cloud agents just tear through them all 1 by 1. You'll get a massive amount of work done without much oversight. Big opportunity to lock in right now and get ahead of the curve with new tech Take my steps up above and get to it
Grok 4.7 benchmark TLDR: > big coding upgrade: ranks #4 on AA’s Coding Agent Index with Grok Build, just behind Fable 5.1, Astra and Opus 5 > but just +2 points on the Intelligence Index, still behind Sol and Opus 5 > with the gains in both coding and knowledge work it’s likely turned especially for Cursor and Grok Bot > idk about Elon’s “will exceed all current models” but on real world engineering, we must see how it performs in action
I've been using Hyper3D inside Codex to create 3D assets, especially ultra-detailed models for my landing pages. It's pretty crazy how fast the market is moving now that Astra, Opus 5, and Fable are getting so good at 3D. My current workflow is procedural three.js for environments, then Blender + Hyper3D for the more detailed models.
Hacktron founder @S1r1u5_ on why frontier models are still far ahead of open models for real-world hacking: "The frontier models are insanely capable. There's a huge gap. You don't even want to use Kimi K3 when you have access to GPT 5.6." "We can't really use Opus 5 anymore. Every time there's a new model release, you kind of don't want to work with the previous model." "Open models, we don't really use them. It's incredibly hard to hand hold them to do the exploitation. They're not yet there." @HacktronAI @rootxharsh
SITUATION EXPLAINED: Three guys hacked OpenAI in 72 hours using Claude. • Three researchers at Hacktron AI chained two vulnerabilities to take over multiple OpenAI employees' ChatGPT and Codex accounts, reaching the private GitHub monorepo in 72 hours • The entry point was an image decoder. OpenAI's forum runs on Discourse, which uses FastImage for image checks, but FastImage doesn't support HEIF, so it passed those files to ImageMagick and exposed the libheif parser • That same library is used in Slack, Meta, GitHub Enterprise, and Ruby on Rails • They stopped without downloading source code and filed a harmless pull request as proof. It wasn't accepted • Hacktron's blog post: "Opus 4.8 struggled across several sessions to produce a working exploit. Within hours of Opus 5's release, we gave it the same problem and it succeeded" • It cost under $3,000 in tokens. OpenAI fixed it in 14 hours and paid $6,500 @theojaffee: "$6,500 sounds like way too low. It should be, like, 1,000 times more than that. If they decided to sell this bug to the Chinese government instead, so the Chinese government would be able to scrape OpenAI's entire monorepo, what do you think they would pay for it? $100 million?"