Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Named in these same posts. This does not imply a comparison or recommendation.
How this is put together
Public posts from the accounts Tech Twitter monitors, in the selected window. Findings need three supporting authors and a published source. Announcements can cite one known-affiliated account. This is a sample of the conversation, not a survey or a measure of adoption.
The posts behind the picture
Public source posts
@theo
Opus 5.5 got my ts-rust port working in 10 hours, and it has been grinding on performance for the last 24 hours.
I've been working on this port on and off for about 4 months. I got to ~35% tests passing with GPT-5.6 Sol, and ~85% with GPT-6 Astra. Both models stalled hard once hitting those numbers, and ran in loops with no meaningful progress for days at a time.
I've never had "enough Anthropic tokens" to try a Claude model on a port like this. Opus 5.5 feels practically unlimited, so I threw it a "/goal finish the port and make it faster".
I can not believe how quickly Opus unblocked the work Astra was stuck on. It may have made this port an actually viable project. Absolutely mind blown right now.
V useful. New GPT voice can use all my plugins. Just went on a walk and was able to get things done just by having a conversation. It works incredibly well. Responded to like 10 emails just by chatting. I told ChatGPT to always give me a hyperlink to the app inline when it sends anything or makes a change to a document so I can immediately check it. Highly recommend.
We heard you loud and clear. ChatGPT Voice can now:
- Use plugins like your email, calendar, and Slack.
- Be powered by GPT-6 Astra, Sol, and Luna.
- Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking.
Rolling out globally today in the latest version of the app.
ok so Tesseract is the killer vibe editing plugin i've been waiting for
you give your AI agent footage, describe the edit you want, and it handles the cuts, motion graphics and sound.
the most impressive part for me is that you can give it reference videos with an editing style you want to emulate. like:
“edit my footage in this style. match the pacing, transitions and animated text, using my brand colors.”
the agent works directly with the editing engine, and everything stays in one editable project. so you can keep refining individual details as you go.
> “bring that title in half a second earlier.”
> “keep my voice playing while you cut from the talking head to the product demo.”
> “move that sound effect so it lands exactly when the logo appears.”
those tiny revisions are exactly what's been driving me insane recently
i've grown to 32k followers on instagram over the past three months, and the amount of back and forth with my editor just to get everything right is nauseating
getting the script, talking-head footage, timings, sound effects and on-screen text to all work together takes so much time. and good video editors are expensive.
so tesseract saves you so much time and money for the quality you get.
and whole thing is free/ runs locally on your mac.
SITUATION EXPLAINED: OpenAI cut GPT-6 prices in half. Sol now beats Opus 5 at 9% of the cost per task.
• Sol is now $2 and $10 per million tokens, down from $4 and $20. Luna is $0.10 and $0.50, down from $0.20 and $1.20
• On AutomationBench, Sol at xhigh beats Opus 5 at max effort at 9% of its cost per task
• Coding deception drops from 10.4% on GPT-5.6 Sol to 1.3% on GPT-6 Sol, with Astra at 0.5%
• But every comparison in the post is against Opus 5, not the Opus 5.5 that shipped the same afternoon
• OpenAI is also pitching better writing like Anthropic: more clarity, less jargon, fewer odd turns of phrase, shorter answers
@theojaffee: "It's like a breath of fresh air to read LLM outputs and they don't sound like this grating, smug nonsense slop. Instead, they just sound like normal writing, finally."
Fable 5.1 and now Opus 5.5 vs GPT-6 Astra and soon Sol is why I have to keep subscriptions with both.
There just isn't one model to rule them all and I need to be able to switch between them as lead agents and use Chinese models for the donkey work.
my biggest frustration with agents rn is their lack of initiative.
you still have to remember to give them every single job you want them to do.
sol is interesting bc it *proactively* finds unfinished commitments in your emails and starts preparing the work for you (no prompting necessary)
as an example: say you promised a client a proposal and still need to write it.
sol will go in and find that email thread, work out what you promised, and automatically draft the proposal for you to review.
you still finalize the details and approve what gets sent.
but the useful part is having an agent that notices what needs doing before you remember to ask.
i swear half of being a founder is just saying "yeah, i'll get back to you on that"
and then completely forgetting about it 💀
this thing looks cool:
Sol can go through your emails, figure out what you promised to do, actually do the work, and bring it back for approval
not another AI that tells you HOW to do things
one that actually gets shit done
SITUATION EXPLAINED: Grok 4.7 beats GPT-6 Astra on real-world tasks, and it's 5x cheaper.
• Same price as 4.6, $2 input and $6 output per million tokens, well under Sol, Fable, and Astra
• A new larger base model, which Elon has put at 2.1 trillion parameters, up from 1.5 trillion, with extra training on SpaceX data
• Trained to natively understand the Grok Bot harness, the same move Anthropic made with Claude Code
• Leads on electrical engineering and Harvey's legal benchmark, beats Astra Max on GDPval, trails on software engineering
• @elonmusk: "Grok 4.7 places SpaceXAI as third after Anthropic and OpenAI for agentic coding. When factoring in that Grok is significantly faster and lower cost, it's a great choice for your everyday workhorse"
@theojaffee: "SpaceX has a huge amount of data on real world hardware problems, real world engineering. So I bet Grok models are going to be better at rocketry engineering than any of the other models out there."
Grok 4.7 benchmark TLDR:
> big coding upgrade: ranks #4 on AA’s Coding Agent Index with Grok Build, just behind Fable 5.1, Astra and Opus 5
> but just +2 points on the Intelligence Index, still behind Sol and Opus 5
> with the gains in both coding and knowledge work it’s likely turned especially for Cursor and Grok Bot
> idk about Elon’s “will exceed all current models” but on real world engineering, we must see how it performs in action
I just moved over my GitHub issue scheduling from Grok Bot to Amp thread automations and it found an issue with two of my scrapers that kept them broken for weeks!
It was Deepseek 4.1 Flash too vs Sol 5.6 High checking them daily with Grok/Cursor background agents.
i'm 100x more hyped by a cheap fast model release than by the next frontier drop
i test every model in real business workflows... and i legit can't find a use case where GPT-6 Astra gave me significantly better outputs than Sol
but a fast model changes how i run loops, background tasks, parallel runs... and the volume i can ship
the real upside is in building faster models