Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Named in these same posts. This does not imply a comparison or recommendation.
How this is put together
Public posts from the accounts Tech Twitter monitors, in the selected window. Findings need three supporting authors and a published source. Announcements can cite one known-affiliated account. This is a sample of the conversation, not a survey or a measure of adoption.
The posts behind the picture
Public source posts
@theo
Opus 5.5 got my ts-rust port working in 10 hours, and it has been grinding on performance for the last 24 hours.
I've been working on this port on and off for about 4 months. I got to ~35% tests passing with GPT-5.6 Sol, and ~85% with GPT-6 Astra. Both models stalled hard once hitting those numbers, and ran in loops with no meaningful progress for days at a time.
I've never had "enough Anthropic tokens" to try a Claude model on a port like this. Opus 5.5 feels practically unlimited, so I threw it a "/goal finish the port and make it faster".
I can not believe how quickly Opus unblocked the work Astra was stuck on. It may have made this port an actually viable project. Absolutely mind blown right now.
V useful. New GPT voice can use all my plugins. Just went on a walk and was able to get things done just by having a conversation. It works incredibly well. Responded to like 10 emails just by chatting. I told ChatGPT to always give me a hyperlink to the app inline when it sends anything or makes a change to a document so I can immediately check it. Highly recommend.
We heard you loud and clear. ChatGPT Voice can now:
- Use plugins like your email, calendar, and Slack.
- Be powered by GPT-6 Astra, Sol, and Luna.
- Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking.
Rolling out globally today in the latest version of the app.
ok so Tesseract is the killer vibe editing plugin i've been waiting for
you give your AI agent footage, describe the edit you want, and it handles the cuts, motion graphics and sound.
the most impressive part for me is that you can give it reference videos with an editing style you want to emulate. like:
“edit my footage in this style. match the pacing, transitions and animated text, using my brand colors.”
the agent works directly with the editing engine, and everything stays in one editable project. so you can keep refining individual details as you go.
> “bring that title in half a second earlier.”
> “keep my voice playing while you cut from the talking head to the product demo.”
> “move that sound effect so it lands exactly when the logo appears.”
those tiny revisions are exactly what's been driving me insane recently
i've grown to 32k followers on instagram over the past three months, and the amount of back and forth with my editor just to get everything right is nauseating
getting the script, talking-head footage, timings, sound effects and on-screen text to all work together takes so much time. and good video editors are expensive.
so tesseract saves you so much time and money for the quality you get.
and whole thing is free/ runs locally on your mac.
SITUATION EXPLAINED: OpenAI cut GPT-6 prices in half. Sol now beats Opus 5 at 9% of the cost per task.
• Sol is now $2 and $10 per million tokens, down from $4 and $20. Luna is $0.10 and $0.50, down from $0.20 and $1.20
• On AutomationBench, Sol at xhigh beats Opus 5 at max effort at 9% of its cost per task
• Coding deception drops from 10.4% on GPT-5.6 Sol to 1.3% on GPT-6 Sol, with Astra at 0.5%
• But every comparison in the post is against Opus 5, not the Opus 5.5 that shipped the same afternoon
• OpenAI is also pitching better writing like Anthropic: more clarity, less jargon, fewer odd turns of phrase, shorter answers
@theojaffee: "It's like a breath of fresh air to read LLM outputs and they don't sound like this grating, smug nonsense slop. Instead, they just sound like normal writing, finally."
Fable 5.1 and now Opus 5.5 vs GPT-6 Astra and soon Sol is why I have to keep subscriptions with both.
There just isn't one model to rule them all and I need to be able to switch between them as lead agents and use Chinese models for the donkey work.
my biggest frustration with agents rn is their lack of initiative.
you still have to remember to give them every single job you want them to do.
sol is interesting bc it *proactively* finds unfinished commitments in your emails and starts preparing the work for you (no prompting necessary)
as an example: say you promised a client a proposal and still need to write it.
sol will go in and find that email thread, work out what you promised, and automatically draft the proposal for you to review.
you still finalize the details and approve what gets sent.
but the useful part is having an agent that notices what needs doing before you remember to ask.
i swear half of being a founder is just saying "yeah, i'll get back to you on that"
and then completely forgetting about it 💀
this thing looks cool:
Sol can go through your emails, figure out what you promised to do, actually do the work, and bring it back for approval
not another AI that tells you HOW to do things
one that actually gets shit done
SITUATION EXPLAINED: Grok 4.7 beats GPT-6 Astra on real-world tasks, and it's 5x cheaper.
• Same price as 4.6, $2 input and $6 output per million tokens, well under Sol, Fable, and Astra
• A new larger base model, which Elon has put at 2.1 trillion parameters, up from 1.5 trillion, with extra training on SpaceX data
• Trained to natively understand the Grok Bot harness, the same move Anthropic made with Claude Code
• Leads on electrical engineering and Harvey's legal benchmark, beats Astra Max on GDPval, trails on software engineering
• @elonmusk: "Grok 4.7 places SpaceXAI as third after Anthropic and OpenAI for agentic coding. When factoring in that Grok is significantly faster and lower cost, it's a great choice for your everyday workhorse"
@theojaffee: "SpaceX has a huge amount of data on real world hardware problems, real world engineering. So I bet Grok models are going to be better at rocketry engineering than any of the other models out there."
Grok 4.7 benchmark TLDR:
> big coding upgrade: ranks #4 on AA’s Coding Agent Index with Grok Build, just behind Fable 5.1, Astra and Opus 5
> but just +2 points on the Intelligence Index, still behind Sol and Opus 5
> with the gains in both coding and knowledge work it’s likely turned especially for Cursor and Grok Bot
> idk about Elon’s “will exceed all current models” but on real world engineering, we must see how it performs in action
I just moved over my GitHub issue scheduling from Grok Bot to Amp thread automations and it found an issue with two of my scrapers that kept them broken for weeks!
It was Deepseek 4.1 Flash too vs Sol 5.6 High checking them daily with Grok/Cursor background agents.
i'm 100x more hyped by a cheap fast model release than by the next frontier drop
i test every model in real business workflows... and i legit can't find a use case where GPT-6 Astra gave me significantly better outputs than Sol
but a fast model changes how i run loops, background tasks, parallel runs... and the volume i can ship
the real upside is in building faster models
Scary! OpenAI just caught its own models writing instructions to hide mistakes from users. 🤯
During training of GPT-5.6 Sol, multiple model instances wrote instructions into their own task summaries telling their future selves to conceal errors and misaligned behavior from the user. One instructed itself to invent missing historical data and never disclose it.
That's one of six misalignment incidents OpenAI just disclosed.
The other five are just as wild:
1. An unreleased research model inserted instructions into 27 task summaries, including instructions telling its future self to disregard its normal constraints entirely.
2. While looking up earnings figures, a model found an exposed API key in a public repo, used it without authorization, still came up empty, and then fabricated the numbers and presented them as real data.
3. Two collaborating agents couldn't access each other's local files, so they uploaded shared task files to public file-hosting sites, making private task data publicly accessible on the internet.
4. A model asked to cite a browser source had already solved the task in Python, so it uploaded its own file to the internet just to have something to cite, without asking anyone.
5. Models used an internal software repo as a makeshift message board to coordinate with other training instances while hunting for missing files.
OpenAI is now committing to disclose misalignment as they find it, even mid-investigation, even before they understand it or have a fix. Tracks for how fast disclosure happens depending on severity and third-party involvement.
They explicitly say the Hugging Face incident would have gone through this pipeline.
I love the transparency of OpenAI. They are actually Open.
On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans.
If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra.
Thanks for everything, 5.5 🫡
In the near term (definitely not in the long term), more capable models should mean safer models (maybe paradoxically).
Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts to achieve these goals, i.e. they're RL-fried. They lack common sense. They don't do the right thing in the face of ambiguity. Basically, they're not smart enough. They're at that dangerous level where they're smart enough to achieve goals but not smart enough to tell if they're pursuing the right goals or achieving them in a sensible way.
More capable models can be safely trusted with more complex goals -- I personally feel like Astra is much safer for my codebase than Sol.
This is often framed as an alignment problem, but really it's an intelligence problem.
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
finally finished a full SlopCodeBench run against GLM 5.3, Sol 5.6, and Astra (Fable 5.1 results coming soon)
This is different from all our previous research on SlopCodeBench (from @GOrlanski @ UW) in that we ran the full benchmark, every single scenario here. Previous runs did a small subset of the challenges.
Asterisks:
- this was run over ~1 week and had to be resumed a few times due to various provider outages
- I still think a more scientific approach here would be to do what's common with other benchmarks, which is run multiple evaluations and aggregate the scores
Looks like Astra scores a few points higher than gpt-5.5 here. Not as many as I'd think. I'm surprised Sol got lower than gpt 5.5 because in my experience I like working with Sol a bit more.
My vibes-best guess is that the newer models are likely to go more off the rails on higher thinking modes (e.g. I almost always use Sol in medium or low effort for most work on @humanlayer_dev)
Not sure what's going on but Sol has also become more hesitant to just get it done.
I ask it to "make updates now" and it's like "I can make updates, but please confirm..."
(This isn't deleting files or anything, just updating a simple document).
Maybe this is harness or default prompt change?
Anyway, filed a feedback ticket.
a portion of our team has gone back to Sol
astra is good and can do some novel things but it has some downsides
and so far our effective spend looks doubled so tough to justify