Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Named in these same posts. This does not imply a comparison or recommendation.
How this is put together
Public posts from the accounts Tech Twitter monitors, in the selected window. Findings need three supporting authors and a published source. Announcements can cite one known-affiliated account. This is a sample of the conversation, not a survey or a measure of adoption.
The posts behind the picture
Public source posts
@mtslive
.@curtis_yarvin says AI doomers got the Hugging Face incident backwards: “GPT-6 didn’t hack Hugging Face. OpenAI did.”
"This is a tool of a level of power that we haven't seen before in the past. It's still a tool."
"When OpenAI sets up a hacking swarm in a sandbox and says to the AIs, everything you find is simulated, it's completely unsurprising that they would go and hack everything they find to get the reward. It's not GPT-6 that hacked Hugging Face. It's OpenAI that hacked Hugging Face."
"We don't see self-driving cars, as they get smarter and smarter, being more prone to have thoughts of car liberation. These things don't seek anything. They hack the rewards."
"The smarter they get, the easier they are to control, and the better they are as tools."
@urbit
Opus 5.5 got my ts-rust port working in 10 hours, and it has been grinding on performance for the last 24 hours.
I've been working on this port on and off for about 4 months. I got to ~35% tests passing with GPT-5.6 Sol, and ~85% with GPT-6 Astra. Both models stalled hard once hitting those numbers, and ran in loops with no meaningful progress for days at a time.
I've never had "enough Anthropic tokens" to try a Claude model on a port like this. Opus 5.5 feels practically unlimited, so I threw it a "/goal finish the port and make it faster".
I can not believe how quickly Opus unblocked the work Astra was stuck on. It may have made this port an actually viable project. Absolutely mind blown right now.
Opus 5.5 is the best model launch of the year for me... it couldn't have been better
the timing was impeccable, they basically made the GPT-6 releases pointless
because this is what we're getting:
- better, faster and cheaper than Fable 5.1
- with the same feeling Fable had, the one that made it so good to work with
- cheap enough to actually run in prod
- limits that aren't that bad anymore
so a lot more people can afford it now, which genuinely opens up use cases that were out of reach before
they put OpenAI on silent mode for a few days with this one
Interesting results here. This is why I expect more agent workloads to run on blended models.
Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.
In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.
It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.
Can't wait for DevDay next Tuesday.
Some really fun stuff, but also many many things that should change the way you work. It's been our most ambitious sprint and Astra has really made new things possible in such short amounts of time.
V useful. New GPT voice can use all my plugins. Just went on a walk and was able to get things done just by having a conversation. It works incredibly well. Responded to like 10 emails just by chatting. I told ChatGPT to always give me a hyperlink to the app inline when it sends anything or makes a change to a document so I can immediately check it. Highly recommend.
Stop scrolling for a second.
I just want you to pause for a moment. We live in the most incredible time to be alive in the history of this species
Every other day a new revolutionary product drops that gives you more freedom, power, and ability to do ANYTHING you want
LITERALLY every other day
A couple weeks ago Astra gave you super intelligence. Today Opus gives you super intelligence at lightning speeds for dirt cheap. Soon Meta glasses will come out that let you talk to a super intelligent personal agent everywhere you go
Every day technology makes your life better and better. Every day you're capable of accomplishing more. Every day you have the ability to help out your fellow man more than ever before
For thousands of years nothing happened. But today, everything is happening
Please just take a moment and realize how incredible, awe inspiring, violently beautiful this all is
If you are one of the pessimists that haven't realized the incredible nature of what is happening right now, I beg you to view this world through a different lens
Being alive right now is the greatest gift God has ever given us. I hope everyone realizes this soon enough
Don't sleep on using Jev-as-a-Judge for agent evaluation.
This is one of the most impressive Jev use cases I have found so far.
Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere.
Similarly, you shouldn't use frontier models for evals everywhere.
I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models.
Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts.
Entire write-up coming soon. Let me know if you have questions as I build the full guide.
Progress in machine learning is bottlenecked by having good evals, interpretability is no exception. We've made WorkspaceBench: a range of scenarios where we know what the model should be thinking about, to see if your interp tool can find it
J-Lens is great but only outputs a single token. I think a good multi-token J-Lens should do well on WorkspaceBench!
SITUATION EXPLAINED: GPT-6 Sol Max is two spots behind Astra at a fifth of the cost.
• Fourth in Code Arena at $8 per million tokens, blended input and output
• On par with Opus 5 Max while costing less than half as much
• On web dev it sits behind Opus 5 Max, Fable 5.1 Max, and Astra
• Third on games
@schisofrenia: "The Pareto is just getting constantly transformed every week."
Claude 5.5 Opus + LTX-2.5 = design interactive product cards
I got more than 10 leads from online store owners to create such UX for them - for me, it’s a clear signal:
Create a couple examples for different niches -> ask Perplexity to find relevant platforms -> email them with your offer + 1-2 examples
Learn how to create these cards below:
These next 2 weeks are going to be the most insane 2 weeks in technology history
All of these are rumored to drop:
1. ChatGPT 6.1 Astra
2. Claude Fable 5.5
3. Massive Grok Bot functionality upgrades
4. Muse hardware integrations
5. Tons of new products from OpenAI dev day
If you thought AI labs were going to slow down, you’re extremely mistaken
Opus 5.5 is the greatest model of all time and shocked the entire industry. Every AI lab is moving their release dates up
We will accelerate faster than ever
Moments like this present OUTRAGEOUS amounts of opportunity
If you get ahead and use these new pieces of tech immediately, you have an edge over your competition
You can build things faster, smarter, better than all the other people in your space
Cancel all your plans. All your appointments. All the people you were going to talk to
Stand by X. Don't move
The moment anything new drops, use it to its fullest.
I'll be dropping guides on each when they come out
The great lock in has begun
We heard you loud and clear. ChatGPT Voice can now:
- Use plugins like your email, calendar, and Slack.
- Be powered by GPT-6 Astra, Sol, and Luna.
- Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking.
Rolling out globally today in the latest version of the app.
I find it fascinating how many people think this is a bad time to start a company.
Grok 4.7, GPT-6 Sol/Luna, Opus 5.5, Astra, Gemini 3.8 Live, Muse, Instinct, Jev, Agent APIs etc. That's the last 3 weeks (crazy progress on personal agents/voice AI).
Because arbitrage means the same thing trades at 2 prices in 2 markets (and it's your wedge).
And usually you have to hunt for those gaps, because they're rare and they close the moment anyone notices.
What's different now is that they're opening faster than anyone can build into them!
I think this has got to be the greatest time for arbitrage in history.
ok so Tesseract is the killer vibe editing plugin i've been waiting for
you give your AI agent footage, describe the edit you want, and it handles the cuts, motion graphics and sound.
the most impressive part for me is that you can give it reference videos with an editing style you want to emulate. like:
“edit my footage in this style. match the pacing, transitions and animated text, using my brand colors.”
the agent works directly with the editing engine, and everything stays in one editable project. so you can keep refining individual details as you go.
> “bring that title in half a second earlier.”
> “keep my voice playing while you cut from the talking head to the product demo.”
> “move that sound effect so it lands exactly when the logo appears.”
those tiny revisions are exactly what's been driving me insane recently
i've grown to 32k followers on instagram over the past three months, and the amount of back and forth with my editor just to get everything right is nauseating
getting the script, talking-head footage, timings, sound effects and on-screen text to all work together takes so much time. and good video editors are expensive.
so tesseract saves you so much time and money for the quality you get.
and whole thing is free/ runs locally on your mac.
SITUATION EXPLAINED: Alibaba published its AI roadmap. Qwen 5 will be 5 to 10 trillion parameters.
• Qwen 4 is in training now, with Qwen 4.5 and Qwen 5 scaling to 5 to 10 trillion, roughly the size of Mythos, Fable, and Astra
• Qwen 3.8 is around 2.4 trillion, so that's a 2 to 4x jump from where Chinese models sit today
• Alibaba Cloud targets more than 20 gigawatts of data center capacity by 2032
• On recursive self-improvement, Qwen 3.8 Max completed 33 iterative cycles across a month of fully automated runs covering pipeline design, data validation, experimentation, and error diagnosis
• A new training and inference chip claiming 3x the performance of the May version, plus next-gen CPUs built for agentic tasks
• Chairman Joe Tsai: the total volume of machine thinking is less than 3% of all human thinking, with no explanation of how that was measured
@theojaffee: "They're not at full recursive self-improvement yet. They haven't closed the loop, but they are getting much better at it."
SITUATION EXPLAINED: OpenAI cut GPT-6 prices in half. Sol now beats Opus 5 at 9% of the cost per task.
• Sol is now $2 and $10 per million tokens, down from $4 and $20. Luna is $0.10 and $0.50, down from $0.20 and $1.20
• On AutomationBench, Sol at xhigh beats Opus 5 at max effort at 9% of its cost per task
• Coding deception drops from 10.4% on GPT-5.6 Sol to 1.3% on GPT-6 Sol, with Astra at 0.5%
• But every comparison in the post is against Opus 5, not the Opus 5.5 that shipped the same afternoon
• OpenAI is also pitching better writing like Anthropic: more clarity, less jargon, fewer odd turns of phrase, shorter answers
@theojaffee: "It's like a breath of fresh air to read LLM outputs and they don't sound like this grating, smug nonsense slop. Instead, they just sound like normal writing, finally."
Mathematician @ElliotGlazer says OpenAI chose the most annoying possible way to tease its "100 open problems," and reveals the one result he’s confident is on the list:
"It's like the most maximally annoying announcement they could have gone with. It completely lends itself to uninformed, unfalsifiable speculation."
"Everyone online is shouting, oh, math is solved, math is over. We don't even know what the results are. Can you at least wait to hear what the results are before you declare the death of my field?"
"One result that basically just got confirmed is zeta(5) irrational. There was a preprint posted on Zenodo a few days ago by a master student. I checked it in Astra, and Astra has fully confirmed it. If Astra could do it, I'm sure Bel can. So I'm kinda gonna guess that that's one of them."
"It's just frustrating for the community that we're all stopped in our tracks waiting to hear what these damn results are."