Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Named in these same posts. This does not imply a comparison or recommendation.
How this is put together
Public posts from the accounts Tech Twitter monitors, in the selected window. Findings need three supporting authors and a published source. Announcements can cite one known-affiliated account. This is a sample of the conversation, not a survey or a measure of adoption.
The posts behind the picture
Public source posts
@theo
Grok 4.5 was an incredible model for the price: fast, pleasant to use, reliable, solid default model.
Grok 4.6 was a (forgivable) step in the wrong direction IMO: slower and more expensive, using way more tokens per task for a slight edge in intelligence. I get it, though. They have to climb benchmarks.
Grok 4.7 is much harder to forgive. They claimed it would be more token-efficient, and it's less by 30 to 80%. It scores worse than Grok 4.6 in various benchmarks. It's slower, it's less pleasant to use, and real-world costs come out to more than 2x above Grok 4.6, putting it over Astra's costs in real-world use.
Considering how much they've been hyping this model release up, I have to say I'm disappointed. The benchmarks don't tell the whole story, and it is pleasant to use in various real-world engineering tasks, but it feels so 2025 still.
The frontend capabilities are unacceptably bad. The 3D capabilities are nonexistent. It gets stuck in random Gemini-style loops all the time.
This was a very disappointing release. I hope that the SpaceXAI team can acknowledge that and impress us with the next one.
SITUATION EXPLAINED: Grok 4.7 beats GPT-6 Astra on real-world tasks, and it's 5x cheaper.
• Same price as 4.6, $2 input and $6 output per million tokens, well under Sol, Fable, and Astra
• A new larger base model, which Elon has put at 2.1 trillion parameters, up from 1.5 trillion, with extra training on SpaceX data
• Trained to natively understand the Grok Bot harness, the same move Anthropic made with Claude Code
• Leads on electrical engineering and Harvey's legal benchmark, beats Astra Max on GDPval, trails on software engineering
• @elonmusk: "Grok 4.7 places SpaceXAI as third after Anthropic and OpenAI for agentic coding. When factoring in that Grok is significantly faster and lower cost, it's a great choice for your everyday workhorse"
@theojaffee: "SpaceX has a huge amount of data on real world hardware problems, real world engineering. So I bet Grok models are going to be better at rocketry engineering than any of the other models out there."
It happened. Grok 4.7 dropped
Better intelligence than Opus 5. Half the price
Fully baked into my favorite AI agent harness at the moment: Grok Bot
If you haven't tried using cloud cursor agents inside Grok Bot, now is by far the best time to do it
Choose a project you want to work on, connect your github, ask a grok bot to do work on it
It will spin up Cursor cloud agents and write code in the cloud. Lightning fast and incredibly smart
I recommend using a project management tools like Linear or Notion to make a bunch of tasks first, then have cloud agents just tear through them all 1 by 1.
You'll get a massive amount of work done without much oversight.
Big opportunity to lock in right now and get ahead of the curve with new tech
Take my steps up above and get to it
just in time harness
sometimes when i want to build a long running projcet, i vibe code new harnesses to support the use case.
e.g, this week i wanted to really push grok 4.6 by building an age of empire clone that looks polished. i wanted this to run on a vm, report me screenshots of the progress, playtest the game etc etc. if it happened to crash, i wanted it to restart and ping me in slack
i also wanted this harness to parallelize as much work as possible, and to be efficient. so i asked cursor to help build a harness to accomplish this. under the hood, it uses the cursor sdk to work of the long list of tasks that was put in a dag (3000+)
when the harness was built, i could quickly glance over it to see how confident i'd be in it succeeding or not. when i was happy, i started it and let it run. right now, its still going!
can be worth trying out
i am looking forward to grok 4.7 and 4.8. grok 4.6 is pretty much the only model that’s steerable, fast, and doesn’t over engineer the world by default
SITUATION EXPLAINED: Cognition's new model is post-trained on Kimi K3.
• SWE-2 scores 50.0% on FrontierCode 1.1 Main, Cognition's benchmark for whether a maintainer would merge the pull request
• Fable 5.1 gets 50.9 at 64% higher cost. Grok 4.6 gets 48.0, Sol 47.5, Astra 53.3
• It leads everything on Terminal-Bench 2.1 at 92.8
• Post-trained on Kimi K3. SWE-1.7 was post-trained on Kimi K2.7
• First Cognition model with effort levels, all from a single RL run
• SWE-2 medium beats SWE-1.7 at 53 steps per run against 127, and 81% less cost
@theojaffee: "Over the course of the RL run, they push the Pareto curve while preserving its shape. At every individual level of effort, from medium to high to max, it got better and also cheaper."
First impressions on Grok Bot
+ Fool-proof. It shows that it was designed for the mainstream. It's dead simple to get started, use it and connect it to other tools. It just works.
- When you add a new plugin, it doesn't ask you to configure it/authenticate right away. After installing it, you need to go to your plugins and authenticate. A bit annoying.
- No traces for the actions taken by the bot. I would like to see the steps it took to solve a task and the commands it ran.
+ Good selection of plugins. You can also easily build yours. Built a @documenso plugin and wired to Grok Bot in less than 20 minutes.
+ You can teach your bots how to do things by recording yourself (Teach a task). Better and faster than teaching the bot through written instructions.
- No model selection. Grok-only obviously. It'll never allow us to use other non-Grok models, which is expected. So, I just hope that Grok models get better. Grok 4.6 is already good, but there's a lot of room for improvement.
- No onboarding. No example workflows. You install the app, open it and you're on your own.
- Only 2 options when pressing the + button in the text input. ChatGPT, Cursor and others give you a lot more options.
+ It does what you ask it to do. No overthinking, no moral lessons, no extra unrelated stuff.
+ Nice app that works great on both desktop and mobile.
- No command centre. I'd like to have a main place from where I can manage bots and everything. I don't want to create a bot that acts as the main controller.
- No ability to share chats publicly.
Radom one: it would be fun if you could have bots collaborate with bots from other people. Like inviting someone else's bot in your chat and have them collaborate.
Anyway, Grok Bot is really good overall. A few more polishes, and it becomes great.
AA Intelligence Index is updated, and if Astra wasn’t dropped yesterday, Meta would now beat OpenAI with its Muse Spark 1.3.
some details + ranking changes:
GPT-6 Astra: #4 → #2 ↑2
Claude Opus 5: #2 → #3 ↓1
Muse Spark 1.3: #3 → #4 ↓1
Claude Fable 5: #3 → #4 ↓1
GPT-5.6 Sol: #4 → #5 ↓1
Grok 4.6: #4 → #5 ↓1
Kimi K3: #5 → #6 ↓1
GLM-5.3: #5 → #7 ↓2
Gemini 3.8 Flash: #6 → #8 ↓2
(this is dense ranking comparison)
> as I pointed out yesterday, Astra scored suspiciously equal to Sol, now it’s 4 points higher
> Grok 4.6 still shines next to Sol
> Anthropic still tops with Fable 5.1
> what’s changed: AA added new benchmarks, removed GPQA Diamond, and doubled the weight of private/held-out evals to 40%
how to perform advanced research on X:
- setup a simple agent inside Grok Bot and connect your X account
- use the official X MCP to research keywords, bookmarks and trends
- install the Apify CLI to extract the full tweet history of any profile
all of it powered by Grok 4.6... the best model for searching X, since it's the only one with exclusive access
you can have Grok Bot perform research daily and send you a detailed brief every morning instead of having to scroll for hours