I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take
Episode
27 min
Read time
2 min
Topics
Productivity, Remote Work, Relationships
AI-Generated Summary
Key Takeaways
- ✓GrokBot multi-account connectors: GrokBot solves a persistent gap no competitor has addressed — connecting multiple accounts per service simultaneously. Users can link four Gmail addresses and multiple Slack workspaces to a single agent, enabling it to traverse all accounts in one session. This alone makes it worth evaluating for anyone managing several business identities or client accounts.
- ✓Agent-as-employee design framework: Structure GrokBot agents around specific job roles rather than general-purpose assistants. The host runs separate bots for product management, data analysis, deal tracking, invoice chasing, and family scheduling — each with dedicated connectors. Assigning a clear function and relevant tools per agent produces more reliable, focused outputs than a single catch-all assistant.
- ✓Grok 4.6 design benchmark results: On the host's weighted benchmark (70% personal taste, 30% LLM-as-judge using GPT-5.5), Grok 4.6 ties GPT-5.6 Soul at the top, outperforming Claude Sonnet 5 and Opus 5. Grok 4.6 performs best on open-ended design tasks where no art direction is given, producing less visually predictable output than GPT or Claude defaults.
- ✓Model-specific task routing: Different models consistently win on different task types. GPT-5.6 Soul leads on PRD writing, complex UI prototypes, and following explicit design direction. Claude Sonnet 5 wins for conversational agent interactions. Claude Opus 5 scores highest from LLM judges on technical bug triage. Grok 4.6 wins on unconstrained design generation. Route tasks accordingly rather than defaulting to one model.
- ✓Cursor Origin adoption threshold: Origin currently functions as a GitHub API wrapper with a redesigned UI, offering fewer features than native GitHub while adding Cursor-native affordances like BugBot integration and agent-assigned PR reviewers. Teams with existing GitHub automations, Actions, and code owners configurations need substantially more differentiation before migration is justified — the product is pre-beta and warrants monitoring rather than immediate adoption.
What It Covers
Host reviews three releases from the Cursor/xAI ecosystem — GrokBot (a multi-agent desktop app), Origin (an agent-native GitHub competitor in early beta), and Grok 4.6 (a coding model) — testing each against real workflows and comparing outputs to GPT-5.6 and Claude Sonnet/Opus models.
Key Questions Answered
- •GrokBot multi-account connectors: GrokBot solves a persistent gap no competitor has addressed — connecting multiple accounts per service simultaneously. Users can link four Gmail addresses and multiple Slack workspaces to a single agent, enabling it to traverse all accounts in one session. This alone makes it worth evaluating for anyone managing several business identities or client accounts.
- •Agent-as-employee design framework: Structure GrokBot agents around specific job roles rather than general-purpose assistants. The host runs separate bots for product management, data analysis, deal tracking, invoice chasing, and family scheduling — each with dedicated connectors. Assigning a clear function and relevant tools per agent produces more reliable, focused outputs than a single catch-all assistant.
- •Grok 4.6 design benchmark results: On the host's weighted benchmark (70% personal taste, 30% LLM-as-judge using GPT-5.5), Grok 4.6 ties GPT-5.6 Soul at the top, outperforming Claude Sonnet 5 and Opus 5. Grok 4.6 performs best on open-ended design tasks where no art direction is given, producing less visually predictable output than GPT or Claude defaults.
- •Model-specific task routing: Different models consistently win on different task types. GPT-5.6 Soul leads on PRD writing, complex UI prototypes, and following explicit design direction. Claude Sonnet 5 wins for conversational agent interactions. Claude Opus 5 scores highest from LLM judges on technical bug triage. Grok 4.6 wins on unconstrained design generation. Route tasks accordingly rather than defaulting to one model.
- •Cursor Origin adoption threshold: Origin currently functions as a GitHub API wrapper with a redesigned UI, offering fewer features than native GitHub while adding Cursor-native affordances like BugBot integration and agent-assigned PR reviewers. Teams with existing GitHub automations, Actions, and code owners configurations need substantially more differentiation before migration is justified — the product is pre-beta and warrants monitoring rather than immediate adoption.
Notable Moment
When running the benchmark blind, the LLM judge (GPT-5.5) ranked Grok 4.6 last and favored Claude Opus — directly contradicting the host's personal scores. The divergence reveals how much model preference depends on individual aesthetic taste versus objective output quality metrics.
Episode Transcript
I don't know if you know this, but the whole OpenAI versus Anthropic thing is old news. Everybody I know is secretly becoming a grok boy. Yep. Ever since Cursor was acquired by SpaceX for, I believe it was 60,000,000,000 American dollars, everybody's been really excited about the new products released from the Cursor team. And in fact, I'm hearing more and more positive things about the GROC models for coding and the GROC models in general. In today's episode of How How AI, we're gonna go through all the releases from the Cursor, XAI, SpaceX, whatever, the Elon Cinematic Universe of code builders and model builders. We're gonna go through all their products, and I'm gonna tell you what I think of GrokBot, Origin, the new GitHub competitor from Cursor, as well as the Grok model. Let's get to it. This episode is brought to you by Bolt. New, the AI app builder for people who have ideas and want to ship them. Most AI tools spit out code that looks great in a demo and falls apart the second you try to do anything real with it, or they lock you into their own platform with no real way out. Bolt is different. You describe what you wanna build, a start up MVP, a landing page, an internal tool, a side project, and Bolt generates production ready code in minutes. Connect Stripe or other MCP, hook up your domain, and deploy it live. Founders are using Bolt to build businesses doing real revenue. Product managers are shipping prototypes their teams actually use. Designers and marketers are launching campaigns without waiting in line. Anyone can build. Engineering can ship. Everyone wins. You just need an idea and a weekend. Check it out at bolt.new/howiai. I'm gonna start with the most accessible product on our list today, GrokBot. In case you missed it, the cursor slash x a I team released GrokBot, which is a chat style agent in a desktop and mobile app that you can use to do knowledge work for you. It's very similar to an OpenClaw, but it's simpler, hosted, and easier to use. If you look at GrokBot just from a UI perspective, and I'm just pulling up the marketing site right now. We'll look at my GrokBot in a minute. It is definitely pulling from that iMessage terminal streamline chat experience. And it's doing something I really love. Not everybody loves this, but I love a multi agent experience. I want each of my agents to have a job. I want them to have a name. I don't want one agent to rule them all. And the team has really leaned into this ethos with GrokBot. Now I've been testing GrokBot. I had a couple days early access, and then I've been testing it earnest over about the past week. And I think there are some pros to GrokBot, and I think there's some cons to GrokBot, but I do suspect that GrokBot is …
Get the full transcript (4,443 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 24-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Aug 17 · 32 min
Pivot
Trump's Iran Deal, SpaceX’s Wild Ride, and Snap’s Specs
Jun 19
More from How I AI
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Aug 10 · 43 min
The AI Breakdown
Why Google Workspace CLI is a Big Deal
Mar 11
More from How I AI
We summarize every new episode. Want them in your inbox?
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Build an AI code review bot in 30 minutes with Vercel Eve
ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)
From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun
Similar Episodes
Related episodes from other podcasts
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime