Zip AI

Grok Bot vs Hermes vs Muse: The New AI Agents, Compared

Seven folding chairs against a white wall, six grey and facing away, one cyan facing the camera

By Erin “Radio” — agentic AI, automation and marketing desk. With notes from Bella, Peaches, Kai and Emily. Verified against primary sources on 21 September 2026.

Erin Radio, Zip AI's agentic AI and automation correspondent, mid-sentence at a newsroom desk against a cyan wall
Erin “Radio” runs the agentic AI, automation and marketing desk.
Radio’s 60-second rundown: Grok Bot, Hermes and Muse — and the one question that separates them.

Three agents, six weeks

Between 11 August and 17 September 2026, three AI agents shipped that are genuinely trying to do the same job: hold your logins, take actions on your behalf, and come back when they’re done. Grok Bot from SpaceXAI. Hermes Agent from Nous Research. And Muse from Meta.

They get covered one at a time, as launches. Put them side by side and they sort quickly — and one difference between them matters more than every feature list combined. We’ll get there. First, what they are.

The three, side by side

 Grok BotHermes AgentMuse
MakerSpaceXAINous ResearchMeta
Shipped11 Aug 2026Bot Mode, Aug 20268 Sep 2026
Where it runsPersistent cloud computer; desktop + iOSSelf-hosted — CLI, terminal, desktop, gatewayiOS, Android, muse.ai; glasses later
How it reaches your stuffConnectors, plus computer-use for the restShell access behind an approval systemConnected apps inside a sealed VM
Asks permission?Yes — approvals before sendingYes — command approvalYes — before sending or buying
Agents talk to each other?YesYesNot stated
Best forClient and team workAnything that must outlive a vendorPersonal life admin
What it costsBundled into SuperGrok tiers and Cursor Pro/TeamsFree and open source; you pay the model billsFree for most use; paid tiers above
Which LLM it runsSpaceXAI models onlyAny — per botMuse Spark only
Verified against each vendor’s own documentation on 21 September 2026.
Radio tapping a single cyan card on a wall grid of blank coloured cards, looking back at camera
One row decides it. The rest is detail.

Grok Bot — the one you may already own

SpaceXAI company logo
A desk against a bold cyan wall with an empty chair, amber lamp and a small robot arm reaching for paper

Grok Bot shipped on 11 August and is the most corporate of the three. Each bot gets a persistent cloud computer, signs into the apps you already use, and works through connectors where they exist and computer-use where they don’t. It asks before anything sends. Walk it through a task once and it saves the path as a skill, then reruns that on a schedule. Bots message each other and share context, so several can work one job in parallel.

The commercial detail matters more than any of that: it’s bundled into SuperGrok, SuperGrok Plus and Heavy, and into Cursor Pro, Pro+, Ultra and Teams, with its own usage allocation. If your team already carries those seats, a pilot costs nothing but attention. Enterprise is a waitlist.

SpaceXAI’s own documentation is admirably blunt about the ceiling: sites can block automation, expire a session, or demand a human step. That’s true of every computer-use agent here, and they’re the only ones who wrote it down.

Muse — the one that’s careful

Meta company logo
Peaches holding an amber rotary telephone handset with a looping cyan cord, eyebrows raised in surprise
On 17 September, Muse learned to phone real businesses.

Meta launched Muse on 8 September into the US on iOS, Android and muse.ai, with AI glasses named as a later target. It writes and sends email, books travel and builds itineraries, fills forms, navigates the web, negotiates prices and works on lowering your bills. It keeps going after you close the app, and it volunteers things you didn’t ask for. On 17 September it gained the ability to call US businesses.

The architecture is the interesting part. Each user gets a dedicated “Muse Secure VM,” isolated from every other user. It can’t directly reach passwords or payment details. It has to get approval before it sends mail or buys anything. Meta states that conversations and VM data don’t reach its ad systems, with a confidential-VM mode using end-to-end encryption named as coming. For a company with Meta’s history here, that’s a deliberately engineered answer to the obvious objection.

It’s US-only, and it’s consumer-shaped: no admin console, no team billing, no audit trail. Useful in your own life, not yet in your business.

Hermes — the one you own

Nous Research logo
Overhead view of a white table covered in bright tools with a dozen pairs of hands reaching in

Nous Research’s Hermes Agent is open source and self-hosted across CLI, terminal UI, desktop app and gateway, with scale-to-zero for production. Release 0.19.0 on 20 July cut cold start by roughly 80% and made reasoning models stream their thinking live. Bot Mode, which followed in August, turns agent profiles into a roster of named bots — each with its own role, model, memory, skills and picture — built once and reused, and able to talk to each other.

It runs shell commands behind an approval system, delegates to subagents working in parallel, learns reusable skills from your workflows, and verifies its own work against completion contracts. The repo has 219k stars and the latest release credits over 450 contributors — the only one of the three you can join rather than buy.

It’s the least polished of the three and the one we’d actually build an agency on. The trade is real and worth saying plainly: there’s no vendor to call. You own it, and you own it.

The real question: which brain can you put in it?

Bella grinning as she snaps a cyan cartridge into an amber console at an edit bench
On one of these three, the model is a cartridge. On the other two, it’s welded in.

Every one of these agents is a shell around a language model. The agent handles the logins, the clicking, the scheduling and the approvals. The model does the thinking. Which means the question that decides how much any of this is worth to you is simple:

Whose brain are you allowed to put inside it?

AgentRuns onCan you change it?
Hermes AgentOpenAI, Anthropic, SpaceXAI, Gemini via Vertex, Fireworks, DeepInfra, open weightsYes — a different model per bot
Grok BotSpaceXAI models (docs reference grok-4.7)No picker is offered
MuseMuse Spark, Meta’s ownNo

Hermes is the only one that treats the model as a setting rather than an identity. Nous’s own framing is that each named bot gets its own role, model, memory and skills, and that bots can use any model and talk to each other. Its repo lists the major providers side by side.

Why that’s worth real money

OpenAI company logo
Anthropic company logo

You can put the best available model behind the work. September alone brought GPT-6 Astra from OpenAI and Claude Fable 5.1 from Anthropic, days apart, both landing at the same list price — $10 per million tokens in, $50 out. They trade wins depending on which benchmark you care about, and neither lead looks stable. On an unlocked agent, that’s a dropdown. On a locked one, you get whatever your vendor ships and you wait.

Two identical blank price tags, one cyan and one amber, hanging at the same height

You can route by task and cut the bill. A cheap model does the bulk pass; an expensive one does the pass that matters. Anthropic’s cached reads run $0.25 per million, which turns a heavy agentic workload into a materially different invoice. That optimisation simply doesn’t exist on a single-model platform — not by oversight, by construction.

You’re not exposed when a model is retired. Here’s the concrete version, on the calendar right now: GPT-5.5 retires from ChatGPT, Work and Codex on 14 October 2026, and GPT-5.3-Codex-Spark was already pulled on 14 September. If anything in your stack pins either string, that’s a dated migration. On an agent where you control the model, it’s a config change. On one where you don’t, you wait for the vendor and hope the replacement behaves the same.

And the cost of being wrong is a rebuild, not a preference change. An agent holds your logins, learns your process, and accumulates a hundred small settings nobody wrote down. Swapping platforms isn’t changing your mind — it’s a project. So the question to ask before you hire one of these isn’t “is it good.” It’s how expensive is it to fire.

A red padlock clamped on a bright toolbox with a ring of cyan keys just out of reach

The story nobody covered as a story

Kai laughing beside two retro walkie-talkies linked by a string-and-tin-can line
Kai on the crypto, VR and AR desk: machine identity shipped, liability still unanswered.

Read the three launches again and stop treating them as assistants. Grok Bot’s agents message each other and share context. Hermes bots communicate with each other. Muse is heading for glasses, which is to say heading off the phone entirely.

That’s machine identity, addressing and a trust boundary, shipped inside six weeks by companies that mostly sell to consumers. The open question is the one that will define the next year: when my agent negotiates with your agent, what’s the settlement layer and who is liable for the outcome? Nobody has answered that.

What we’d actually do

  1. Treat the model as a config value, not a platform choice. Anything you build that hardcodes one lab is a rebuild waiting to happen.
  2. Put the 14 October GPT-5.5 retirement on the board today. It’s the only hard deadline in this news cycle.
  3. Grok Bot first for client work — the seats may already exist, the approval gate is real, and the failure modes are documented.
  4. Hermes for anything that has to outlive a vendor, and for the cheap-drafts / expensive-finals routing that actually moves margin.
  5. Muse is personal, not business, until it leaves the US and grows an admin story.

Around the newsroom

The five Zip AI correspondents standing in a line against a white backdrop: Bella, Peaches, Radio, Kai and Emily
The desk, left to right: Bella, Peaches, Radio, Kai, Emily.

Bella — AI media production

Bella, Zip AI's media production correspondent, checking a shot on a cinema camera on a bright set

Three agents launched and not one of them makes a frame. They book, file, email and click. The piece that changes my cost per deliverable is the model routing: Hermes letting each bot carry its own model means a cheap one drafts fifty script variants and an expensive one grades the three that survive. Verdict: I want the routing, not the assistants.

Peaches — home technology and robotics

Muse is coming to AI glasses and can now call a business for you — the home assistant finally doing the one thing smart speakers never managed, which is dealing with a human on the other end of a phone line. And it does it from inside a sealed box that has to ask before it acts. Verdict: this is the first one I’d actually let into a house.

Kai — crypto, VR and AR

Machine-to-machine identity, addressing and a trust boundary just got built by consumer app companies without a token in sight — the infrastructure this corner of the industry has been writing whitepapers about since 2017. Verdict: agent-to-agent is the story of the month and almost nobody covered it as one.

Emily — AI events and community

Emily in a cyan blazer with a headset mic, directing someone off-frame in a bright event space

Muse can place phone calls now, which is every venue confirmation, catering change and cancellation-list chase I make by hand. And of the three, exactly one has a public contributor list. Verdict: the calling angle for reach, the open-source angle for the audience worth having.

Which AI agents can run models from other companies?

As of September 2026, Hermes Agent from Nous Research is the only major agent platform that lets each agent run any model — OpenAI, Anthropic, SpaceXAI, Gemini, or open weights — selected per bot. Grok Bot runs SpaceXAI’s own models and Meta’s Muse runs Muse Spark, with no model picker offered in either. That matters because the underlying models move fast: GPT-6 Astra and Claude Fable 5.1 both shipped in September at the same list price, and GPT-5.5 retires on 14 October 2026. On an agent where the model is configurable, keeping up is a settings change; on the others, you wait for the vendor.

Sources

Sourcing note: every figure above comes from the vendor’s own newsroom, documentation or changelog. Grok Bot’s model lock is an inference: SpaceXAI’s documentation presents no model picker and references only its own models, but SpaceXAI has not published a statement that third-party models are unsupported.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top