September 22, 2026·5 min read·AIgentic.media

SpaceX's Grok 4.7 Is Two Things at Once

spacexxaigrokmodel-releaseai-coding
SpaceX's Grok 4.7 Is Two Things at Once

One Model, Two Jobs

Every few months, the AI industry produces a release that quietly rewires an assumption nobody had questioned. This week's candidate: SpaceX : a company whose primary business is strapping humans to controlled explosions and hurling them beyond the atmosphere : released a language model that beats OpenAI's GPT-5.6 Sol on coding benchmarks.

The assumption being rewired is that frontier AI models are built by companies that do one thing: AI. OpenAI does AI. Anthropic does AI. Google DeepMind does AI. SpaceX builds rockets, launches satellites, and now, as of Monday, runs one of the most capable coding models on the market.

Grok 4.7 is the first major model release since xAI's consolidation into SpaceX, a merger that turned a standalone AI lab into a division of an aerospace manufacturer. The result is a model whose corporate identity is unlike any frontier competitor: a language model whose safety reviews are signed off by the same engineering chain that certifies Dragon capsule life support.

Benchmarks and Pricing: The Numbers

SpaceX evaluated Grok 4.7 using CursorBench 4.0, a coding benchmark developed by Cursor : another of SpaceX's recent acquisitions. The model completed challenges at an average cost of $4.69 per task, putting it ahead of GPT-5.6 Sol and Anthropic's Fable 5.1, both evaluated in their fastest, most cost-efficient configurations.

The model also outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which test legal reasoning and chip design capabilities respectively. On EEBench, Grok 4.7 fell behind OpenAI's GPT-6 Astra, the current leader on hardware-oriented tasks.

SpaceX credits the performance to a new base model architecture and enhanced reinforcement learning. Engineers gave the model tougher training tasks and longer training cycles compared to its predecessor Grok 4.6. The result is a model that trades some raw benchmark dominance in specialized domains for broad competence across coding, legal analysis, and technical knowledge work : the kind of versatility a company that does everything from launch vehicles to satellite internet would need.

Pricing is unchanged from Grok 4.6: $2 per million input tokens and $6 per million output tokens. A faster variant doubles processing speed at twice the cost, aimed at latency-sensitive workloads.

The Grok Bot Harness

The model ships with what SpaceX calls the Grok Bot harness : a multi-agent framework that lets Grok 4.7 split complex tasks across multiple AI agents running in parallel. Each agent can take on a subtask, and agents verify each other's outputs. This is the same architectural pattern that powers OpenAI's agent systems and Anthropic's tool-use models, but SpaceX has built it around a model that costs roughly half as much as its direct competitors.

The harness positions Grok 4.7 not just as a coding assistant but as an orchestration layer for agentic workflows : a market that every frontier lab is chasing but that SpaceX enters with the unusual advantage of owning its own inference hardware and data centers (also inherited from xAI's pre-merger infrastructure).

Safety at the Intersection of Rockets and Chat

SpaceX claims Grok 4.7 set records on LatchBio and HackerBench, two benchmarks testing a model's ability to refuse malicious biology research requests and block cybersecurity attacks. For a model that might one day interface with autonomous spacecraft systems : even indirectly through code generation : safety evaluations carry higher stakes than they do for a pure chatbot.

The tension is real: safety for a rocket means predictable, deterministic behavior within tight tolerances. Safety for a chatbot means refusing harmful requests while remaining helpful. Grok 4.7's benchmarks suggest SpaceX optimized for both, but the dual-use nature of the model means that every safety decision has implications far beyond the chat interface.

What It Means

The Grok 4.7 release marks the moment when the AI industry's neat categories : "AI company," "rocket company," "social media company" : stop being useful. SpaceX now operates one of the world's most capable coding models. Meta runs Muse, a consumer agent that out-downloaded ChatGPT's mobile debut. Every major technology company is becoming an AI company, and the ones that started as AI companies are becoming everything else.

For developers evaluating their next coding assistant, Grok 4.7 offers a compelling price-performance ratio. For the rest of the industry, it raises a quieter question: what happens to AI safety when the models are no longer built by organizations whose only job is to think about it?

Sources

Want to learn more?

Let's discuss how AI can transform your business.

Get in Touch