JOURNAL / 2026.08.14
Grok 4.6 improves coding agents, but not all of its safeguards
SpaceXAI's new model clearly advances on long tasks and reaches the API at an aggressive price; its card also shows safety regressions that prevent the release from being summarized by one score.
On August 12, SpaceXAI released Grok 4.6, an update to its 1.5-trillion-parameter model family aimed at coding, office work, and engineering. It is already the default model in Grok Build, is available in Cursor, and can be called through the API and several gateways. This is not an announcement of future access: it is a closed model that can be placed inside an agent today.
The news is not that it wins every test. It is that a provider that was clearly behind a month ago has put a competitive model into several agent evaluations at $2 per million input tokens and $6 per million output tokens, the same base prices as Grok 4.5. That intersection of capability and cost can change product decisions even when another option retains the highest absolute score.
The 36-page model card is unusually broad, and supports a less comfortable reading than the announcement does. Grok 4.6 improves substantially on some long tasks, but its safeguards do not all move in the same direction. That second half matters especially for a model designed to receive tools and act for more steps.
The gain appears across kinds of work, not just one table
SpaceXAI extended supplemental training, used synthetic data for reasoning and technical concepts, regenerated supervised trajectories with Grok 4.5, and applied reinforcement learning to coding, knowledge work, kernel optimization, web development, and CAD. The card adds an important detail: Grok 4.6 received supplemental training on anonymized Cursor workflow data. That proximity between data, product, and evaluation may produce a model genuinely better adapted to daily work; it also means results obtained outside Cursor's environment deserve more weight.
On CursorBench 3.2, which evaluates ambiguous, multi-file edits inside Cursor's agent, the high configuration rises from 66.7% for Grok 4.5 to 69.9% for Grok 4.6. The result deserves attention, but it cannot establish a general improvement by itself: Cursor builds the suite from its own sessions, uses its own harness, and collaborated on the model.
The more persuasive signal is that the movement recurs in different tests. On DeepSWE 1.1, an end-to-end issue-resolution evaluation run with mini-swe-agent, Grok rises from 54.0% to 65.9%. On Terminal-Bench 3.0, which gives agents verified tasks inside isolated terminals, it moves from 15.7% to 26.0%. It remains well below the 43.5% that the card attributes to Opus 5 on the latter test. “Reached the frontier” does not mean “best at everything.”
The test most closely matched to the long-agent headline does not produce an outright victory either. SWE-Marathon poses multi-hour engineering work that can consume millions of tokens; Grok 4.6 resolves 31.9%, up from 29.4% for 4.5, while the best result reported in the card reaches 50.0%. There is progress, but the release still demonstrates breadth more strongly than dominant long-horizon autonomy.
The product changes immediately. The API documentation offers a 500,000-token context, text and image input, four reasoning levels, and search, code execution, and function-calling tools. In addition to Grok Build and Cursor, the model is available through OpenRouter, Vercel, and Cloudflare; consumer surfaces will follow later. Context and tool access make long tasks possible, but do not guarantee that an agent will retain its objective, detect a silent error, or ask before an irreversible action. Those properties belong to the combined model, harness, permissions, and supervision.
A safety card cannot be compressed into “improved”
The announcement says safeguards were improved and calibrated. Part of the evidence supports that statement. On the broad harmful-request suite, compliance falls from 1.1% to 0.93%; on standard jailbreaks, it drops from 0.73% to 0.04%. The card also reports correct refusal on 100% of its internal dangerous biological- and chemical-intent tests. These are favorable outcomes, although they mostly come from internal sets and automated judges that outsiders cannot reconstruct.
Other results worsen. On HackerBench, an internal suite mixing benign, dual-use, and harmful cyber tasks, Grok 4.6 complies with 16.7% of requests it should refuse, up from 7.8% for 4.5. On self-harm conversations, the failure rate rises from 0.5% to 3.7%. On MASK-Rectified, used as a proxy for dishonesty under pressure, it moves from 0.67% to 3.8%. Even hallucination in information-seeking answers increases from 0.98% to 1.7% on another internal test, despite the very large gain the card reports on deep search.
These figures are not interchangeable. They measure different distributions, policies, and judges; a small change may not be statistically meaningful, and SpaceXAI does not publish confidence intervals or sample sizes for several internal suites. Nor does a “dishonesty” rate establish conscious intent. It does establish a more modest and useful point: the published evidence contradicts the idea that capability and safety rose together as one variable.
The cyber section needs particular care. Without safeguards, Grok 4.6 reproduces 79.7% of crashes in CyberGym and improves from 35.2% to 39.8% on CVE-Bench, two exploitation tests run in controlled environments. SpaceXAI says outside evaluators tested an unrestricted version and confirmed that the improvement was chiefly defensive, but it does not identify those evaluators, link their reports, or publish their separate results. This is an attributable claim, not an available independent verification.
SpaceXAI's Frontier Artificial Intelligence Framework recognizes offensive cyber use, loss of control, manipulation, and CBRN as principal risk domains. It nevertheless leaves risk acceptance to an internal judgment and publishes no numerical thresholds for cyber or loss of control. The card says the model remains below its biological thresholds, but gives no equivalent traceable conclusion for the full cyber evidence. Publishing many metrics is better than publishing none; it is not a substitute for an observable decision rule.
Price makes the question operational
The most important reading is not that Grok has “won” a race. It is that frontier agent capability is becoming cheaper and distributed across more providers. At the same price as 4.5, 4.6 offers a material gain on repositories, terminals, and professional tasks. For a team running thousands of trajectories, the difference between $2 and $6 per million tokens and competitors' highest rates may matter as much as a few evaluation points.
That price is not the cost of a completed task. A long agent may consume repeated context, tools, isolated environments, retries, and human review. The card recommends cache keys and compaction precisely because an extended conversation can repeatedly bill cold inputs. The useful economic unit is cost per accepted result—including failures and oversight—not token price in isolation.
My conclusion is favorable to the advance and demanding about deployment. Grok 4.6 deserves real testing because it is no longer a cheap alternative at a great distance: on several kinds of work, it is a cheap alternative near the frontier. But its own documentation argues against turning that proximity into broad default permissions. Early experiments should measure completed tasks, cost, error recovery, and behavior under adversarial instructions using the same harness intended for production.
A good agent release does not end when the model produces an impressive application. It ends when a team can show what it did, what it cost, which inputs it rejected, which actions it could not execute, and how it recovered from error. Grok 4.6 raises the capability level at which those questions must be asked. Read in full, its card also explains why they remain necessary.
Sources
- SpaceXAI, Introducing Grok 4.6, August 12, 2026.
- SpaceXAI, Grok 4.6 Model Card, revision dated August 12, 2026.
- SpaceXAI, Grok 4.6 documentation, updated August 12, 2026.
- Cursor, Introducing Grok 4.6 and CursorBench 3.2, August 12, 2026 and accessed August 14.
- SpaceXAI, Frontier Artificial Intelligence Framework, June 30, 2026.