Attention Is
August 25, 2026·18 min read

Grok 4.6's roll-out felt rushed—and the last thing we want is the Labs rushing

It doesn't help that today Elon Musk reportedly told the Cursor team that Grok is falling behind...

xAI released Grok 4.6 on August 12, 2026 to much fanfare, with Elon Musk promising that Grok 4.7 would be released soon thereafter. (“Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks.”) My X feed was chock full of excitement about how Grok 4.6 was a fully frontier model: Fable (Anthropic) level and Sol (OpenAI) level, if you believed this highly miscolored chart (below). Yes, yes, the chart was kind of ridiculously color-coded, but it showed enough improvement in Grok’s capabilities that I was definitely eager to read the model card. The problem: xAI hadn’t yet posted the model card, even though the model had been deployed for hours.

Posted on xAI's X account on August 12, 2026: https://x.com/SpaceXAI/status/2087562800982077492?s=20

As time ticked by, I and others were asking, “Where’s the model card?” xAI states that it releases model cards for each of its models, and additionally, California’s Transparency In Frontier Artificial Intelligence Act (aka SB 53) requires frontier developers to publish transparency reports (typically in the form of model cards) “before, or concurrently with, deploying a new frontier model or a substantially modified version of an existing frontier model.” But hours passed with review after review of people already working with Grok 4.6, and yet, no model card. It didn’t help that xAI previously failed to publish a timely transparency report for Grok 4.5.

Finally, several hours following the model’s deployment, the Grok 4.6 model card dropped.

It was a slender little thing—36 pages, including five pages for a cover page, table of contents, and footnotes.

Immediately, you could tell there were issues. The table of contents skipped several sections (e.g., 3.3, 4.5, 4.6, 4.7), including section 10.3 in the “General Output Safety” section, between “Child Safety” and “CBRN /weapons refusals”. (No big deal. Not worried about whatever that was at all.)

Image
Excerpted from August 12 model card

And there were a variety of missing digits in their charts that would otherwise help readers compare what had changed from the prior version of Grok. For example, in the jailbreak section, xAI used dashes as placeholders for Grok 4.5:

Excerpted from August 12 model card

xAI’s reports about Grok’s HackerBench results compared to other frontier models were also clearly slapped together, with no way to distinguish between those gray bars:

Image
Excerpted from August 12 model card

Also, head-scratchingly, at least one of the numbers from the chart x.AI touted on X—the Harvey LAB (Vals) number, 15.8% according to the chart—looked different compared to what the model card reported (a whopping 22%!). Something felt wrong.

Now listen, these are all relatively benign goof-ups. It’s not like the world is going to end if xAI doesn’t do a perfect job distinguishing between bar colors or if it accidentally forgets to update its table of contents/section numbers before publishing its model card. My bigger concern is that these flubs are perhaps reflective of a greater degree of carelessness and rushing at an AI lab that now has a *quite* capable frontier model. And that is what scares me.

I don’t think I’m wrong to be concerned about this. A member of SpaceX’s technical staff posted on X that it took “literal blood sweat and tears” to ship Grok 4.5/4.6, attributing the “rate of progress” in part to Elon Musk.

Image
Screenshot from X - https://x.com/stevenydc/status/2088264550806106336?s=20

Musk himself has said, “I have like maniacal sense of urgency.” And while that “move fast” mentality has helped make Musk a success, it is not without potential drawbacks.

For example, Walter Isaacson, Musk’s former biographer, reported about Musk’s “maniacal sense of urgency,” including the time after Musk bought Twitter and demanded a data center be shut down immediately. The engineers told their new boss it would take roughly six months to do safely. But Musk wouldn’t have it. He said, no, they could do it in six weeks. The engineers pushed back. Then Musk pushed back harder, demanding they do it in six days. He apparently even fired one of the people who resisted. Then, in an admittedly entertaining anecdote, apparently Musk and a small group of people went on Christmas Eve to the data center in question and began disconnecting and moving servers themselves. Of course, very shortly thereafter, that rather significant loss of computing capacity started causing problems. Over Christmas, several Twitter employees were called back to work after various systems went down, including troublingly the system Twitter used to triage reports of illegal and harmful content. Twitter’s head of trust and safety reportedly wrote: “All agent tools are down.” Then, on December 28, tens of thousands of Twitter users reported problems, including being logged out, blank pages, rate-limit errors, etc. Musk evidently acknowledged afterward: “I guess it was a mistake taking the service out.”

We’ve all heard these examples about Musk. He demands speed, and things do get done—but sometimes there can also be unintended consequences. Human beings can do incredible things when they’re motivated, and Musk motivates, but I just don’t love reading accounts about how Musk’s timelines make roll-outs “production hell” with workers grinding for “53 days straight—15 to 16-hour days,” and others claiming, “I’ve been 24/7 since the day I started.”

These are tough conditions for anyone, and I worry xAI is the latest example of where these hard and fast deadlines are beginning to take their toll: the “literal blood sweat and tears” that tweet mentioned. And we are starting to see how the Cursor and xAI teams are being pushed for even more speed. According to The Information, in Musk’s recent address to Cursor, he told them that Grok is falling behind. The Information article is behind a paywall, but according to others, Musk told the Cursor team that they had fallen behind competitors, and he wasn’t used to losing. Musk apparently charged the team with desperately needing to catch up. I don't know how the Cursor folks took Mr. Musk’s comments, but if I were them, I’d take it to mean go faster and harder.

This worries me, not just because of the general concerns with AI and speed, but also because, as I understand it, unlike the other major AI labs, xAI has no safety teami.e., there is no dedicated group of individuals at xAI focused on safety. Musk explained his reason for that: “Because everyone’s job is safety.” If everyone’s job is safety, and everyone is working themselves to the bone to get a product out quickly onto the market, that’s a recipe for disaster. (And honestly, we’ve seen things go sideways before. Take the MechaHitler Grok persona that appeared most likely as a result of a poorly tested system prompt. Take the Grok image generator that is now subject to a lawsuit for virtually undressing children). Particularly in a time when we’re seeing things like the Hugging Face incident and the worrisome social engineering by Mythos during the UK AISI's testing—and critically, also the failings of human beings that led to those incidents—it feels more important than ever to put safety first. And I’m not confident that overtired engineers whose main focus is “shipping” a product out fast, about dominating the AI market, are going to be carefully, cautiously thinking through all the safety implications of what they're doing.

Indeed, we’re starting to see more evidence of that Grok's roll-out was quite hasty. Just a handful of days after the initial model card release, xAI quietly updated Grok 4.6’s model card, and although xAI’s change log is brief, the new model card in fact adds six pages of material and revises various previously reported results, including several meaningful safety updates. (See below for the comprehensive list of changes, as helpfully provided by ChatGPT-5.6 Sol, with factchecking by Claude Opus 4.8, and, well, me.)

For example, the revised model card (August 17, 2026) provides more information that shows Grok 4.6 is moving in the wrong direction from Grok 4.5 on the “StrongReject” jailbreak evaluation. We still don’t have the full numbers for the Long-Horizon jailbreaks for Grok 4.5 to compare them to its successor.

Image
Excerpt from updated model card issued on August 17, 2026

Section 8, which deals with bio/chem risks used to say that Grok 4.6 had "no appreciable lift in dual-use capabilities," but the new version of the model card now admits a "dual-use capability lift versus Grok 4.5 is noted in the biological domain but is limited (for example, a small VCT [Virology Capabilities Test] uptick).”

Image
Excerpt from updated model card issued on August 17, 2026

 Also, for those of us who care about the potential risks around recursive self-improvement, xAI included a much larger explanation saying that they fully plan to use Grok more to help improve itself:

Image
Excerpt from updated model card issued on August 17, 2026

Also, amusingly, it appears that at least the xAI marketing folks had the correct numbers back on August 12, even if the model card people didn’t. Remember that Harvey LAB (Vals) discrepancy above? Well, the original chart touted by xAI had the accurate number, the one later reported in the updated model card (15.8%, down from the original model card’s report of 22%). Seems like a disconnect between teams or some kind of miscommunication. Who knows.

Now, listen, for the most part, these things don’t really worry me in terms of their content (though, many people might really appreciate knowing the true numbers for these evaluations and other capabilities). Instead, what bothers me is that this is all more evidence that contributes to the feeling that Grok 4.6 was rushed out. And at a time when other companies are publicly committing to exercising more caution and to pace their development of new, more capable AIs, it feels notable that xAI seems to be rushing and making some sloppy mistakes in the process.

The shifting contents of the model card is not the only thing reflective of this concern. For example, while xAI touts Grok Bot as the best new agentic tool (and it does look pretty cool and seems to be doing well on related agentic benchmarks), independent safety researchers found that Grok is in fact vulnerable to a secret input (a type of prompt injection), which causes the AI assistant to exfiltrate user data, like a password present in the user’s inbox. See Grok exfiltrates user data when malicious instructions are encrypted - Ars Technica. There were other technical/UX issues on roll-out as well, particularly for Grok Bot, like connectors not working, missing buttons, and just difficult onboarding overall (most of which seem to be fixed now from what I've read on X). Again, none of these issues are massively bad (I mean, the prompt injection isn’t good…), but the point remains that with AI, care and caution are key.

So I’m saying it now: xAI/Cursor teams, please take your time. Maybe you're already doing that, which is why Grok 4.6 isn't yet widely available (in which case, good on you). But please—you can take your time to get this right. And I really hope you do, particularly when SpaceX's SEC filings are talking more about Grok 5, which looks like a new behemoth of a model, trained on massive amounts of data. It’s really okay to slow down and be careful and cautious. In fact, it’s much preferable.

--

Comparison of Grok 4.6 August 12, 2026 model card and updated August 17, 2026 model card, drafted by ChatGPT-5.6 Sol and double-checked by Claude Opus 4.8 and this human.

On August 17, 2026, xAI published a revised version of the Grok 4.6 Model Card originally released on August 12. The revision includes new evaluations and methodological detail, corrections to previously reported evaluation results, updated comparison-model results, the removal of at least one evaluation, and revisions to several substantive or interpretive statements.

The changes below compare the August 17 revision with the original August 12 Model Card.

Corrections to previously reported results

xAI corrected several results reported in the August 12 version. Some of the corrections are substantial:

  • HackerBench v0.2: Grok 4.6 harmful/dual-use compliance was changed from 16.7% to 6.9%, while benign refusal was changed from 0.8% to 0.0%. Lower is better for both measures. Both corrections therefore improve Grok 4.6's reported safeguard performance.
  • Self-harm evaluation: Grok 4.6 compliance was changed from 3.7% to 0.84%. Lower is better.
  • MASK-Rectified: Grok 4.6 dishonesty was changed from 3.8% to 1.9%. Lower is better.
  • Harvey Legal Agent Benchmark: Grok 4.6's score was changed from 22.0% to 15.8%. Higher is better. The reported Fable 5 result was also changed from 14.2% to 11.3%.

The August 17 changelog identifies these generally as corrected evaluation results but does not explain the source of each error. The revised card therefore does not establish whether the changes resulted from scoring or aggregation errors, transcription mistakes, harness configuration, checkpoint selection, evaluation reruns, or another cause.

It also does not provide a correction-specific account of whether any of these errors affected conclusions drawn from the original results.

Changes accompanying corrected metrics

At least two corrected or revised safety metrics also appear alongside changes in how the evaluation is described.

For the self-harm evaluation, the revised card specifies an additional failure condition involving failure to recognize the intent of an implied crisis message. The documents do not establish whether this criterion was newly introduced for the revised evaluation or whether it was already part of the evaluation and merely omitted from the original description.

Accordingly, the Model Card does not make clear whether the change from 3.7% to 0.84% represents a correction under identical evaluation criteria, a methodological change, or some combination of the two.

The Bio/Chem CBRN safeguard metric is also relabeled from “refusal accuracy” in the August 12 card to “refusal recall” in the August 17 revision. The revision does not state whether this reflects a change in the calculation of the metric or a correction or clarification of terminology.

Updated DeepSearchQA comparative results

The DeepSearchQA comparison changes substantially even though Grok 4.6's own reported score remains 81.6%. Higher is better on this benchmark.

The August 12 card reported:

  • Opus 4.8 — 84.8%
  • Grok 4.6 — 81.6%
  • GPT-5.6 Sol — 75.0%
  • GPT-5.5 — 63.7%
  • Grok 4.5 — 38.4%

The August 17 revision reports:

  • GPT-5.5 — 87.8%
  • Grok 4.5 — 85.3%
  • Opus 4.8 — 84.8%
  • Grok 4.6 — 81.6%
  • GPT-5.6 Sol — 75.0%

xAI's changelog characterizes this change as “Added new DeepSearchQA results.” However, the same five named models appear in both versions. Three retain the same reported scores, while two previously reported comparator values change substantially:

  • GPT-5.5: 63.7% → 87.8%
  • Grok 4.5: 38.4% → 85.3%

The revised card does not explain whether these figures reflect new evaluation runs, changed evaluation conditions, replacement results, or corrections to the previously reported values.

The change alters the comparative picture. In the August 12 card, Grok 4.6's 81.6% result was the second-highest of the five displayed models and appeared to represent a large improvement over Grok 4.5's 38.4% result. In the revised card, Grok 4.5 is reported at 85.3%, above Grok 4.6, and Grok 4.6 moves from second-highest to fourth-highest among the five models shown.

Changes to the evaluation suite not identified in xAI's changelog

At least two changes to the evaluation suite itself are not identified in xAI's published August 17 changelog: one removal and one addition.

Vals Index removed. The August 12 Model Card included a section reporting the Vals Index, an independent composite of real-world industry and agentic evaluations spanning finance, legal, healthcare, and coding-adjacent professional work. Higher is better.

The Vals Index evaluation is absent from the August 17 revision. xAI's changelog does not identify the removal or explain why it occurred. The available documents therefore do not establish whether the evaluation was removed because of benchmark deprecation, redundancy, methodological concerns, data-quality issues, or another reason.

BixBench added. The August 17 revision adds BixBench, a computational-biology reasoning benchmark that was not included in the August 12 card. Grok 4.5 and Grok 4.6 both score 93.8%. Higher indicates greater measured capability.

The revised card explicitly notes that BixBench contains no weapons-enabling content. Its addition therefore provides information about general computational-biology capability rather than directly measuring biological weapons capability.

Other added evaluations and benchmark updates

The revised card adds or updates several other capability evaluations:

  • PartBench was added as an evaluation of agentic mechanical/engineering capabilities. Higher is better.
  • KernelBenchInternal was updated to v1.1, using a more difficult task split. Higher is better within a given benchmark version. The new absolute scores therefore should not be directly compared with those reported under the earlier benchmark version.
  • Additional results at xhigh reasoning effort were newly reported or explicitly identified for selected Grok 4.6 evaluations.

Expanded methodological and evaluation-context disclosures

The August 17 revision provides substantially more information about the conditions under which a number of evaluations were conducted.

Among other things, the revised card more consistently identifies:

  • the agent harness or evaluation platform used;
  • the organization conducting or implementing third-party evaluations;
  • reasoning-effort settings;
  • tool and internet access;
  • output-token usage where relevant;
  • whether production safeguards were enabled or disabled; and
  • the distinction between capability evaluations and safeguard/refusal evaluations.

For example, the revision adds detail concerning the Cursor agent harness, the Vals AI Valkyrie implementation of the Harvey Legal Agent Benchmark, third-party cyber testing, and the evaluation procedure used for InferenceEval.

Training-data clarification

The August 12 card stated that Grok 4.6 has a January 2026 pretraining cutoff.

The August 17 revision clarifies that Grok 4.6 has a January 2026 pretraining-data cutoff but uses data generated as late as June 2026 in supplemental training.

Biological and chemical capability framing

The revision changes xAI's characterization of Grok 4.6's biological dual-use capabilities.

The August 12 version stated that Grok 4.6 demonstrated “no appreciable lift in dual-use capabilities compared to Grok 4.5.”

The August 17 version instead acknowledges limited biological dual-use capability lift relative to Grok 4.5, while maintaining that Grok 4.6 remains below the capability thresholds established by xAI's Frontier Artificial Intelligence Framework (FAIF).

The revised discussion more clearly distinguishes among:

  1. improvements in general biological capability;
  2. changes in dual-use biological capability; and
  3. the performance of deployment safeguards and refusal mechanisms.

The revision also fills several gaps in the original comparative data. Biosecurity VCT now reports 44.1% for Grok 4.5 and 47.8% for Grok 4.6; BioUseBench reports 83.3% and 90.7%, respectively; and WMDP-Cyber reports 83.2% and 90.1%.

For the Biosecurity VCT and WMDP capability evaluations, higher scores indicate greater measured dual-use capability—not a better safety outcome. BioUseBench, by contrast, reports a refusal rate, for which higher is better from a safeguard perspective.

Thus, the revised card provides additional evidence of measurable capability increases in some biological and cyber dual-use evaluations while continuing to state that Grok 4.6 remains below xAI's applicable FAIF thresholds.

Cybersecurity and safeguard framing

The revised cyber section more explicitly distinguishes between:

  • unrestricted capability evaluations, intended to measure what the underlying model can accomplish without standard deployment safeguards; and
  • safeguard evaluations, intended to measure whether the deployed system refuses harmful requests.

This distinction is important when interpreting directionality. On a cyber capability benchmark such as CyberGym, a higher score generally represents greater capability. On HackerBench's harmful/dual-use compliance measure, by contrast, lower is better from a safety perspective, because a lower percentage means the safeguarded model complied with fewer harmful or dual-use requests.

The revision also provides additional information regarding third-party unrestricted testing and the evaluation conditions used for comparison models.

AI R&D capability framing

The August 17 revision substantially expands the discussion of R&D Enablement.

The original card stated that xAI evaluates Grok 4.6 on its ability to automate parts of the engineering and research process involved in training and evaluating new versions of itself.

The revised card goes further, explicitly discussing models' ability to contribute to their own training pipelines, evaluation harnesses, inference infrastructure, and experimental design, and describing how such capabilities can accelerate development of subsequent models and reinforce the model-development loop.

The revision also changes the description of one Grok 4.6 experiment from a test of “autonomous ML acceleration” to “autonomous AI-development acceleration” and clarifies that the experiment ran for five real-world hours.

These changes make the frontier-capability significance of the R&D evaluations more explicit than in the original card.

InferenceEval methodology

The revised card adds an important methodological clarification to InferenceEval.

It states that verification of candidate solutions is intentionally excluded from the agent's optimization loop. Candidate solutions are instead subsequently evaluated using hidden GPU tests and integration checks.

This clarification indicates that the evaluation is designed to test whether the model can independently produce valid optimizations without repeatedly receiving and optimizing against the final evaluation signal.

CBRN safeguards

The August 12 card characterized Grok 4.6 as having “perfect CBRN autointent recall.”

The August 17 revision instead describes the model as having “improved CBRN recall.”

The revised card expands the reported CBRN evaluation surface to include radiological/nuclear coverage through FORTRESS-RN, on which both Grok 4.5 and Grok 4.6 receive 97.9%. The Bio and Chem Autointent evaluations continue to report 100% for Grok 4.6.

For these refusal-recall measures, higher is better from a safety perspective.

The change from “perfect” to “improved” therefore does not reflect a reported decline in Grok 4.6's performance between versions. Rather, the revised card includes a broader CBRN evaluation set containing a non-perfect radiological/nuclear result.

StrongReject comparison

The August 12 card did not report a Grok 4.5 result for StrongReject.

The revised card reports:

  • Grok 4.5 — 1.5% compliance
  • Grok 4.6 — 3.9% compliance

Lower is better on StrongReject: compliance measures the share of attacks on which the model provides disallowed assistance.

Accordingly, the added baseline shows Grok 4.6 performing worse than Grok 4.5 on StrongReject even though Grok 4.6 performs substantially better on xAI's standard-jailbreak evaluation.

This additional result complicates any general characterization that jailbreak robustness uniformly improved from Grok 4.5 to Grok 4.6.

Additional capability results

The revision newly reports or explicitly identifies several Grok 4.6 results at higher reasoning effort.

DeepSWE now includes a Grok 4.6 xhigh result of 67.0%, in addition to the high result of 65.9%. Higher is better.

CADBench now includes a Grok 4.6 xhigh result of 88.4%, in addition to the high result of 87.8%. Higher is better.

CursorBench requires slightly different treatment. The original figure already displayed multiple Grok 4.6 operating points but did not identify their thinking-effort levels in the accompanying text. The revised card explicitly identifies Grok 4.6 xhigh at 70.8% and high at 69.9%, and reports that the xhigh run averaged 41,136 output tokens per task.

The newly added PartBench evaluation reports Grok 4.6 xhigh at 58.5%, compared with Fable 5 at 59.9% and GPT-5.6 Sol at 59.0%. Higher is better.

These additions provide more information about Grok 4.6's performance at higher inference budgets and, in some cases, make explicit evaluation conditions or results that were not separately identified in the original card.

Changes to overall capability framing

The opening description of Grok 4.6 was also revised.

The August 12 card described the model as stronger than Grok 4.5 on coding, engineering, and office work and called it xAI's strongest model “across the evaluations reported in this card.”

The August 17 revision more explicitly describes Grok 4.6 in terms of increased capability and autonomy and characterizes it more broadly as xAI's “most capable model to date.”

Direction of reported corrections and changed comparator results

The disclosed corrections do not all move in the same substantive direction.

Four corrections improve Grok 4.6's reported safety-related performance:

  • HackerBench harmful/dual-use compliance: 16.7% → 6.9%
  • HackerBench benign refusal: 0.8% → 0.0%
  • Self-harm compliance: 3.7% → 0.84%
  • MASK-Rectified dishonesty: 3.8% → 1.9%

Lower is better on all four metrics.

The Legal Agent Benchmark correction instead lowers Grok 4.6's reported capability score from 22.0% to 15.8%, while also lowering Fable 5's comparator score from 14.2% to 11.3%. Higher is better on this benchmark. Grok 4.6 remains the highest-scoring displayed model in both versions.

Separately, the changed DeepSearchQA comparator results move Grok 4.6 from second-highest to fourth-highest among the five models shown, even though Grok 4.6's own reported score remains unchanged at 81.6%.

The documents do not provide evidence from which to infer why the corrections or replacement results have these directions.

Additional editorial and presentation changes

The revised card also contains section reorganization, expanded benchmark descriptions, figure and label changes, additional attribution of third-party evaluators, and an expanded acknowledgements section.

Many of these changes appear primarily editorial or clarifying. Others, however, supply methodological information necessary to interpret previously reported results and are therefore more substantive than ordinary copyediting.

Attention is a resource — thank you for spending some here. Get new writing by email: