xAI publishes a Grok 5 evaluation card after months of criticism

The document covers capability benchmarks, red team findings, and refusal behavior. It omits training data description and pre-deployment testing timelines, which is where critics focused.

Younes Bekrar10 min read
ShareXLinkedInFacebook
xAI publishes a Grok 5 evaluation card after months of criticism

xAI published a 31-page evaluation card for Grok 5 on Wednesday, four months after the model shipped and following sustained criticism from researchers and from two government bodies about the absence of any safety documentation. The document covers capability benchmarks across 24 evaluations, describes a red team exercise conducted with an unnamed external firm, and reports refusal rates across a taxonomy of harmful request categories. It does not describe training data sources, does not state when testing occurred relative to deployment, and does not include the dangerous capability evaluations that Anthropic, OpenAI, and Google all publish for frontier models.

What the document contains

The capability section is thorough and unremarkable, reporting scores on standard benchmarks with methodology notes that are more detailed than most labs provide, including the exact prompting used and the number of samples. Grok 5 performs competitively, leading on two mathematics benchmarks and trailing on agentic coding tasks. The transparency about methodology here is genuinely good and worth acknowledging.

The safety section reports refusal rates on a taxonomy covering weapons, cyber operations, self-harm, and a category xAI labels as speech restrictions, which it uses for requests other labs refuse and xAI does not. The company presents its lower refusal rate as a feature reflecting a commitment to fewer restrictions, which is a defensible product position and makes comparison with other labs' numbers meaningless without careful reading.

The omissions researchers focused on

Frontier model documentation from other labs includes evaluations of whether a model meaningfully uplifts a non-expert attempting to create biological or chemical weapons, and whether it can autonomously conduct cyber operations. Those evaluations exist because governments asked for them and because the labs committed to them at the 2023 and 2024 international summits. xAI made similar commitments and this document does not report them.

The training data omission is more consequential legally than technically. xAI faces litigation over training data in two jurisdictions, and any description would be discoverable. Every lab has the same incentive and most publish at least a category-level description. A researcher at the Ada Lovelace Institute called the omission conspicuous and noted that xAI's document explicitly says data sources are proprietary rather than declining to address the topic, which at least is direct.

The story is rarely the launch. It is what breaks, what ships, and who owns the mess at 2 a.m.
Younes Bekrar

Why this arrived now

Two pressures converged. The European Union's AI Act obligations for general purpose models with systemic risk took effect for new deployments last year and require technical documentation to be provided to the AI Office. xAI has European users and has been in correspondence with the office since spring. The United Kingdom's AI Security Institute also published a note in June observing that it had been unable to obtain pre-deployment access to Grok 5, which several members of Parliament raised.

Commercial pressure matters too. xAI has been selling enterprise API access and enterprise buyers ask for safety documentation as a procurement checklist item. A model with no evaluation card fails that check regardless of how good it is. Several enterprise buyers we spoke with confirmed they had raised it, and one said the absence of documentation had been disqualifying in a bake-off earlier this year.

How the industry norm developed

Model cards started as a research proposal in 2018 and became a de facto requirement through a combination of voluntary commitments, competitive pressure, and regulatory anticipation. There is no standard for what they must contain, which means labs publish what flatters them and omit what does not, and comparing across labs requires reading carefully enough to notice what is missing.

The European Union's AI Act will change that for models above a compute threshold, specifying required content. That specification is being drafted through a code of practice process with industry participation, and the current draft is weaker than researchers wanted and stronger than nothing. Enforcement begins in stages through 2027. Until then, publication remains voluntary and the quality varies enormously.

What to take from it

For anyone evaluating models, the practical advice is to read what is absent as carefully as what is present. A capability section with detailed methodology and a safety section with vague categories tells you where a company's attention went. Comparing refusal rates across labs is close to meaningless without reading each taxonomy, and the taxonomies differ deliberately.

For the broader question of whether voluntary disclosure works, this is a data point on the pessimistic side. xAI published after regulatory and commercial pressure, four months late, with the sections that would have been most useful to safety researchers absent. That does not mean voluntary approaches never work. It does suggest they work about as well as one would expect from companies with strong incentives to disclose selectively.

The next test arrives with Grok 6, which xAI has said is in training. If an evaluation card ships alongside the model rather than four months later, and if it includes the dangerous capability sections, this document will look like the start of a change in practice. If it does not, the pattern will be clear enough that the European AI Office and the UK institute will have a straightforward case for treating xAI differently from labs that publish on schedule. The company has some control over which of those happens and appears not to have decided.


Skarvonix will keep following this beat with reporting grounded in how systems behave outside the launch keynote.

  • LLMs

Keep reading