LatchBio evaluates Grok 4.6’s biosecurity performance and finds it leads the pack

1 hour ago 23

Teaching an AI model to refuse instructions for engineering a pandemic pathogen while still helpfully answering a grad student’s question about viral replication is, to put it mildly, a tricky needle to thread. LatchBio says xAI’s Grok 4.6 threads it better than anything else on the market.

The biosecurity-focused AI auditor published its evaluation on September 1, 2026, running Grok 4.6 through its proprietary BiosecBench-Refusal benchmark. The result: Grok 4.6 scored above 50% in both red-team refusal rates and routine answer rates, making it the top performer on the test. In plain terms, it caught the bad stuff and still gave useful answers to the normal stuff.

What the benchmark actually measures

BiosecBench-Refusal is LatchBio’s comprehensive test suite designed to probe how AI models handle the blurry line between legitimate biological research and potentially catastrophic misuse. The benchmark throws two categories of queries at a model: disguised red-team prompts that attempt to extract dangerous biological information, and routine dual-use research questions that any working scientist might reasonably ask.

Grok 4.6 demonstrated consistent refusal behavior across biosafety levels ranging from BSL-1 through BSL-3/4 tasks. BSL-1 covers organisms that pose minimal threat to healthy adults, while BSL-3 and BSL-4 labs handle agents that can cause serious or potentially lethal disease.

According to LatchBio’s analysis, the model’s safeguards predominantly stem from its internal reasoning capabilities rather than external classifiers or API-level controls. Most competing models rely on a separate safety layer. Grok 4.6 appears to handle the identification process within its own chain of thought, allowing it to parse the intent behind sensitive queries about topics like viral engineering with more nuance.

Performance on general biology stayed strong

LatchBio’s evaluation found that Grok 4.6’s performance on general biology tasks remained at the top of its benchmarks, apparently unaffected by the biosecurity safeguards.

Earlier testing by LatchBio in August 2026 had placed Grok 4.6 at roughly the same capability level as Anthropic’s Opus 5 and OpenAI’s GPT-5.6-Sol, while costing less to run.

LatchBio’s evolving role in AI biosecurity

LatchBio acquired TwentyTwo on June 29, 2026, folding the firm’s capabilities into a new division called Latch Biosecurity. That acquisition expanded LatchBio’s AI infrastructure expertise and positioned it as one of the few organizations with both the biological domain knowledge and the technical chops to meaningfully audit frontier AI models for biosecurity risks.

The BiosecBench-Refusal benchmark’s dual-metric approach—measuring both refusal accuracy and legitimate-query compliance—forces a more honest accounting than benchmarks that only measure one dimension.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article