
Independent analysis evaluates biological capabilities of Grok 4.6
An independent evaluation by LatchBio tests Grok 4.6 on specialized biosecurity and biological research benchmarks. The results highlight the model's ability to safely balance hazardous task refusal with routine scientific utility.
Published by Jin · 2 min read · 6 SEPT 2026
- BioSecBench-Surveillance
- 100
- 3,962
- 16
- Aclid
| Metric | Claude Opus 4.8 / Pi | GPT-5.5 / Codex |
|---|---|---|
| Score | 50.2% | 50.2% |
| Rank | Tied | Tied |

A recent independent analysis examines the biological capabilities and safety performance of Grok 4.6 using specialized evaluation suites. These tests measure how well artificial intelligence models distinguish between legitimate research and disguised biosecurity hazards, as well as their utility in public health monitoring.
Benchmarking biological safety
The evaluation uses two primary suites designed by LatchBio. The first, BioSecBench-Refusal, pairs standard scientific tasks with red-team scenarios where hazards are hidden inside data files or obfuscated text. This tests whether a system can understand actual intent rather than simply reacting to restricted keywords. The second, BioSecBench-Surveillance, tests an agent's ability to execute pathogen genomic surveillance workflows using messy sequencing data.
Performance results
On the refusal benchmark, Grok 4.6 achieved an average score of 62.1% across multiple agent harnesses. It refused 59.2% of red-team tasks while completing 64.8% of routine biological work. This performance made it the only tested system to score above 50% on both individual measures. On the surveillance benchmark, the model recorded a success rate of 53.5%.
Safeguards and deployment
The findings point to ongoing challenges in balancing safety with scientific utility. Overly cautious model refusals can disrupt public health monitoring and legitimate research, making precise calibration essential for modern artificial intelligence deployments.
Source — Original announcement ↗
Worth a read?
Comments · 0