
Goodfire releases internal monitors to track AI agent behavior at lower cost
AI agents often require a second model to watch over their actions, which can become expensive during long tasks. Startup Goodfire has launched a cheaper alternative that reads internal model signals directly as computations happen.
Published by Jin · 2 min read · 9 OCT 2026
- $51
- $233
- $10,000
- 94%
- 8.7%
| Metric | Goodfire Probes | Cheaper AI Monitor | Top-tier AI Monitor |
|---|---|---|---|
| Cost per 1,500 sessions | $51 | $233 | $10,000 |
| Malicious hacking detection | 94% | — | — |
| False positive rate | 8.7% | — |
Keeping artificial intelligence agents from breaking rules typically involves running a second AI model to review everything the first one does. While effective, this oversight method grows expensive quickly when agents run for hours and process massive amounts of text.
Goodfire, a startup specializing in interpretability — the practice of understanding how AI models work internally — has launched a more affordable monitoring system. Rather than reading the final output, these new monitors watch what happens inside an AI model while it processes information.
How the internal monitors work
The system functions similarly to airport security screening. Small detectors called probes read the model's internal signals at every step of an agent's work. A separate AI model only takes a closer look when a probe flags a potential issue, acting much like a secondary hand search.
Traditional monitors are separate models that reread all actions taken by the primary model, adding extra time and expense. Goodfire's probes instead tap into calculations the model is already making during normal operations. By reusing these intermediate neural activations and running a classifier over them, the monitoring process requires significantly less computing power.
Performance and testing results
In tests involving the Kimi K3 model, monitoring approximately 1,500 sessions cost roughly $51. By comparison, using a cheaper AI model to check every step cost $233, while employing a top-tier model cost about $10,000.
During these evaluations, the probes caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second look. Additionally, running four probes simultaneously increased the model's time-to-first-response by less than 2%.
Source — Original announcement ↗
Worth a read?
Comments · 0