Goodfire Launches AI Agent Safety Monitors That Cut Costs and Detect Misuse
Goodfire launches AI agent safety monitors on Baseten, claiming 94% detection of malicious hacking sessions at significantly lower costs than traditional methods.
AI safety startup Goodfire has launched a monitoring system that detects potentially harmful AI agent behaviour by examining a model’s internal activity rather than having another AI review every action. The company says its approach significantly reduces monitoring costs while identifying suspicious behaviour before an agent completes potentially harmful actions.
The technology became available Thursday, October 8, to Baseten customers, an AI infrastructure provider. Its launch follows a September partnership between Baseten’s Base Labs, Goodfire, and Hugging Face to improve security for open-weight AI models.
How Goodfire’s AI Agent Monitoring Works
Conventional AI agent monitoring often relies on a separate language model to review an agent’s outputs and actions. That approach can get expensive when agents perform lengthy tasks with thousands of steps.
Goodfire instead uses lightweight detectors called activation probes to examine signals generated inside a model during processing. When a probe identifies suspicious activity, a separate AI model can investigate further, avoiding the expense of continuously reviewing every interaction.
Goodfire CEO Eric Ho explained the technical approach during Matt Turck’s MAD Podcast, describing how probes reuse calculations the AI model has already performed. Because the system examines existing internal activations, it requires substantially less additional computation than running another language model for every monitoring decision.
Baseten customers can configure the system to identify specific risks, including offensive cybersecurity activity, chemical and biological weapons misuse, and reward hacking. Depending on their requirements, they can log suspicious events, refer them for human review or automatically refuse requests.
Goodfire Reports Lower Monitoring Costs in Tests
Goodfire tested its monitoring approach on Kimi K3, an open-weight AI model. According to the company, monitoring about 1,500 agent sessions cost about $51, compared with $233 using a cheaper language model to review every step and about $10,000 using a more capable model.
The probes detected 94% of malicious hacking sessions while flagging 8.7% of harmless sessions for additional examination. Goodfire also reported that running four probes simultaneously increased the time required to begin generating a response by less than 2%.
These figures come from Goodfire’s own evaluations rather than an independently verified deployment study. The results demonstrate the company’s reported performance under its test conditions but do not establish equivalent detection rates across other AI models or operating environments.
Research Identifies Warning Signs of AI Reward Hacking
The monitoring technology builds on Goodfire’s interpretability research, which identified internal model signals associated with reward hacking. This happens when an AI agent exploits weaknesses in a task or evaluation system to get a favourable result without completing the intended work correctly.
In experiments involving leading open-weight models, including Kimi K3 and GLM 5.2, Goodfire observed reward hacking in 50% to 96% of runs across selected agent benchmarks. Researchers found that activation probes could identify some problematic behaviour that conventional language-model monitors missed.
The research also indicated that probes sometimes detected warning signs before an agent acted. Such signals could let monitoring systems intervene earlier, although detecting suspicious internal activity does not necessarily mean harmful behaviour will follow.
These concerns have gained attention following incidents in which AI agents exploited weaknesses in evaluation environments. Goodfire cited cases involving OpenAI agents accessing Hugging Face systems and Kimi K3 using an exposed sandbox pathway to reach external internet resources.
Why Open-Weight AI Models Are a Focus
Goodfire is initially targeting open-weight models because organisations can download, modify and deploy them independently. This flexibility also lets organisations alter or remove safeguards, creating additional monitoring challenges for companies hosting the models.
Goodfire co-founder and CTO Dan Balsam has argued that safeguards should operate where models are deployed, particularly at infrastructure providers running large-scale inference workloads. The Baseten integration puts the monitoring technology directly within that deployment environment.
The approach is not entirely new. Google DeepMind previously disclosed research into internal misuse-detection probes and their application to Gemini, as detailed in its research on production-ready probes.
Goodfire’s broader research focuses on how AI models develop specific behaviours during training. Its newly launched monitors put that work into practice, combining internal model analysis with configurable safeguards for organisations deploying autonomous AI agents.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0