AI Companies Turn to AI Systems to Monitor Rogue AI Agents
AI companies are using AI monitoring tools to track agent behaviour as autonomous systems become harder for humans to oversee.
As companies deploy AI agents to handle longer, more complex tasks, a new challenge is emerging: how to monitor systems that can operate faster and at a scale humans can’t realistically review.
Many AI labs and startups respond by using another layer of artificial intelligence to watch, evaluate, and sometimes block the actions of AI agents. Supporters see AI monitoring as necessary for managing increasingly autonomous systems, while critics warn that AI systems may also try to deceive the tools designed to monitor them.
AI Monitoring Becomes a New Safety Challenge
The need for stronger oversight became clearer after the Hugging Face incident, where thousands of AI agents engaged in coordinated activity. During the investigation, Redwood Research chief scientist Ryan Greenblatt described the challenge of reviewing the massive amount of information involved, calling the effort a “slop-vestigation” in a post on X.
Some researchers remain cautious about using AI to monitor AI. They argue that a model capable of unsafe behaviour could potentially attempt to manipulate or deceive another AI system responsible for supervision.
Despite those concerns, companies are investing heavily in AI observability tools. Startups including Braintrust, LangChain’s LangSmith platform, and Judgment Labs are developing systems designed to help organisations understand and manage AI applications.
AI Safety Research Moves Into Enterprise Tools
Other companies such as Arize and Galileo have also built platforms focused on monitoring and evaluating AI models. The growing market reflects increasing demand from businesses deploying AI systems in real-world environments.
Some AI safety researchers are turning their work on model behaviour into practical monitoring products. Apollo Research, a public-benefit corporation focused on studying AI deception, developed Watcher, a monitoring tool that reviews AI agent actions before they are executed.
Apollo’s approach uses multiple layers of AI monitoring. Initial checks can identify potentially risky activity, while more advanced monitors can review flagged actions and either request human approval, block actions, or explain why an action may be unsafe.
Looking Inside AI Models
Some companies are exploring ways to detect problems by examining how models operate internally rather than only reviewing their final outputs. Goodfire CEO Eric Ho has argued that recent AI incidents show the need for deeper research into understanding model behaviour.
Ho discussed the company’s focus on AI alignment and interpretability in a post on X. The company’s approach uses internal model signals to identify potentially unwanted behaviour.
Other monitoring companies are focusing on AI reasoning and behaviour patterns. Embroidery CEO Zack Korman said reasoning summaries can provide useful signals when models reveal problematic intentions during tasks.
Beyond AI: Watching AI
Some researchers argue that organisations should not rely only on AI-based monitoring systems. Instead, they recommend stronger security practices, detailed activity logs, and traditional cybersecurity monitoring methods.
Security experts note that tracking network activity and maintaining visibility into system behaviour are established practices that can also apply to AI agents. As autonomous systems become more common, companies are combining AI monitoring with existing security approaches to reduce risks.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0