The case for AGI safety products
TLDR: I think there is a category of AI safety tools that is both useful for a broad market and genuinely improves the safety of future models.
Definition of AGI safety products
Products that both meaningfully increase AGI safety and are profitable
Desiderata for AGI products include:
Directly and differentially speed up AGI safety, e.g. by providing better tooling or evaluations.
Are “on the path to AGI,” i.e. there is a clear hypothesis why these efforts would increase safety for AGI-level systems. For example, architecture-agnostic mechanistic interpretability tools would likely enable a deeper understanding of any kind of frontier AI system.
Lead to the safer deployment of frontier AI agents, e.g. by providing monitoring and control.
The feedback from the market translates into increased frontier safety. In other words, improving the product for the customer also increases frontier safety, e.g. instead of pulling work away from the frontier.
Building these tools is profitable.
There are multiple fields that I expect to be very compatible with AGI safety products:
Evaluations: building out tools to automate the generation, running and analysis of evaluations at scale.
Frontier agent observability & control: There will be hundreds of billions of frontier agents being deployed in the economy. Companies developing and deploying those agents will want to understand the failure modes of these agents and get fine-grained control over them.
Mechanistic interpretability: Enabling developers and deployers of frontier AI systems to understand them on a deeper level to improve alignment and control.
Red-teaming: Automatically attacking frontier AI systems across a large variety of failure modes to find failure cases.
Computer security & AI: Developing the infrastructure and evaluations to assess frontier model hacking capabilities and increase the computer security of AGI developers and deployers.
There are multiple companies and tools that I would consider in this category:
Goodfire is building frontier mechanistic interpretability tools
Irregular is building great evaluations and products on the intersection of AI and computer security.
AI Underwriting Company creates standards and insurance for frontier AI safety risks.
Gray Swan is somewhere on the intersection of red-teaming and computer security.
At Apollo Research, we are now also building AI coding agent observability and control tools in addition to our research efforts.
Inspect and Docent are great evals and agent observability tools. While they are both developed by non-profit entities, I think they could also be built by for-profits.
Solar power analogy: Intuitively, I think many other technologies have gone through a similar trajectory where they were first blocked by scientific insights and therefore best-placed in universities and other research institutes, then blocked by large scale manufacturing and adoption and therefore better placed in for-profits. I think we’re now at a phase where AI systems are advanced enough that, for some fields, the insights we get from market feedback are at least as useful as those from traditional research mechanisms.
Sufficient Incentive Alignment
I think there are various AI safety subfields where the incentive to reduce extreme risks from AGI are close enough to the incentives coming from direct market feedback.
From my own experience, the incentive landscape has rapidly changed since late 2024, where we saw the first “baby versions” of theoretically predicted failure modes like situationally-aware reward hacking, instrumental alignment faking, in-context scheming, and more. Because it is now possible to study early versions of these phenomena in practice, it is possible to do empirical research and test various hypotheses and mitigations.
Also, probably every company on the planet is currently figuring out how to integrate AI agents into their workflows and that comes with a lot of practical safety challenges, some of which transfer in meaningful ways to future safety risks:
Transfer in time - AGI as scaled up versions of current systems: I expect that AI systems that can automate AI research itself come from some version of the current paradigm. Concretely, I think they will be transformer-based with large pre-training efforts and massive RL runs on increasingly long-horizon tasks. I expect there will be additional breakthroughs on memory and continual learning, but they will not fundamentally change the paradigm.
Therefore, a lot of safety work done today directly translates to increased safety for more powerful AI systems. For example,
Improving evals tooling is fairly architecture agnostic or could be much more quickly adapted to future changes than it would take to build from scratch in the future.
Many frontier AI agent observability and control tools and insights translate to future systems. Even if the chain-of-thought will not be interpretable, the interfaces with other systems are likely to stay in English for longer, e.g. code.
Many mechanistic interpretability efforts are architecture agnostic or have partial transfer to other architectures.
AI & computer security is often completely architecture agnostic and more related to the affordances of the system and the people using it.
Transfer in problem space - overlap of widespread safety problems and frontier problems: There are some problems that are clearly relevant to AGI safety, e.g. ensuring that an internally deployed AI system does not scheme. And there are some problems that have large-scale demand, e.g. ensuring that models don’t leak secure information from companies or are not jailbroken.
There are some clear cases where these overlap. I think most of the relevant examples have one of two properties:
They are blocked by breakthroughs in methods: For example, once you have a well-working interpretability method, it would be easy to apply to all kinds of problems including near-term and AGI safety related ones. Or if you build a well-calibrated monitoring pipeline, it is easy to adapt it to different kinds of failure modes.
Solutions to near-term problems also help with AGI safety: For example, various standard improvements for access control that are used to protect against bad actors inside and outside of an organization are also useful to protect against misaligned future AI models.
Taking AGI & the economy seriously
If you assume that AI capabilities will continue to increase in the coming years and decades, the fraction of the economy in which humans are outcompeted by AI systems will continue to increase. Let’s call this process “AGI eating the economy”.
When AGI is eating the economy, human overseers will require tools to ensure their AI systems are safe and secure. In that world, it is plausible that AI safety is a huge market similar to how IT security is about 5-10% of the size of the IT market. So if you think that AI is eating up the economy, your TAM is some meaningful percentage of the entire economy.
Additionally, a for-profit strategy might enable you to get access to large amounts of real-world data and feedback that you would otherwise have a hard time getting. Therefore you’re able to understand real-world failure modes much better and improve the safety strategies at all levels of the stack.
Automated AI safety work requires scale
Almost any safety plan of a frontier AI company contains significantly automating a meaningful chunk of their safety work at some point. I’ve also previously argued that we should already try to automate more AI safety work and I think there are many ways in which we’re currently not plugging low-hanging fruit as a community.
I specifically envision strategies that can meaningfully scale with compute, i.e. where spending more money on compute continuously and meaningfully gives you higher gains. For example,
A largely automated eval stack that is able to iteratively design, test and improve evaluations. Additional compute would enable you to create more scenarios or run more agents to improve the evaluations.
A largely automated monitoring stack, where additional compute allows you to generate more synthetic data, train for longer, or run more extensive monitoring pipelines.
A largely automated red-teaming stack, where additional compute would allow you to run more attacks or search through a wider space of strategies.
I’d argue that this is currently not quite the case because automated pipelines still require too much human intervention, but it’s already obvious that we can scale automated pipelines much further than a year ago. Extrapolating these trends, I would expect that you could spend $10-100M in compute cost on automated AI safety work in 2026 or 2027. I think the people who will be able to do that are already starting to build automated pipelines now and make conceptual progress on how their problem can be decomposed into more manageable subtasks that can be individually automated.
Conclusion
There are many limitations to this thesis, e.g. many subparts of safety would not benefit from adding the profitability constraint, the transfer arguments are really important, and there are many ways in which you can get “pulled sideways” away from the core problem toward short-term profitability.
However, I think there are meaningful cases where the incentives for profit and safety pull in the same direction, and therefore, markets provide great feedback and allow you to make faster progress. We’ve thought quite a bit about this topic and 2025 and I think AI agent monitoring is the most obvious field where this is true. It likely has broad market demand because many economic actors will want to monitor their agents and it has significant transfer because better methodology generalizes to new failure modes.
You can find a slightly more detailed version of this post here.


