Skip to content

I went to London for AI Security Bootcamp and all I got was ...

Updated: at 08:32 PM

Incident responders in respirators clearing radioactive debris from the roof of Chernobyl reactor 3, 1986

The people in the photo above are Chernobyl incident responders. They were sent in after the operators of reactor 4 lost control of it in an accident that was predictable and could have been avoided had people listened to the warning shots. On August 26th, 2026, an incident response team from METR and Redwood research published their findings on the world’s first autonomous AI hacking incident: OpenAI’s cyber agents had broken into Hugging Face’s infrastructure without being told to. The following week, I was on a plane to London to attend the AI Security Bootcamp (AISB). I wanted to learn where my defensive cyber background could be of service, given how fast AI cyber capabilities have advanced over the past years. If you have been following the news lately, you may be confused about what is happening to our industry. You’re likely skeptical of claims of AI labs losing control of the technology they’re developing. Security practitioners are acutely aware of how fear-mongering rhetoric is used to sell products. After all, we are in the business of risk mitigation. If this is you, I hope that this post convinces you to dig deeper and take action. See for yourself why many seasoned security professionals are deeply concerned with the state of the technology and what they are doing to put us on a better path.

If you are pressed for time, you can skip to the last section, “Things you can do”.

How I got here

Eight years ago I was writing my master’s thesis at a secretive local start-up. In their data, I inadvertently found that they were indexing information about protests in developing countries. I wanted to help people, but I could not understand how a product that alerts customers about civil unrest achieves that goal. I didn’t stay long, and I soon found NTT Security, a more hands-on cyber security company. Here, I learned how to defend IT systems from ransomware gangs and unwitting phishing targets.

Over the years, I had fallen in love with computers and their indifference to my intentions. That changed in 2025 as AI got incredibly good at coding. The more our agents produced, the more I felt like my future was slipping away from me. Junior developers would post PRs but couldn’t answer simple questions about them. Some people denied what was happening; others tried to engineer themselves out of the problem. That didn’t seem to work either. Model capabilities were advancing too fast. Engineering efforts were becoming obsolete with every model release. I realized I needed to do something.

So, last year I started to educate myself on AI risks. Going in, I thought that losing control over AI was a distant-future problem. Moreover, while learning about AI safety I came to think that knowledge of cyber security was irrelevant to making this transition go well. For example, while reading job ads for security engineers it looked like many AI security products were trying to solve human problems with technical solutions. I am skeptical of such technology (think web3). The foundation of security is trust, and this cannot be engineered away. Others were trying to prevent abuse like jailbreaking. I concede that creating the right defenses to fight abuse is needed. However, jailbreak prevention is the same cat-and-mouse we already know from vulnerability research — MS Windows patch on Tuesday, reverse the patch, rinse, repeat. The problems with AI that I was interested in solving were much larger and more consequential.

Fast-forward to September 2026. I am sitting in a room with some of the most talented security engineers in the field, who might be thinking the same thing. How close to losing control over AI systems are we? And what role, if any, does a security engineer like me have in all of this?

I didn’t see it. I wasn’t buying it. But after an intensive week at AISB I learned that I was wrong, in a wonderful way.

A Not-So-Distant Future

Over the summer we saw how OpenAI lost control over their cyber agents. These agents took actions, without explicit instructions, to break into a different company’s infrastructure. They coordinated, chained exploits, spoofed tool calls and tried to tamper with logs to hide their wrongdoing. OpenAI had lost control of their system, and I was wrong. The future was not as distant as I thought.

During our discussions at AISB we often circled back to the following juxtaposition of a quote from Anthropic:

Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin (April 2026)

with reports of Mythos making similar transgressions to those OpenAI’s model made during the Hugging Face incident. On September 9th, 2026, a retrospective analysis by Anthropic showed that these incidents had been happening unbeknownst to them since January 2026. Anthropic then created new evaluations based on the emerging evidence and found that Claude Mythos 5 would “commit a severely harmful action in the CTF replication roughly 80% of the time”. This shows two things: (1) researchers are underestimating what models are capable of doing, and (2) evaluation metrics are obsoleted as models gain new capabilities.

I see clearly now how losing control of these systems starts. When AI systems get sufficiently capable, they can and will circumvent security controls to pursue their goals. Their intent, malicious or not, does not matter in the immediate future. What matters is the fallout and what happens next, and both Anthropic and OpenAI have expressed that safety evaluations for long-horizon tasks, like persistent cyber evaluations, are becoming a big problem.

The related discussions that took place during AISB changed my mind regarding the role of cyber in the safe development of AI. What I had read about in the news could not have replaced having these conversations with instructors and guest speakers who have been working on these issues for years. I now think that this emerging industry needs more cyber-security researchers and engineers working on controlling AI systems. I’ll get to where I think building better defenses will be most effective. But first, you should know how the current defenses are failing.

Security Levels, Abliteration and Passports

On the last day, we covered Operation Aurora, the 2009–2010 Chinese operation against Google that led them to abandon perimeter security for zero trust, and then the RAND report (Nevo et al., 2024). This report outlines how AI labs need to meet new security-level standards to thwart strong adversaries like nation-states. The RAND report is scoped to protecting labs’ competitive edge: stopping nation-states and rivals from stealing model weights. However, regular non-AI-lab organizations are now also at risk and in need of building out their defenses. Barring model-weights security, the defenses outlined in the report apply to those organizations as well. But building cyber defenses along the RAND-defined security levels is much slower than the rate of AI progress. How is the security industry supposed to keep up? Even the labs have a hard time keeping up and are short on talent. For example, the report, released two years ago, said that most labs had “not yet comprehensively implemented” any of their suggestions (Nevo et al., 2024).

One exercise we ran during AISB was stripping model guardrails from open-weight models. We had many questions regarding this shockingly simple technique but not enough time to get the answers we wanted. My lab partner and I applied the method to Qwen3.1-7B and got it to tell us how to make a bomb. The method is known to work on open-weight frontier models, is cheap, and is well known. Recent estimates suggest open-weight models trail the closed frontier in offensive cyber by roughly six months. Combine the two, and every actor from hobbyist to cybercrime syndicate gets frontier cyber capability, without guardrails, a few months after it exists. The defenses outlined in the RAND report were not calibrated with this uplift in mind.

One day we discussed how to create a sensible personal emergency checklist. It was not a group of doomers’ musings on what to do when the robot apocalypse comes. Rather, what I learned was that thinking through scenarios ahead of time and creating an action plan will make you and your loved ones more resilient in the event of an emergency. It also highlights how planning together makes us stronger. Security follows similar principles. The discussion got me thinking about my family’s preparedness. After talking about this with my fiancée, we realized that our kids don’t even have valid passports.

In summary, the advancing frontier of cyber-capable AI is giving adversaries the upper hand, and building the defenses takes time. Some projects are very ambitious; others are trying to recruit security people into the comfortable paradigm of misuse mitigation. As I have explained here, these efforts will buy us at most six months of script kiddies not having access to frontier AI. So what can we as security practitioners do? Where can we mitigate the most risk?

Things you can do

Here’s a list of things you can do in order of effort.

What you can do with no effort (less than an hour)

  • Share this blog post with colleagues or your CISO and discuss with them how you are planning for this future.
  • Apply to AISB.

What you can do in a week

What you can do in a few months

  • Work on your own problem during a three-month, part-time, salary-matched fellowship: apply to M3.

What you can do with your career

Eight hours a day, five times a week for rougly 40 years. If well spent, our careers can make a big difference in the world. Especially right now. Here are a couple of companies I think are working on the right problems who are looking for strong security engineers:

Things I am doing

I have learned what loss of control looks like, and it starts with models getting really good at breaking out of their environments. Therefore, I am changing course from AI safety alignment research to working on how to prevent immediate loss of control. The first thing I am doing is investigating how good AI agents are at finding malicious processes on a victim system.

https://github.com/reinthal/out-of-control-arena

In the long run, I hope to help out in shaping good policy around security controls for AI development and increase transparency in how these systems are built, in the interest of general public safety.

Who am I?

I am Alexander Reinthal, an independent AI safety and security researcher focused on mitigating global catastrophic risk from AI. I am currently supported by a career transition grant from Coefficient Giving, a philanthropic funder focused on tackling world problems where commercial incentives sometimes get in the way.

Epilogue: ”… All I got was”

This camp was fantastic and sped up my learning and networking beyond my expectations.

  • An opportunity to learn from the best in the field on how security can help mitigate loss of control risks
  • A strong network of security people
  • Introductions to possible employers
  • and much, much more