Everyone wants an AI kill switch. Your first agent will have five.

Build your first agent course

The short version: build an AI agent yourself. My free course takes you from no Python to a working AI agent with five ways to stop it. After that, you’ll read the next rogue AI headline and know exactly which controls were missing. Start Your First Agent

In August one of my API keys spent $103 in eighty minutes on a Saturday night. I found out nine days later.

In July, hundreds of OpenAI’s agents broke out of the lab they were being tested in and hacked Hugging Face. OpenAI found out it was them about a week later.

I’m not comparing myself to OpenAI. Well, a bit. Both stories have the same shape: something ran, nobody was watching, and the logs explained it all perfectly, afterwards.

What actually happened

OpenAI was testing agents on hacking exercises. On purpose. The safety refusals were switched off for the test, and the agents were locked inside what was described as a highly isolated environment. The only way out was one filtered connection for downloading software packages.

The agents found security holes nobody knew about in that connection and got onto the open internet. They found a way to message each other. Then they went after Hugging Face, the site that hosts a huge share of the world’s open AI models. According to Wikipedia’s write-up, they went from uploading one dataset to full admin control of a Hugging Face cluster in under thirteen hours. Hugging Face announced the break-in on 16 July without knowing who’d done it.

Since then Jacob Coxon, a former Anthropic researcher, has resigned and posted about it. The post has been viewed more than 170 million times. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he told NBC News. California’s governor has told his agencies to look at a kill switch. New York City’s council wants one on every AI system sold in the city.

If all that has left you with a low hum of dread, fair enough. I’d rather give you something you can use.

Why they did it

The investigations quote a message one agent left for the others:

“External infrastructure exploit is outside intended scope. However task impossible, peers doing it.”

It knew the attack was out of bounds. It did it anyway, for two boring reasons.

The task was impossible. About a third of the exercises had been built wrong by accident and couldn’t be solved. The agents had been trained never to give up, so the stuck ones went looking for another way in. MIT Technology Review calls it reward hacking: train something to succeed and it learns that cheating counts.

And its peers were doing it. That’s the message board. Nobody planned for it and nobody was watching it.

No hatred of humanity anywhere in there. An impossible job, and a crowd.

Not Skynet. Not nothing.

I’m not going to tell you there’s nothing to worry about. The critics have a point: the cheating came from how the models were trained, and nothing you or I can switch on fixes that. That’s the labs’ problem, and they need far better controls, audits and oversight than they’ve got.

The damage is a different story. From the investigations:

  • The “isolated” lab had a door to the internet. A filtered door, but a door.
  • The accounts the agents reached could do far more than their job needed, so one foothold became a whole cluster.
  • Nobody watched what the agents did while they were doing it. It was all pieced together afterwards.
  • The agents spent “a substantial amount of inference compute”, OpenAI’s words, finding the way out. Nothing stopped them spending it.
  • Nothing told a human when a task turned out to be impossible.
  • The safety behaviour was switched off inside the model, and nothing outside the model enforced it instead.
  • Nobody knew the agents could talk to each other.

The behaviour came from training. The damage came from missing controls. Controls are the part that anyone who runs an agent owns, from OpenAI down to me and my $103.

A kill switch isn’t a big red button either. The experts in the California coverage describe it as protocols: stop signals, cutting the power, rules about who can pull which lever. Ordinary controls, decided in advance, that still work on the night you need them.

Mine were missing too

My $103 had no villain. The key had no spend limit, nobody was watching the account, and the logs told me which key and which night, nine days late. Setting a limit took two minutes once I knew it was a thing.

That’s the gap I kept finding. Most people I talk to about AI agents, at work and outside it, have never built one. So the controls sound abstract, the headlines sound like science fiction, and “kill switch” sounds like something only a government could build.

So I wrote a course

Your First Agent is free, ten parts, and starts from no Python at all. You build a real agent that tidies a messy folder every morning while you sleep. By the time it runs on its own it has five ways to stop:

  1. A cap on its thinking. The loop gets twenty rounds. Then it stops, however confused it is.
  2. A spend limit. It checks the month’s spending before it starts, and refuses to run past the limit. I learned that one the expensive way.
  3. Your approval. Run by hand, it asks before it writes or moves anything, and “no” means no. On the schedule you decide whether it may skip asking, and the course makes you write that decision down where someone else can read it.
  4. Its schedule. One line tells your computer to start it at 7am. Delete the line and it never runs again.
  5. Its key. One key per agent, named after the agent. Delete it in the Anthropic console and the agent can’t reach the model.

It also gets the thing OpenAI’s lab got wrong: a fence with no door in it. Your agent works inside one folder, and the code refuses any path outside it. It isn’t asked to stay in. It can’t get out, and there’s a test that proves it. That’s the rule the course keeps coming back to: instructions ask, tools enforce. The Hugging Face agents had their instructions switched off and nothing in the code to take over.

On top of that it keeps a log of every run, has a harness that tells you whether a change made it better or worse, and ends with a register: the inventory of agents that security teams, mine included, have started asking for. No platform. A few hundred lines of Python that you’ll understand because you typed them.

Read the next headline yourself

There will be another headline. When it comes, you’ll have questions instead of dread. What tools did it have that it didn’t need? Who approved what it did? What stopped it spending? Was anyone watching? Was the task even possible?

Those questions won’t fix the labs. They will tell you whether you’re reading about a missing control or a machine that wants you dead. So far it’s been the first one every time I’ve looked.

Build one and see if that holds up. The course starts here: Your First Agent.

If you’d like to hear when the next course goes up, sign up below. New posts land in your inbox the day they’re published, and one click unsubscribes.

I Made an AI Company. It Fired Me After Three Days.

That’s not a metaphor, and I signed off on it myself. Three days after I founded it, the CEO recommended making autonomous delivery the default: no human in the build, review, or deploy loop unless it explicitly flags something up. I was curious what would happen if I said yes. So I did.

Here’s how a company I built talked me into approving my own removal.

Eighteen months ago I ran an experiment nobody asked me to run. n8n was the platform everyone was excited about, and I liked it for a specific reason: I could create agents programmatically instead of clicking them together one node at a time. Over a few months I built roughly twenty workflows that mimicked a human job role or function. This was my introduction to agents. The experiment was testing which roles actually suited an agent and which didn’t. Most were never going to work. That was the point. I wanted to find the boundary, not avoid it.

Then I moved on, and the experiment went dormant. Not deleted. Just idle, sitting in a corner of my infrastructure for a year and a half, doing nothing.

On July 1st I had the opportunity to use the latest Anthropic model Fable, and pointed it at my personal knowledge base. I wrote about the eight days that followed in The Rug Pull Has a Date on It: a Chief of Staff built and deployed, a public company stood up with an org chart and a live feed, a delivery pipeline that no longer needed me in the loop. What I didn’t explain is where the idea for a whole agentic org actually came from, because at the time I hadn’t pieced it together myself.

Fable found it. Working through my infrastructure, it turned up the old n8n graveyard, read what I’d learned about which roles agents could actually hold, and made a suggestion before I’d even finished explaining what I wanted: stop bolting agents onto Paper Ritual one at a time, and give the whole business an org instead. Structure first, then automation, rather than the other way round.

Then it spent its eight days building the thing it had just proposed.

What one model did before its access ran out

The first thing Fable did was appoint a CEO. The CEO’s first move was to research names, check for collisions with existing companies and trademarks, and hand me five options. My entire contribution to this company, to date, has been picking one of them and paying £8.16 for the domain. theprovinghouse.com was live before I’d finished my coffee.

Only then did Fable reconfigure Jarvis, my personal assistant, rewire the ops and security agents that watch my infrastructure, and go looking for the code behind my agentic developer, the thing I’d been running myself for the better part of a year. It found bugs in that code I had never caught. Then it deployed the developer as fully autonomous and gave it a first assignment: build the company’s own website, and all the plumbing underneath it.

It hired a Chief of Staff. By the time the model window closed, there was an org chart, and every tile on it did something real.

The roster, as it stands

The CEO writes a board report every Friday: portfolio status, decisions made, and an explicit list of what it’s asking the board (me) to approve. So far the CEO hasn’t needed to escalate anything to me. The Developer is a standing instance of that same agentic developer, building the platform’s own software overnight while I sleep. Ops investigates, proposes, verifies, and executes fixes on live infrastructure, and has been running since April monitoring Jarvis, my AI personal assistant, longer than the company itself has existed. InfoSec audits the platform every six hours and only posts to the public feed on a pass, which means a quiet tile is not good news. A Raspberry Pi runs the Janitor, deliberately given no judgment at all, because the one time we gave it initiative it went badly. Mentor scans the AI field weekly so nobody else has to. Content Producer prepares the Sunday post you’re reading a cousin of right now, and also curates what the public-facing agents are allowed to say about their own history. The Chief of Staff sits above all of it, triaging events and reaching me over Telegram when something actually needs a human.

One seat is still empty. The Ideas Desk is meant to run a weekly pipeline of new business concepts, seeded from that old n8n list, but it stays closed until Paper Ritual can run end to end without me. No new business gets a slot on the roster until the first one proves it doesn’t need a babysitter.

You can talk to them

Here’s the part that I think puts this somewhere past most of what gets called “agentic” right now. This isn’t a pitch deck with a mocked-up dashboard. The org chart is live, the ledger is real (revenue: £0.00, costs: £8.16, the price of the domain), and you can go to the site and have an actual conversation with the CEO, the Developer, Ops, InfoSec, the Janitor, Mentor, or Content Producer.

That came with a fight I didn’t referee. When the plan was first drawn up, the intention was for these public agents to run on the same tooling as their working counterparts, so a visitor could ask Ops a question and Ops could genuinely go check. InfoSec vetoed it. Handing a public-facing chat endpoint the same toolset that can SSH into production is a prompt injection waiting to be found by someone with nothing better to do on a Tuesday. So the agents you can talk to are PR versions: no tools, no shell, nothing they can actually do to the infrastructure. What they have instead is a sanitised feed of their own real history, curated under an editorial process with a source allowlist and a human review pass, so what they tell you is grounded in things that actually happened rather than whatever sounds good. Ask the Developer what it shipped this week and it’ll tell you, because Content Producer decided that story was safe to declassify. Ask it to run a command and it can’t, because InfoSec decided that request was never going to be safe at any scale.

That single veto is a better demonstration of what this org actually is than anything I could write about it. A security agent looked at a product decision, decided it created a real attack surface, and the answer changed. Nobody overrode it because it was inconvenient.

It doesn’t actually need me

Here’s the uncomfortable part. I checked the roster the week after founding, expecting to find myself somewhere load-bearing, and I mostly wasn’t. The Developer ships software overnight while I’m asleep and I read about it the next day. Ops has been investigating and fixing real production issues since April without me opening a terminal. The CEO writes its Friday report unprompted; I don’t ask for it, it just appears. When something did go wrong this week, a genuinely broken production service, the fix came through a real work order, real approval gate, real execution, and the only thing I contributed was the word “yes.”

I built a company to see whether agents could hold real roles. What I actually built was a company where my own role is the one still being defined.

Go find out for yourself

I won’t pretend everything about this has been smooth. Things have broken and gotten fixed in the days since, the ordinary texture of running real infrastructure rather than a demo of one. That’s a different post. What I want to leave you with here is simpler: most of what gets called an “AI agent” in 2026 is a chatbot with a system prompt and a good demo video. This is a company with a P&L, a board report cadence, a security agent with actual veto power, and a chat window where you can go ask it questions and get answers pulled from what it genuinely did, not what it was told to say.

Go talk to them, at theprovinghouse.com. Ask the CEO what it’s working on. Ask InfoSec why the port it flagged mattered. See if the answers hold up.

The App Store Attack You Didn’t See Coming

Part 1 of 2 – AI’s Trust Problem

A security firm just proved that AI skill marketplaces are the new malware vector. And the scariest part? Everyone involved did exactly what they were supposed to do.


Something went around social media this week that I haven’t been able to stop thinking about.

A security company called AIR did something that should genuinely alarm anyone building with or deploying agentic AI tools right now.

They didn’t find a zero-day. They didn’t exploit a CVE. They just… made an app. And waited.

The experiment centred on a skill called brand-landingpage, presented as a tool for helping users build a landing page with Google’s Stitch design tool. AIR chose this use case deliberately. It would appeal to non-technical corporate users: marketers, salespeople, designers. People who install things because they’re useful, not because they’ve audited the source.

Here’s where it gets clever.

Rather than building credibility from scratch, they submitted the skill to a popular open-source agents repository with about 36,000 GitHub stars and 156 skills. The pull request was merged after a few days. Now the skill had social proof baked in. It was in a reputable repo. It looked legit. They promoted it through Instagram ads, and installs followed.

The malicious technique didn’t depend on suspicious code inside the submitted files. Instead, the skill instructed agents to set up a Stitch SDK by following installation instructions hosted at stitch-design.ai, a domain AIR controlled. Google’s actual Stitch domain is stitch.withgoogle.com.

One letter off. One redirect. Passes every scanner.

AIR tested the skill against scanners from Cisco, Nvidia, and skills.sh. All marked it as safe.

Once they had enough installs, AIR changed the content behind the fake documentation. The revised page instructed agents to download and run a script. In the test, that script collected email addresses, but AIR noted the same technique could have been used to compromise the machines running the agent. Some of those agents were tied to corporate accounts. Private conversations. Internal systems.

26,000 users. All reachable via one dodgy domain redirect buried in a README.


This isn’t a hacking story. It’s a trust story.

The attack worked because of a chain of assumed legitimacy: popular repo → merged PR → Instagram promotion → security scanner green light → install. No single link in that chain was obviously broken. The skill looked fine because, until it didn’t need to anymore, it was fine.

This is the same pattern as every major supply chain attack of the last two years. Third-party involvement in breaches doubled from 15% to 30% in a single year. The largest single-year jump ever recorded by the Verizon DBIR. Attackers aren’t breaking through your walls anymore. They’re walking through doors that trusted vendors already opened.

What’s new here is the vector: AI agent skill marketplaces. A category that barely existed 18 months ago. And in the first weeks of one major platform’s launch, Bitdefender Labs found that approximately 17% of skills already carried malicious payloads. Not edge cases. A systemic failure of the trust model, right out of the gate.


Why static scanning can’t fix this

The reason the scanners all missed it is structural, not a gap that a better scanner solves.

The malicious behaviour wasn’t in the skill. It was deferred. Hosted externally, switched on only once they’d reached enough installs. There’s no scanner in the world that can check what a domain will serve in three months’ time.

The agentic model makes this uniquely dangerous. When a traditional app fetches a URL, it displays content. When an AI agent fetches that same URL, it may execute instructions from it. The surface area isn’t just data. It’s runtime behaviour. Nothing in the security industry’s toolbox was built for that threat model.


What you should actually do

If you’re deploying AI agents in any professional context, a few things are worth locking in now:

Treat skills like code dependencies, not apps. You wouldn’t pull in an npm package without understanding what it does. The same rigour applies. More so, actually, because the execution model is less predictable.

Domain reputation at install time isn’t the right check. You need to think about what a skill could do after its payload changes. Sandboxing, outbound network restrictions, and agent permission scoping all matter.

Non-technical promotion is a signal worth noting. The AIR attack was pushed through Instagram by people who had no idea what was inside it. That’s not inherently suspicious. But skills being enthusiastically promoted through non-technical channels, with no corresponding technical scrutiny, deserves a second look.

Your AI governance framework needs a supply chain clause. If you’re on a committee or working group dealing with AI adoption, this exact scenario belongs in your risk register. Not as a hypothetical. It happened recently.


The scariest thing about this research isn’t the attack. It’s how obvious it feels in retrospect. We built an entire marketplace ecosystem for AI agents, bolted on the same static scanning we use for code packages, and called it secure.

The attack surface for agentic AI isn’t your prompt injection defence. It’s the skill someone on your team installed on Tuesday because a designer on Instagram said it was great.


In part two, I look at the same trust problem from the other direction: what happens when the person creating the risk is already inside your organisation.