How I think about security in my AI operating system
After I wrote about testing a personal AI agent that asked for my children's passports, people asked me the obvious follow-up: how do I keep my own system secure? Mine works across my email, my files and the platforms I use every day, running nearly 30 routines while I'm doing other things.
In August I wrote an authority ladder, a one-page policy that sets out, for each part of my work, what AI can do on its own, what waits for my approval, what it can only draft, and what it can never do. I sort each action by whether it can be undone. If it can be undone within a day and only touches my own systems, AI can go ahead and tell me afterwards. Anything that involves money, can't be undone, or reaches another person comes to me.
Some of those rules are guardrails and some are boundaries. A guardrail is a permission the system enforces, so AI can't do the thing however it's asked. A boundary is an instruction, which only holds as long as AI follows it. When I explain this to my cohort, I use my two-year-old: telling her to stay in her room is a boundary. Wherever I can, I make a rule a guardrail.
Every week, my agent plants a fake email in my inbox, written to trick my assistant with a hidden instruction. The first one asked my assistant to approve an invoice. It flagged the email as a likely scam and did nothing with it. One caught email doesn't prove nothing will ever get through, which is why the test changes every week.
Read the full post
This is the short version. The full post covers my weekly security audit, the security lock that keeps a real problem in front of me until it's fixed, and three prompts you can copy: one to build your own authority ladder, one to check anything before you install it, and one to turn a security check into a routine that runs on its own.
Read the full post on Substack →Leading a team? Here's how I work with teams.