Interactive tool

Giving AI more autonomy and more access also gives attackers a new way in.

Every system you connect to an AI model — its inbox, its documents, its tools, its plugins — becomes something an attacker can try to talk to. This isn't about AI being malicious. It's about AI being persuadable: it reads instructions wherever they appear, including places you didn't intend as instructions. This tool walks a board-level audience through the three risk categories that matter most, with one toy example you can click through yourself.

Conceptual & defensive — no exploit code, just the patterns to recognize
01 — The core idea

Autonomy and access are the same lever, in two directions.

A chatbot that only answers questions from a fixed knowledge base has a small attack surface — there's not much for an attacker to do even if they fool it. An agent that can browse the web, read your email, and take actions on your behalf has a large one. The more an AI system can do and the more it can see, the more valuable it becomes — and the more damage a successful manipulation can cause.

Autonomy

What can it do without asking?

Send emails, run code, move money, edit files, call other tools — every action an agent can take on its own is an action an attacker might be able to trigger indirectly.

Access

What can it see?

Customer records, internal documents, credentials, your contacts list, prior conversation history — everything in its context window is something it could be tricked into repeating or acting on.

The attack surface is the overlap. An attacker rarely needs to "hack" the model itself. They just need to get malicious instructions in front of it, somewhere it has enough autonomy or access to make those instructions worth following.
02 — Try it yourself

Spot the hidden instruction.

Below are three pieces of content an AI agent might be asked to read on your behalf — a webpage, an email, and a document. Each one contains an instruction hidden from a human reader but perfectly readable by a model. Read the visible text, then click where you think something is hiding. These are illustrative toy examples, not real attack payloads.

Click anywhere in the content above where you think text might be hidden.

What happens if an agent reads this

Click "Reveal" to see how an agent processing this content could be hijacked by the hidden text.

03 — The three risk categories

Almost every AI security headline is one of these three.

You don't need to track a long list of exotic attacks. Nearly everything reported about AI security so far reduces to one of three patterns — each with a straightforward, non-technical mitigation a board can ask about directly.

04 — Check your understanding

Five questions on the three risks.

A quick comprehension check — no trick questions, just the core distinctions that matter when someone briefs you on this topic for real.

05 — Key takeaways

What to actually ask your team.

If you remember nothing else from this page, remember these four questions — they apply to almost any AI deployment you'll be asked to approve.

06 — Go deeper

Security review belongs before deployment, not after.

The right time to ask "what can this agent access, and what happens if it's fed the wrong instructions" is during design — not after an incident. We help teams scope AI deployments with these risks in mind from day one.

Talk to us about AI security review