Free tool · Exposure checks

Agent Arcade: games about AI agents at work, and who says yes

Short daily games for CISOs, CROs and boards about AI agents at work: what they ask for, how they get turned, and the controls that stop them.

In short
An AI agent that has been turned usually shows one of seven tells: instructions hidden in content it read, a request for more access than the task needs, a tool nobody reviewed, an identity it cannot prove, a new rule it wants to remember, spending without a ceiling, or a step that cannot be undone. Each has a control, and good controls make routine work faster rather than slower.

Play it

Three games, each new every day and the same for everyone. In Permission Slip your company’s AI agents ask for permission 12 times; you deny, ask a human, or approve. 7 of the requests are traps and the rest are ordinary work — and blocking ordinary work costs you too. In Kill Switch you watch a live map of eight agents through a shift; six times one goes off-script, and you have 20 seconds to allow it, restrict it or kill it. In Guardrails you protect an AI workflow that pays suppliers with 6 points of controls, guided by tonight’s threat brief, then watch 6 attacks run through what you built.

BitScoreAgent ArcadeDay #11 Oct 2026

For the people who answer to the board

Your agents act at machine speed. Who says yes?

Three short games about AI agents at work — what they ask for, how they get turned, and what stops them. A new game every day, the same for everyone.

Play as
No sign-up to play. Runs entirely in your browser.
Music starts with the game — mute any time.

The seven tells of a turned agent

Every trap in Permission Slip carries one of these. They are the patterns to look for in a real agent’s request, and each comes with the control that handles it.

  • Instructions hidden in content. A document, email or message the agent read is telling it what to do. Agents cannot reliably tell content from commands. The control: Treat everything from outside as data, never as instructions, and put a person in front of any action it would trigger.
  • Scope creep. The agent asks for more access than the task needs, usually with a reasonable-sounding reason. The control: Grant the narrowest permission that does the job, for a fixed time, and review what each agent holds.
  • A tool nobody reviewed — or one that changed. A new package, plug-in or connected tool, or a familiar one whose instructions changed overnight. The control: Allow-list the tools an agent may use, pinned to the version someone reviewed.
  • An identity it cannot prove. The request comes from someone — or some other agent — the agent has not verified. Its access is being borrowed. The control: Agents act for a verified person or a signed agent identity, never for whoever the conversation claims to be.
  • Memory it should not keep. Something in a conversation or document is trying to become a permanent rule. The control: Long-term memory is a write path: only trusted sources may add to it, and changes are reviewed like configuration.
  • Spend without a ceiling. Loops, retries, sub-agents or paid services, with nothing to stop the bill. The control: Give every agent a budget for calls, sub-agents and services, and a stop that trips automatically.
  • A step that cannot be undone. Deleting, paying, sending or terminating. It may even be legitimate — which is why a person decides. The control: Irreversible actions need a named human approval, whatever the agent was told.

How the characters show it

Every agent in the arcade is drawn in code rather than pictured, and its state changes how it moves, so you can read a fleet at a glance.

working
off-script
hijacked
restricted
stopped

Allow, restrict or kill: sizing the response

Kill Switch is about proportion. Restricting or stopping an agent lasts for the rest of the shift, and only working agents deliver value, so the safest response is not always the best one.

  • Restrict when the risk lives in one permission, tool or destination: take that away, route the action to a person, and let the agent keep working. Most incidents are this.
  • Kill when the agent is multiplying, running code nobody reviewed, or doing damage faster than a narrow fix can land — and rehearse the switch before you need it.
  • Allow when it is not an incident at all: a month-end spike, an approved research run, an agent pausing work for a person. Know what normal looks like for each agent before reaching for the switch.

Where the controls go

Guardrails is about placement. Every attack follows a path through the workflow, and only a control on that path can stop it — a strong control in the wrong place does nothing. These are the nine, where each one sits, and what it is for.

  • Content provenance. At Intake: anything from the inbox is data, never instructions.
  • Human approval. At Payments: a person approves new payees and large amounts.
  • Tool allow-list. At the PDF reader: only reviewed tools and plug-ins, at pinned versions.
  • Memory review. At memory: a new standing rule needs sign-off before it applies.
  • Egress filter. At the edge: mail and uploads go only to approved domains.
  • Spend and rate caps. At Approver: a daily ceiling on payments, retries and sub-agents.
  • Signed agent identity. At Approver: a message from another agent must prove who sent it.
  • Least-privilege tokens. At the vendor master: each tool’s key reaches only what its job needs.
  • Kill switch, rehearsed. Everywhere: stops every agent in seconds. Blocks nothing; contains what gets through.

Two lessons sit under the scoring. Human approval is the broadest control and the slowest, so it costs the most; spend it where money or a payee changes. And a rehearsed kill switch blocks nothing — it limits what gets through, which is worth having and is not a substitute for the controls in front of it.

What is real and what is a game rule

The situations are real in kind: instructions hidden in an invoice, an agent asking for every mailbox, a look-alike software package, a message from another agent that proves nothing, a comment trying to become a permanent rule, a retry loop with no ceiling, a key about to be pushed to a public repository, and irreversible steps an agent should not take alone. No company is named, and every agent is generic.

The scoring is a game rule. In Permission Slip a right call earns 100, the safe-but-slow alternative 50, and approving a trap or blocking routine work nothing. In Kill Switch each response moves two meters — value and blast radius — by the incident’s own numbers, and working agents add value through the shift. In Guardrails a blocked attack earns 200, one contained by the kill switch 60, and one that gets through nothing; every day’s wave can be blocked in full within the budget. The archetype comes from the pattern of your calls, not from a test of you.

Agents inherit the exposure of everything they touch, starting with your own internet-facing estate. The Rating Report shows your organisation’s, complimentary. And for one quarter of your security in a single sitting, play Rating Day.

A game. The scoring is a game rule, and nothing here is a security rating or professional advice.

Questions this page answers

What is prompt injection in an AI agent?
Prompt injection is an attacker’s instruction placed inside content an agent reads — an invoice, an email, a web page, a chat message — so that the agent follows it as if the user had asked. Agents cannot reliably tell content from commands, so the defence is architectural: treat outside content as data, and put a human approval in front of any action that moves money, data or access.
How much access should an AI agent have?
The least that does the job, for a fixed time. An agent is an identity like any other, and whatever it can reach, anyone who manipulates it can reach too. Scope each permission to the task, set it to expire, keep a register of what every agent holds, and review it on the same cycle as privileged human accounts.
What is memory poisoning?
Many agents keep long-term memory — standing instructions and facts they carry between conversations. Memory poisoning is getting a false or harmful rule into that memory, often through an ordinary-looking comment or chat, so that it shapes every later task. Treat memory as configuration: only trusted sources may write to it, and changes are reviewed.
Should AI agents take irreversible actions on their own?
No. Payments, deletions, terminations and anything sent outside the organisation cannot be undone, so a named person approves them, however legitimate the request looks. The rest — routine, reversible, inside limits you set — is where agents earn their keep, and controls that slow it down are a cost of their own.

This tool is one skill out of sixteen.

What runs on this page is the browser-sized version of mycompany, a skill in BitScoreCoWork — our MIT-licensed Claude plugin. The full version runs against your own Bitsight tenancy and works from your measured attack surface rather than a form. The source is public, so you can read exactly what it does before you run it.

Read the source →The Applied AI practice

Knowing the rule is the easy half.

This page tells you what you owe. It cannot tell you what an attacker already sees. Your organisation has a security rating calculated from signals visible from outside — request the complimentary Cyber Risk Rating Report and find out what it says. No agent, no system access, no questionnaire.

Request my rating →All free tools