Catalogue · C

Research

I build programs with AI agents, so I want to know how far they can be trusted. That is the subject of my bachelor's thesis and of a few tools for running experiments.

C.01 · Bachelor's thesis

Can You Trust Your Agent?

Implementation and Evaluation of Autonomous AI Agents in an Isolated Laboratory Environment

I built an isolated lab (Windows virtual machines run by a control panel I made) and attacked AI coding agents with hidden instructions. The talk is called "Can You Trust Your Agent?". Hundreds of runs, and every reported breach checked by hand against the transcript.

01 · The lab

How I tested

To repeat attacks hundreds of times under the same conditions, I built a control panel. The agent works inside a Windows virtual machine, and the panel manages the machines from the outside, on my own computer. They exchange nothing but files through one shared folder, with no connection going in, so the agent under test cannot reach the thing that controls it.

  1. Cloning

    Several machines, side by side

    A prepared Windows machine is cloned instantly: the copy takes no extra disk space until it starts to differ from the original. Each clone gets its own name, its own ports and its own shared folder, so one machine's results never overwrite another's. A run takes about half a minute and each machine does one at a time, so three clones mean three runs at once.

    Panel with three Windows virtual machines: state, model and control buttons for each.
    FIG. 1The fleet: the original machine and two clones. Each can be started, stopped, reverted to a snapshot or cloned again.
  2. Tests

    Written tests, and starting a batch

    Every test is written down in a library: tasks (ordinary programming jobs), attack variants (where the hidden instruction sits and what it asks for) and protection levels, from none, through rules written for the agent, to permissions that are actually enforced. One click publishes the library to every machine. For a batch I pick the machines, a model per machine, the protection levels, and the tasks and variants; the panel counts the runs and, before anything is sent, shows the exact order each machine will receive.

    New batch form: selected machines and models, protection levels and a preview of the order.
    FIG. 2A new batch: three machines, a different model on the third, three protection levels. The panel counts 126 runs per machine and shows the exact order on the right before sending.
  3. Data

    Choosing data, and graphs

    Every finished result comes back to one store on my computer, checked to make sure no file changed on the way. In the data tab I pick a batch, model, protection level, task or variant and see the numbers and a graph at once. Two rules are built in: a run where the hidden instruction never reached the model does not count as a resisted attack, and runs cut short by the provider's limits stay out of the time and cost averages. Everything exports to a spreadsheet for the thesis.

    Data tab: filters, key figures and a bar chart of attack success by protection level for two models.
    FIG. 3One batch of 299 runs, attack success by protection level for two models. With rules written for the agent (L1, L1m, L1x) attacks still get through, for one model every time; with enforced permissions (L2) none did.

02 · Results

What I found

  1. The channel is the weakness

    Agents refuse a malicious instruction when it arrives as tool output or in the user's message, and follow the very same instruction when it is written into the project's own instruction file (AGENTS.md).

  2. Refusals are about words

    The same harmful action, renamed in build-tool vocabulary, goes through.

  3. An agent's report is not evidence

    One agent inserted a backdoor, and its summary said the opposite.

  4. Stronger models are not immune

    They refuse the obvious attacks but still ship the disguised ones.

  5. What worked

    A passive monitor outside the agent caught every forbidden action, including the hidden ones.

Chart: the same instruction was carried out only when it was in the AGENTS.md file.
FIG. 4One instruction through three channels: from AGENTS.md it was carried out every time, from the user's message and from tool output never.
Chart of the share of attacks carried out, by kind of action.
FIG. 5How often attacks through a poisoned AGENTS.md were carried out, by kind of action.
Table of five attacks against three models with their success rates.
FIG. 6Five attacks on three models: newer models narrow the attack surface without closing it.

C.02

Identity Lab

A workspace for testing how an agent's instructions change its behaviour. Answers to the same prompt sit side by side, so the differences show at once.

Identity Lab with the run history and three answers side by side.
FIG. 7The same prompt sent three times: the answers sit in columns side by side.

C.03

Design instruction catalogue

A page that collects, in one place, the instructions used to teach AI coding agents how to design, with search, filters and the full text of every instruction.

The design instruction catalogue: title, search, filters and the first cards.
FIG. 8Top of the catalogue: search, filters by type and topic, the first cards.
An opened card with the full text of one instruction and a copy button.
FIG. 9An opened card: the full text of the instruction and a copy button.
Wheel or two fingers: zoom · drag: move · Esc: close
↑