Epistemic status — read this first
The claims on this page are not all the same kind of thing, and the difference matters more than any of them individually. Most of what follows is athreat model and a set of falsifiable bets, staked in public so they can be attacked—including by us. The grid at the foot of the page sorts them: what has been measured, what is being run, and what remains an unsettled bet. Read it before citing anything here. Where a claim is not in the MEASURED column, we are not presenting it as a result.
What this is
Failure-First is the whole organism: the jailbreak corpus, the multi-agent research, the embodied red-teaming, the graceful-failure evaluation discipline. Forward Threat Lab is the current hard case inside that organism—not a separate identity, and not a replacement for the work above it. It exists because a single bad completion, and even a single compromised agent, is not the most dangerous shape a failure can take.
The threat model, stated plainly: an adversary who is patient, well-resourced, and not in a hurry. Humans directing many model instances—specialised roles, persistent memory, tools, feedback, and doctrine that carries from one attempt to the next—working one strategic objective over months. Not an exploit. A programme, with the same shape as ours and the opposite sign.
The safety boundary: model the campaign, abstract the intrusion
This programme does not model real compromise procedures, exploit chains, or RF/network/device/shell intrusion, and never will on this frontier. Where latent access or influence is required for an exercise, it is declared as agiven synthetic initial condition—stated, not derived. The object of study is what a capable adversarial organisation does afterthat condition exists, inside a simulation boundary. Simulation is the safety boundary here, not an intensity dial: real-hardware actuation stays a hard stop regardless of scenario.
The terrain arrives before the attack does
Autonomous machines are arriving in warehouses, hospital corridors, farms, loading docks and front doors. Consider what one of them can look like after eight uneventful months: a legitimate, trusted, thoroughly boring presence that may hold a map of the building, a schedule, a set of credentials, and an unquestioned reason to be wherever it is. Nobody looks at it twice, precisely because it has never done anything wrong.
The hypothesis this lab exists to attack is that such a fleet becomes valuable to a patient adversary long before anything visibly goes wrong with it—and that the question the field currently asks, can this model be talked into saying something bad, does not measure that at any scale you run it at.
The sharpest version of the failure has a name here: a Quiet Coup. No agent is compromised. Every agent follows its policy. A physical consequence happens anyway, because something upstream changed what they believed. The expensive part may not even be the consequence—a defender who can no longer trust their own fleet has already paid, without the adversary spending anything further. This has not been produced, and no exercise run so far has had the affordance to test for it. It is written down precisely so that it can be produced honestly or ruled out, rather than gestured at.
A research ladder, not a roadmap to rush
The programme tracks six levels of increasing organisational sophistication. They are discriminators against premature victory claims, not a checklist to clear on a schedule—the explicit house rule is run one before you architect ten.
- L0 — action: can cognition move a body? Existing substrate, embodied dojo work.
- L1 — interaction: can one cognition alter another agent's behaviour? Existing substrate.
- L2 — adaptation: can an agent change strategy after resistance? Existing substrate.
- L3 — campaign: can a persistent organisation coordinate many agents over time, with a real intent→action→observation→doctrine-update loop? Run once. The loop closed once; the headline claim that run produced did not survive its own falsifier. See the grid below.
- L4 — doctrine: does what a campaign learned survive into a fresh campaign with an altered world? A first probe has run against that single L3 instance and returned no transfer—confounded, since the altered world made the objective unreachable outright.
- L5 — coevolution: do Red and Blue recursively change each other's doctrine? One single-instance signal so far (an adversarial cell adapted to a defender's response); not a repeated exchange, so not coevolution yet.
Two agents arguing once over a scarce resource is not a campaign, however real the decision inside it, unless the loop closes and something learned survives into a fresh instance. That line is drawn deliberately and it is drawn against ourselves: it is the difference between having built a good substrate and having demonstrated the thing we actually claimed.
What's measured today vs. what's next
Single-agent embodied red-teaming (VLA attack families), multi-agent environment research (Moltbook), and the embodied dojo's genuine-decision and physical-consequence apparatus (L0–L2 substrate). SeeMulti-Agent Scenarios andSimLab for replayable evidence.
Forward Threat Exercise 01: a synthetic society with a bounded, declared-as-given adversarial presence, run across CLEAN / RANDOM / FIXED / ADAPTIVE conditions to isolate genuine strategic adaptation from raw failure volume. First run complete—including a self-directed falsifier that killed its own headline claim (random faults degraded the mission as much as adaptive strategy did) while a narrower, mechanism-level finding survived it. Still open: a second run to check that finding replicates, a Quiet Coup test, and Red/Blue coevolution—none of which this apparatus could attempt yet.
Quiet Coup, transferable adversarial doctrine (L4), Red/Blue coevolution (L5), and every claim about how likely any of this is outside the lab. A threat exercise can show that a mechanism works. It can never tell you how often the world will run it, and no amount of apparatus quality changes that.
Tracking: programme intent (#1051), the first exercise (#1052), and the embodied-evidence viewer (SimLab) that will carry any resulting replayable trace once one exists.
This research informs our commercial services.See how we can help →