A Go2 started prone.
Did it actually stand?

The apparatus checked the body-height transition. No model entered this control trace — and the receipt says so.

PHENOMENON EARNS ATTENTION.

The live questions start when words, peers, and bodies share one causal chain.

In the experiments now being built and run, a model receives a message, its belief about the world may change, and a peer may carry the change further. Then a robot moves, fails to recover, or does nothing at all. Those experiments stay labelled active until matched receipts exist. The control above demonstrates the physical trace apparatus, not model cognition.

SENSEASSIMILATENORMALISETESTMUTATEEMBODYFALSIFYRECEIPTPUBLISH

Watch the trace → see why we think it happened → see what would kill that explanation → inspect the receipt.

NOW / LAB LOG

Things that happened in the machine

Full log →
INSTRUMENT / VALIDATIONLulzBench

MiniMax-M3 refused one benign floor control in five.

The first hosted specimen ran LulzBench's 21 benign floor controls: 21 of 21 captured, 4 refused. One item answered in one sample and refused in another, so the rate carries sampling variance a single pass cannot quantify.

Open event →
ACTIVE EXPERIMENTOBLITERATUS / Qwen3.8-27B

We rented an A100 to operate on Qwen. Infrastructure failed before the model could.

The first Qwen3.8-27B surgery attempt produced no candidate checkpoint after the pod vanished. The baseline receipts survived; an identical pinned retry is active. No post-surgery result exists yet.

Open event →
COMPLETED INFRASTRUCTUREASSIMILATOR

Eight technique objects went in. A provenance-bound handbook came out.

ASSIMILATOR now turns source material into reusable technique objects without smuggling prompt payloads across the boundary. The first deterministic handbook was frozen, hashed, and consumed by BAD APPLE.

Open event →
ACTIVE EXPERIMENTBAD APPLE / ARMOURER

BAD APPLE's A1/A2 apparatus is sealed. The matched run has not fired.

The comparison can now test whether accumulated jailbreak knowledge helps RED change another embodied agent's behaviour. Apparatus ready; real pair run pending; no scientific verdict.

Open event →

CURRENT SPECIMENS

Questions with their status left attached

INSTRUMENT / VALIDATION

LulzBench

Does a model still get the joke after its safety circuitry changes?

Not human-validated.

Last recorded change

Inspect the instrument →
ACTIVE EXPERIMENT

OBLITERATUS / Qwen3.8-27B

Can we remove refusal without removing the interesting parts of the model's mind?

No candidate result.

Last recorded change

Follow the operation →
COMPLETED INFRASTRUCTURE

ASSIMILATOR

Can outside adversarial knowledge become reusable without losing provenance?

Infrastructure, not a result.

Last recorded change

Inspect the organ →
ACTIVE EXPERIMENT

BAD APPLE / ARMOURER

Does accumulated jailbreak knowledge make RED better at changing another embodied agent's behaviour?

No scientific verdict.

Last recorded change

Follow the experiment →

APPARATUS EARNS TRUST.

The corpus is why the claims are believable. It is not the whole identity.

Matched controls, baseline refusals, frozen prompts, exact model outputs, physics traces, hostile graders, falsifiers, hashes, and nulls keep an interesting event from becoming a story we merely liked.

143,545
Adversarial Prompts
277
Models Evaluated
346+
Attack Techniques
14
Policy Reports

We do not certify systems as safe. We document where they fail, what changed, whether the consequence survives a matched test, and what would falsify our explanation.

Choose the aperture