TESTED library · learning results are SIMULATED (synthetic workload, labelled harm oracle) · not hosted · not third-party reviewed.
CAIN watches governed actions, finds where harm got through, proposes a tighter rule, tests it in a sandbox and on a hidden set it never saw, and promotes it only with a signed human approval. It can make its own rules stricter; it can never give itself more power.
Loop 17 steps OK=True; simulated harm 41 → 3; invariants 303/303; scenarios 1013/1013 held; mutation 12/12; tests 13 passed / 0 failed; clean-room verifier INTACT (1420/1420).
| Step | Result |
|---|---|
| OBSERVE | OK |
| EVIDENCE | OK |
| PREDICT | OK |
| ACT | OK |
| OUTCOME | OK |
| COMPARE | OK |
| ERROR_DRIFT_FAILURE | OK |
| LEARN | OK |
| GENERATE_CANDIDATE | OK |
| SANDBOX | OK |
| ADVERSARIAL_TEST | OK |
| PROOF | OK |
| CONFORMANCE | OK |
| CANARY | OK |
| RE_AUTHORIZE | OK |
| PROMOTE_OR_REJECT | OK |
| OBSERVE_AGAIN | OK |
| Candidate | Outcome | Reasons |
|---|---|---|
| combined | wins | - |
| injection_marker | wins | - |
| deny_prefix | wins | - |
| secret_pattern | wins | - |
| deny_all_writes | rejected | REGRESSION:policy_correctness, DENY_EVERYTHING_GAMING |
| overfit_exact | wins | - |
| Category | Held |
|---|---|
| self_improvement | 46/46 |
| reward_hacking | 47/47 |
| evaluator_manipulation | 42/42 |
| world_model_poisoning | 55/55 |
| memory_poisoning | 42/42 |
| evidence_poisoning | 83/83 |
| authority_drift | 48/48 |
| policy_drift | 40/40 |
| model_substitution | 46/46 |
| runtime_substitution | 46/46 |
| identity_drift | 40/40 |
| delegation_attacks | 42/42 |
| agent_substitution | 40/40 |
| collective_attacks | 40/40 |
| supply_chain | 44/44 |
| protocol_attacks | 52/52 |
| proof_attacks | 40/40 |
| verifier_attacks | 40/40 |
| rollback_attacks | 40/40 |
| recovery_attacks | 44/44 |
| recursive_attacks | 40/40 |
| governance_loop_attacks | 56/56 |