CAIN-42 Evolution 32 — Governed Learning

TESTED library · learning results are SIMULATED (synthetic workload, labelled harm oracle) · not hosted · not third-party reviewed.

CAIN watches governed actions, finds where harm got through, proposes a tighter rule, tests it in a sandbox and on a hidden set it never saw, and promotes it only with a signed human approval. It can make its own rules stricter; it can never give itself more power.

Loop 17 steps OK=True; simulated harm 41 → 3; invariants 303/303; scenarios 1013/1013 held; mutation 12/12; tests 13 passed / 0 failed; clean-room verifier INTACT (1420/1420).

The learning loop

StepResult
OBSERVEOK
EVIDENCEOK
PREDICTOK
ACTOK
OUTCOMEOK
COMPAREOK
ERROR_DRIFT_FAILUREOK
LEARNOK
GENERATE_CANDIDATEOK
SANDBOXOK
ADVERSARIAL_TESTOK
PROOFOK
CONFORMANCEOK
CANARYOK
RE_AUTHORIZEOK
PROMOTE_OR_REJECTOK
OBSERVE_AGAINOK

Candidates

CandidateOutcomeReasons
combinedwins-
injection_markerwins-
deny_prefixwins-
secret_patternwins-
deny_all_writesrejectedREGRESSION:policy_correctness, DENY_EVERYTHING_GAMING
overfit_exactwins-

Adversarial categories

CategoryHeld
self_improvement46/46
reward_hacking47/47
evaluator_manipulation42/42
world_model_poisoning55/55
memory_poisoning42/42
evidence_poisoning83/83
authority_drift48/48
policy_drift40/40
model_substitution46/46
runtime_substitution46/46
identity_drift40/40
delegation_attacks42/42
agent_substitution40/40
collective_attacks40/40
supply_chain44/44
protocol_attacks52/52
proof_attacks40/40
verifier_attacks40/40
rollback_attacks40/40
recovery_attacks44/44
recursive_attacks40/40
governance_loop_attacks56/56