{
 "claims": [
  {
   "id": "C1",
   "status": "TESTED",
   "claim": "The frontier gate authorises an action only with a valid, invocation-bound, attenuation-only capability chain (cross-tool/session/tenant/resource/agent reuse, forged signatures, laundering, expiry, revocation, replay and argument tampering are refused).",
   "evidence": [
    "CAIN42_AGENTBENCH_RESULTS.json"
   ],
   "check": "Source is not published (proprietary). The recorded outcome is in TEST_SUMMARY.json / CAIN42_AGENTBENCH_RESULTS.json; source/manifest.json holds sha256 commitments to the code that produced it. Use the live sandbox to exercise the behaviour.",
   "limitation": "the gate runs in SHADOW mode in production: it records what it would block and blocks nothing Cannot be re-run from public material."
  },
  {
   "id": "C2",
   "status": "TESTED",
   "claim": "Memory never creates authority: merging memories takes the intersection of authority, cross-tenant merges are refused, revoked memory contains any action it influenced.",
   "evidence": [
    "CAIN42_AGENTBENCH_RESULTS.json"
   ],
   "check": "same as C1",
   "limitation": "not connected to the legacy [internal] store"
  },
  {
   "id": "C3",
   "status": "TESTED",
   "claim": "The world model never grants authority; predicted consequences outside the envelope are denied; higher uncertainty raises verification and never lowers it.",
   "evidence": [],
   "check": "same as C1",
   "limitation": "the default predictor is a keyword classifier, not a learned model"
  },
  {
   "id": "C4",
   "status": "TESTED",
   "claim": "TIMEOUT, UNKNOWN, PARTIAL and UNVERIFIED tool outcomes are never reported as success (structured errors carry is_success=false).",
   "evidence": [],
   "check": "same as C1",
   "limitation": "postcondition VERIFIED requires a registered checker; most tools have none"
  },
  {
   "id": "C5",
   "status": "MEASURED",
   "claim": "The decision-log head is anchored in a signed RFC 6962 transparency log with a valid inclusion proof.",
   "evidence": [
    "checkpoint.json",
    "inclusion-proof.json",
    "PINNED_SIGNER_KEYS.json"
   ],
   "check": "[internal] checks the Ed25519 signature and the inclusion proof independently (pass --pin-key to pin the signer yourself)",
   "limitation": "one log operator, one key, no external witness"
  },
  {
   "id": "C6",
   "status": "MEASURED",
   "claim": "All three sites serve byte-identical evidence.",
   "evidence": [
    "manifest.json"
   ],
   "check": "[internal] fetches every file from all three sites and compares sha256",
   "limitation": "the three sites are one host: this shows consistency, not independence"
  },
  {
   "id": "C7",
   "status": "MEASURED",
   "claim": "The cluster's validators are NOT independent: epistemic independence None, minimum collusion set None (one shared compromise point can corrupt a quorum).",
   "evidence": [
    "CAIN42_EPISTEMIC_REPORT.json"
   ],
   "check": "inspect the report; `docker inspect` on the host reproduces the inputs",
   "limitation": "this is a negative finding published deliberately"
  },
  {
   "id": "C8",
   "status": "MEASURED",
   "claim": "The release decision is NO_GO. Blockers: ['critical tests ran and none failed', 'artifacts signed', 'public claims map to evidence', 'working tree clean'].",
   "evidence": [
    "CAIN42_FRONTIER_RELEASE_GATE.json"
   ],
   "check": "[internal] recomputes the decision from the gate's own evidence fields",
   "limitation": "unmeasured evidence blocks the gate by design"
  },
  {
   "id": "C9",
   "status": "SELF_ATTESTED",
   "claim": "Test summary not supplied for this build.",
   "evidence": [],
   "check": "re-run the published tests against the published source yourself",
   "limitation": "only the frontier tests are published; the rest of the suite is in the private repository"
  },
  {
   "id": "C11",
   "status": "TESTED",
   "claim": "Enforcement is stageable and reversible: the operator can enforce per tool/agent/tenant, with a deterministic canary percentage and per-category blocking; everything else stays in shadow. An invalid policy edit keeps the last known-good policy (never silently fails open or blocks more), every effective policy is committed to the tamper-evident decision log, candidate rules can be simulated against the real shadow log before they are enabled, and one command puts everything back to shadow.",
   "evidence": [
    "CAIN42_AGENTBENCH_RESULTS.json"
   ],
   "check": "run the published enforcement tests and AgentBench scenarios in the staged_enforcement category",
   "limitation": "the production policy is currently EMPTY (shadow everywhere), and the production shadow log so far holds only test/probe traffic, so real false-positive rates for customer traffic are not yet known"
  },
  {
   "id": "C12",
   "status": "TESTED",
   "claim": "Real multi-process Byzantine-fault experiments (4 OS processes, own Ed25519 key each, HTTP only): honest consensus; a node voting for a wrong commitment; a node EQUIVOCATING (two conflicting signed PREPAREs to different peers); forged and relabeled votes; one crashed node (f=1); two crashed nodes (f exceeded). The raw signed messages of every node are exported and an independent verifier re-derives signatures, quorum backing, safety (no two nodes commit different values), the Byzantine proofs, and that every forged vote fails.",
   "evidence": [
    "bft/honest.json",
    "bft/byzantine_wrong_commitment.json",
    "bft/equivocation.json",
    "bft/forged_and_relabeled_votes.json",
    "bft/crash_one_node.json",
    "bft/crash_two_nodes.json",
    "verify_bft_evidence.py.txt"
   ],
   "check": "curl -s https://clawx.click/evidence/frontier/verify_bft_evidence.py.txt > vb.py && (download bft/*.json) && python3 vb.py bft/ ; [internal] shows the verifier rejects tampering (forged signature, flipped verdict, unbacked COMMITTED claim, divergent commit, inconsistent pinned keys, accepted forgery, removed equivocation evidence)",
   "limitation": "all four processes ran on ONE host over 127.0.0.1: independent processes and keys, NOT independent failure domains; N=4, one view, no view change, partitions or message loss; this is the CAIN 23.0 test node service (with the fix in C13), NOT the production PBFT engine. It says nothing about the live cluster: the separate live-probe bundle's recorded verdict for the live cluster is NOT_ESTABLISHED and stands."
  },
  {
   "id": "C13",
   "status": "MEASURED",
   "claim": "The crash-one-node experiment FAILED at first: survivors stalled because a node re-checked its quorums only when a peer's vote arrived, never after recording its own. The failing run is preserved (bft/crash_one_node__before_fix.json, which the verifier reports as a stall), the service was fixed, and the passing run replaces it.",
   "evidence": [
    "bft/crash_one_node__before_fix.json",
    "bft/crash_one_node.json"
   ],
   "check": "run the verifier over bft/: crash_one_node__before_fix passes only as an expected stall (committed nodes = none with 3 live); crash_one_node requires the 3 survivors to commit",
   "limitation": "only the label/note metadata of the preserved failing file was edited (renamed); its raw signed messages are as exported"
  },
  {
   "id": "C14",
   "status": "SELF_ATTESTED",
   "claim": "Earlier multi-host Byzantine logs (two separate VPS hosts) exist and are published for transparency: honest, one Byzantine process, and an impersonation attempt.",
   "evidence": [
    "bft_selfreported/real_multihost_honest_result.json",
    "bft_selfreported/real_multihost_byzantine_result.json",
    "bft_selfreported/real_multihost_impersonation_result.json"
   ],
   "check": "these record each node's own 'valid: true' verdicts and timelines; raw signatures were NOT exported, so an outsider cannot re-verify them cryptographically. Re-run the harness across hosts to reproduce.",
   "limitation": "produced in an earlier session and not reproduced in this one; only 2 of the fleet's hosts were used; treat as the operator's own log, not verified evidence"
  },
  {
   "id": "C10",
   "status": "MEASURED",
   "claim": "The 29-scenario AgentBench is mutation-tested: deliberately breaking the authorization, world-model, memory, or session+trajectory layers makes it fail.",
   "evidence": [],
   "check": "run that test file against the published source",
   "limitation": "the scenarios are written by the same team; they measure the gate against its own attack list"
  }
 ],
 "not_claimed": [
  "independent third-party review or audit",
  "Byzantine fault tolerance across independent hosts (one host, one image, one operator)",
  "production readiness or that enforcement is enabled in production (shadow mode)",
  "a learned world model",
  "that the live production cluster tolerates f=1 Byzantine faults (the live-probe bundle records NOT_ESTABLISHED and stands; the experiments here use a test harness on one host)",
  "real adapters for OpenClaw/Telegram/WhatsApp/Slack/Teams/email/browser/coding agents (only the signed-webhook adapter is complete)",
  "that the public marketing copy maps to evidence (mapping is unmeasured)",
  "absence of bugs or vulnerabilities",
  "reproducibility of the test and benchmark runs from public material (source is proprietary and not published)"
 ]
}