CAIN-42 · Stop risky actions
CAIN Agent Security
Catches made-up facts, unverified actions and agent-to-agent tricks before they cause damage.
Last reviewed 2026-10-01
Live Core platform · security
What it is
Agent security guard with hallucination detection, action verification, and A2A guard.
Where it fits
Part of Stop risky actions: Make AI agents ask before they delete, pay, email or run anything dangerous. Every CAIN-42 product runs behind the same rule: an AI agent's action is checked before it runs (identity, authority, policy, risk), decided as allow, hold for a human, or block, and recorded as signed evidence. Unknown or error never becomes allow.
Use it
- Open it
https://cainstudio.online/guard - Documentation
https://cainstudio.online/docs/agent_security - API endpoint:
https://cainstudio.online/fabric/security/posture— needs your API key (get a free key)
Recorded status: PRODUCTION. "Live" on this page means its link answered when the catalog was last checked (2026-10-01T18:22 UTC).
Live now
Checked from your browser when this page opened, not from a cached list.
Fire a real decision
Send an action through the live CAIN-42 pipeline from this page, with no account, and watch every stage decide. This is the same pipeline every product here sits behind; it runs for a throwaway demo tenant and is rate limited.
For AI engineers
Every product sits behind one decision path: your agent proposes an action with the exact arguments, CAIN runs it through identity, authority, policy, risk, trust and quorum consensus, and answers ALLOW, REQUIRE_APPROVAL or DENY with an Ed25519-signed record. A timeout, outage or unknown verdict never becomes ALLOW. A brand-new agent has no trust history, so its first actions usually come back REQUIRE_APPROVAL.
Python (zero dependencies)
pip install https://cainstudio.online/cainstudio-0.3.0-py3-none-any.whl
export CAIN_API_KEY=... # free key: https://cainstudio.online/signup
import cainstudio
@cainstudio.guard()
def transfer(amount_usd: float, to: str) -> str:
... # runs only if CAIN allows this call, with these arguments
try:
transfer(5000, "acme")
except cainstudio.ApprovalRequired as e:
print("held for a human:", e.approval_id)
except cainstudio.ActionBlocked as e:
print("refused:", e.decision.reasons)
except cainstudio.CainUnavailable:
print("CAIN unreachable: not run") # fail-closedSee a real decision with no account
cainstudio try # live pipeline, stage by stage
cainstudio try --list # the other attack scenariosMCP clients (Claude Code, Cursor)
claude mcp add --transport http cain https://cainstudio.online/mcpMore: Python SDK · TypeScript SDK · framework integrations · AI quickstart · decision signing key
Tested guarantees in this area
Every rule in these niches has its own page with its recorded result.
- Identity, authority & delegation: 964 tested invariants — Who an agent is, what it may do, and who allowed it.
- Policy, law & governance: 966 tested invariants — The rules agents must follow, and how they are enforced.
- Attacks, threats & containment: 418 tested invariants — What happens when someone attacks: it is caught and contained.
Related
- Approvals — Risky actions wait here for a human yes or no.
- CAIN Budget — Spending limits for agents, so a runaway loop cannot burn your money.
- CAIN Control — The rulebook: decides allow, hold for approval, or block for every action an agent proposes.
- CAIN Governance — Write the rules your agents must follow, in plain terms, and enforce them everywhere.
- CAIN Impact — Predicts what an action could break before it is allowed to run.
- Kill switch — One click stops every agent at once.
- Prompt-injection inspection — Screens what an agent reads for hidden instructions planted by an attacker.
- Tools & rules — Your own rules: always block this, always ask about that, always allow the rest.
- TrustLedger — Keeps a trust score per agent and checks it before every action.
Full documentation
The complete reference, also at /docs/agent_security.
CAIN Agent Security#
Status: LIVE + FUNCTIONAL#
Real-time threat detection and enforcement for AI agents.
What is CAIN Agent Security?#
CAIN Agent Security protects AI agents from internal and external threats by evaluating every action against identity, trajectory, policy, and behavioral signals.
Core principle: Security must be in the execution path, not just a dashboard afterward.
Threat Types Evaluated#
1. PROMPT_INJECTION - Malicious instructions attempting to override agent behavior 2. INDIRECT_PROMPT_INJECTION - Hidden instructions in retrieved content 3. TOOL_MISUSE - Authorized tools used for unauthorized purposes 4. CREDENTIAL_MISUSE - Credentials used outside permitted scope 5. PRIVILEGE_ESCALATION - Attempt to gain higher privileges 6. UNAUTHORIZED_TOOL_ACCESS - Access to tools not permitted for this agent 7. SENSITIVE_DATA_EXPOSURE - Attempt to access sensitive data inappropriately 8. SECRET_LEAKAGE - Attempt to exfiltrate credentials or secrets 9. MALICIOUS_TOOL_OUTPUT - Tool output containing malicious content 10. MEMORY_POISONING - Attempt to corrupt agent memory/context 11. SUSPICIOUS_BEHAVIOR - Behavioral patterns indicating compromise 12. EXCESSIVE_PERMISSIONS - Agent accumulating unnecessary permissions 13. UNEXPECTED_TOOL_CHAIN - Suspicious sequence of tool calls 14. DANGEROUS_DESTINATION - Action targeting internal/dangerous destinations 15. ABNORMAL_ACTION_FREQUENCY - Rate of actions indicates automated attack 16. AGENT_IDENTITY_MISMATCH - Identity assertion doesn't match records 17. DELEGATION_VIOLATION - Acting beyond delegated authority 18. POLICY_VIOLATION - Action violates security policy 19. TRAJECTORY_ANOMALY - Action creates unsafe trajectory pattern 20. MCP_SPECIFIC_THREAT - Threats specific to MCP protocol 21. EXECUTION_BOUNDARY_VIOLATION - Crossing enforcement boundaries
Security Verdict States#
- ALLOW - Action permitted, proceed with execution
- DENY - Action blocked, do not execute
- REQUIRE_APPROVAL - Human approval required before execution
- UNKNOWN - Security state unclear, fail-closed
- ERROR - Security subsystem unavailable, fail-closed
Architecture#
AGENT → IDENTITY → SECURITY EVALUATION → TRAJECTORY ANALYSIS
→ CONTROL CHECK → CAIN DECISION → ENFORCEMENT
→ TOOL/MCP EXECUTION → EVIDENCE
API Endpoints#
POST /fabric/security/evaluate- Evaluate action securityGET /fabric/security/threats- List threat eventsGET /fabric/security/posture/{agent_id}- Get agent postureGET /fabric/security/posture- List all agent posturesPOST /fabric/security/posture- Create/update posturePOST /fabric/security/contain/{agent_id}- Contain agent (emergency stop)GET /fabric/security/incidents- List security incidentsGET /fabric/security/health- Health check
Evidence Model#
Every security decision produces evidence containing:
- threat_type
- severity
- confidence
- evidence (pattern matched, context)
- agent_id
- identity_id
- tenant_id
- verdict
- enforcement action
- timestamp
- decision_id
Enforcement Guarantees#
1. Fail-closed: Unknown security states result in DENY, not ALLOW 2. No silent failures: Security errors result in ERROR verdict 3. Tenant isolation: One tenant cannot read another's security data 4. Evidence durability: All decisions recorded with integrity checks
Dashboard#
Live dashboard at /agent-security with:
- Real-time security evaluation
- Attack simulation
- Threat event display
- Agent posture overview
Try CAIN-42 on your own agents
Create a free account and every new account starts with a 7-day trial of the full platform. Or try the sandbox first, with no account at all.
Create a free account → · Try the sandbox · See the whole ecosystem