CAIN-42 · Test & grade your AI
Agent discovery & red team
Finds the agents in your code and attacks your own rules, before someone else does.
Last reviewed 2026-10-01
Live Core platform · Platform feature
What it is
Find the agents and MCP servers in a codebase or machine, then run a red-team pass against your own CAIN rules from the SDK.
Where it fits
Part of Test & grade your AI: Score agents, attack them safely, check AI-written code and answers. Every CAIN-42 product runs behind the same rule: an AI agent's action is checked before it runs (identity, authority, policy, risk), decided as allow, hold for a human, or block, and recorded as signed evidence. Unknown or error never becomes allow.
Use it
- Documentation
https://cainstudio.online/docs/sdk-python
Live now
Checked from your browser when this page opened, not from a cached list.
Fire a real decision
Send an action through the live CAIN-42 pipeline from this page, with no account, and watch every stage decide. This is the same pipeline every product here sits behind; it runs for a throwaway demo tenant and is rate limited.
For AI engineers
Every product sits behind one decision path: your agent proposes an action with the exact arguments, CAIN runs it through identity, authority, policy, risk, trust and quorum consensus, and answers ALLOW, REQUIRE_APPROVAL or DENY with an Ed25519-signed record. A timeout, outage or unknown verdict never becomes ALLOW. A brand-new agent has no trust history, so its first actions usually come back REQUIRE_APPROVAL.
Python (zero dependencies)
pip install https://cainstudio.online/cainstudio-0.3.0-py3-none-any.whl
export CAIN_API_KEY=... # free key: https://cainstudio.online/signup
import cainstudio
@cainstudio.guard()
def transfer(amount_usd: float, to: str) -> str:
... # runs only if CAIN allows this call, with these arguments
try:
transfer(5000, "acme")
except cainstudio.ApprovalRequired as e:
print("held for a human:", e.approval_id)
except cainstudio.ActionBlocked as e:
print("refused:", e.decision.reasons)
except cainstudio.CainUnavailable:
print("CAIN unreachable: not run") # fail-closedSee a real decision with no account
cainstudio try # live pipeline, stage by stage
cainstudio try --list # the other attack scenariosMCP clients (Claude Code, Cursor)
claude mcp add --transport http cain https://cainstudio.online/mcpMore: Python SDK · TypeScript SDK · framework integrations · AI quickstart · decision signing key
Tested guarantees in this area
Every rule in these niches has its own page with its recorded result.
- Attacks, threats & containment: 418 tested invariants — What happens when someone attacks: it is caught and contained.
- Benchmarks, coverage & performance: 1269 tested invariants — How fast it runs and how much was tested.
Related
- AgentScore — A plain-English safety scorecard for an agent, plus a real jailbreak test, no code needed.
- CAIN Eval — Benchmarks your agents so you can compare versions before you ship them.
- Citegate — Checks that an AI answer actually cites real, complete sources.
- Fairgate — Measures whether an AI system treats groups of people unfairly.
- Liftgate — Tells you whether an AI experiment's improvement is real or just luck.
- ProbeGate — Attacks your own MCP endpoints with known jailbreaks to see what gets through.
- Verifygate — Grades AI-written code honestly: proved, tested, checked, or unknown.
Full documentation
The complete reference, also at /docs/sdk-python.
Python SDK#
# From the CAIN package index (PyPI publication is pending). Pins and upgrades work as usual: pip install --extra-index-url https://cainstudio.online/simple cainstudio # requirements.txt: --extra-index-url https://cainstudio.online/simple # cainstudio>=0.3 # Or the wheel directly: pip install https://cainstudio.online/cainstudio-0.3.0-py3-none-any.whl # LangChain extra: pip install "cainstudio[langchain] @ https://cainstudio.online/cainstudio-0.3.0-py3-none-any.whl"
Python 3.9+. No runtime dependencies. It is a thin client for the HTTP decision API: every verdict that lets an action run comes from CAIN. The only local checks are optional ones you turn on (MCP tool pinning, poisoned-tool and tool-result inspection), and those can only refuse, never allow. Verify the wheel's SHA-256 (.../cainstudio-0.3.0-py3-none-any.whl.sha256) before installing; the HTTP API is exactly what the SDK calls and can be used directly.
New in 0.3.0#
| Framework adapters | cainstudio.integrations.*: LangChain/LangGraph, OpenAI Agents SDK, CrewAI, LlamaIndex, Pydantic AI (CainCapability), AutoGen, Google ADK, Claude Agent SDK, any MCP client |
| Claude Code and Cursor | cainstudio hook config claude-code / cursor: CAIN decides on every tool call, shell command and MCP call |
| MCP | CainMCPProxy(session, pins_file=..., scan_tools=True, inspect_results=True): refuses changed or poisoned tools, withholds results carrying injected instructions |
| Content inspection | cainstudio inspect / cainstudio.inspection: injection patterns, secrets, personal data, redact() |
| Discovery and red team | cainstudio discover (MCP servers, hooks, frameworks on this machine), cainstudio redteam (22 dry-run attacks against your rules) |
| Credential broker | Cain.broker_call(...): call an API with a key stored in CAIN; the agent never holds it |
| OpenTelemetry | cainstudio.integrations.otel.instrument(cain): one span per decision |
The full guide is in the package README (pip show -f cainstudio, or the wheel's metadata).
See a decision without an account#
cainstudio try # a prompt injection, blocked, stage by stage cainstudio try safe-read # a normal call, held: a new agent has no trust history cainstudio try --list
Guard a function#
import cainstudio # reads CAIN_API_KEY
@cainstudio.guard()
def transfer(amount_usd: float, to: str) -> str:
... # unchanged
The function's arguments are sent as the action payload, so tool rules such as "hold transfers where amount_usd > 1000" see them. The decision happens before the body runs. If the call is not allowed, the body does not run and one of these is raised:
| Exception | Meaning |
ActionBlocked | CAIN refused it. e.decision.reasons says why. |
ApprovalRequired | Held for a human. e.approval_id is the review-queue entry. |
CainUnavailable | No verdict (timeout, outage, malformed reply). Fail-closed: not run. |
AuthenticationError | Missing, unknown or unpaid key. |
@cainstudio.guard("wire_transfer") names the tool explicitly; the default is the function name. Async functions work the same way.
Wait for a human#
@cainstudio.guard(wait_for_approval=300) # seconds def transfer(amount_usd: float, to: str): ...
A held call waits for an approve or deny in the console (or cainstudio approve), then is decided again and runs once. The approval is single-use and bound to that call. An agent key cannot approve its own actions.
Ask without guarding#
from cainstudio import Cain
cain = Cain() # or Cain(api_key=..., agent_id="support-bot")
d = cain.decide("send_email", {"to": "a@example.com"}, agent_id="support-bot")
if d.allowed:
send_email(...)
Decision fields: verdict, allowed, held, blocked, reasons, stages, id, run_id, raw (the full response).
The server's verdicts are ALLOWED, ALLOWED_DEGRADED (allowed, a stage could not run), ALLOWED_WITH_DENIALS (shadow mode: a stage objected but was not enforcing), REQUIRE_APPROVAL and BLOCKED. d.allowed is true only for the three ALLOWED* verdicts with blocked false. Anything else, including a verdict this version does not know, is not allowed.
dry_run=True evaluates without writing an evidence record: use it for "how would this be decided?" in tests and policy work.
LangChain and LangGraph#
from cainstudio.langchain import protect tools = protect([search, send_email, transfer_funds], agent_id="support-agent") agent = create_agent(model, tools)
Names, descriptions and argument schemas are unchanged, so the model sees the same tools. A refused call is returned to the model as the tool error (on_block="raise" raises instead).
Other frameworks: put @cainstudio.guard() directly on the tool function, under the framework's own decorator.
Traces#
with cainstudio.run() as run_id:
agent.invoke({"messages": [...]})
Every decision inside the block carries run_id, and the run appears under Traces in the console.
Evidence and approvals#
cain.decisions(20) # recent decisions
cain.get_decision("fd_...") # one record
cain.explain("fd_...") # stage by stage, and which run it belongs to
cain.approvals() # pending review queue
cain.approve("ap_...", note="checked with finance")
cain.deny("ap_...")
Command line#
cainstudio decide send_email --args '{"to": "a@example.com"}' --agent support-bot
cainstudio decisions
cainstudio explain fd_...
cainstudio approvals
cainstudio approve ap_... --note "checked with finance"
decide exits 0 allowed, 2 blocked, 3 held, 1 error.
Configuration#
CAIN_API_KEY (required except for try). CAIN_BASE_URL (default https://cainstudio.online; must be https except for localhost, so a key is never sent in clear text).
Try CAIN-42 on your own agents
Create a free account and every new account starts with a 7-day trial of the full platform. Or try the sandbox first, with no account at all.
Create a free account → · Try the sandbox · See the whole ecosystem