In-process guardrails for AI agents. Intercept an agent's tool and MCP calls inside its own process, and block or modify a bad action before it runs, not after the damage is done.
Unlike a network proxy, interlock sits inside the process and sees the real call with real arguments. Unlike prompt filtering, it acts at the point where the agent actually does something. And for the actions a guard cannot see — a shell-out, a spawned subprocess — BRACE confines the process itself using the same policy.
flowchart LR
agent([agent]) --> toolcall["tool / MCP call"]
subgraph inproc["in-process: what interlock sees"]
direction LR
interceptors["interceptors"] --> engine["policy engine"]
engine -->|"allow"| ok["allow"]
engine -->|"modify"| ok
engine -->|"block"| no["block"]
end
subgraph outproc["out-of-process: what it cannot see"]
direction LR
brace["BRACE sandbox"] --> backends["bubblewrap · Seatbelt · soft"]
end
toolcall --> interceptors
interceptors --> brace
engine --> brace
M0 core is shipped and stable: tri-state verdicts, a fail-closed rule engine,
the @guard decorator, and in-process MCP + LangChain interception.
This tree goes well beyond M0. Also shipped: hierarchical policy scopes with
delegation chains (and an optional native accelerator), 105 detectors, signed
execution receipts with an append-only sink, tool-result sanitization, MCP
capability leases with server attestation and tool-surface drift detection, the
BRACE confinement layer with its escape-resistance matrix, and a Node.js client
(js/) that guards MCP tool calls, native fs/exec/fetch
effects, and model-SDK tools against the same engine over a sidecar socket. Still
open:
declarative YAML/JSON policies, automatic model-SDK hooks, Verdict.ESCALATE,
and more — see ROADMAP.md for the full list, each item with its
reasoning and a falsifiable done criterion.
- Zero runtime dependencies.
pip install interlock-guardpulls nothing. The MCP adapter duck-types onsession.call_tooland never imports the MCP SDK; LangChain is imported lazily, only when you call for it. - Fail-closed by default. A rule that raises denies. A confinement profile that grants nothing gets no network and no writable paths.
- Tri-state, not binary.
allow/block/modify. A modify rewrites the arguments and proceeds — cap an amount, redact PII — instead of failing shut.
pip install interlock-guard # core: no dependencies
pip install "interlock-guard[langchain]" # extra, only if you guard LangChain toolsRequires Python 3.9+. The native accelerator in rust/ is
optional and never required:
pip install maturin && cd rust && maturin develop --release
python -c "from interlock.policy import SCOPE_BACKEND; print(SCOPE_BACKEND)" # -> rustimport subprocess
import interlock
from interlock import Blocked, deny_when, guard, span
@guard(policy_id="shell.exec.v1")
def run_shell(cmd: str) -> str:
return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout
def looks_destructive(event) -> bool:
cmd = str(event.args.get("cmd", ""))
return any(t in cmd for t in ("rm -rf", "mkfs", "dd if=", "> /dev/sd"))
interlock.install(
rules=[deny_when(looks_destructive, reason="destructive command", policy_id="shell.exec.v1")]
)
with span(principal="my-agent"):
run_shell("echo hi") # allowed
run_shell("rm -rf /important") # raises Blocked; the command never runsRun the bundled demo: python examples/block_shell_tool.py
A Node.js agent can be guarded by the same engine — not a port of it. The
policy stays Python; a interlock-sidecar process owns the one engine and answers
a narrow RPC over a unix socket, and the interlock-guard npm package speaks that
wire so a JS MCP client is blocked or rewritten before a tool call reaches the
server. There is no second implementation to drift.
The daemon requirement, before anything else: this is fail-closed. With no
daemon running — or one that dies mid-session — every guarded call is denied
with a BlockedError. An agent with no daemon does not act at all, and monitor
mode is not exempt: an unreachable daemon is an infrastructure failure, not a
policy verdict. A Python process must be deployed beside the JS agent.
interlock-sidecar --rules my_service.policies:RULES --socket /tmp/il-guard.sockimport { RemoteEngine, guardMcpSession } from "interlock-guard";
const engine = new RemoteEngine("/tmp/il-guard.sock", { timeout: 5 });
guardMcpSession(client, { engine }); // client: any object with callTool({name, arguments})Also shipped in js/: installInterceptors() guards the
process's fs, child_process.exec/execFile, and fetch before the effect
(fs.readFile, child_process.exec, http.fetch), and guardToolMap /
guardToolRunner guard a model SDK's tool seam without importing any SDK. The
interceptors report what they cannot reach — ESM named imports, the *Sync
variants, the handle-returning openers, and result-phase rules on natives — via
installInterceptors().uncovered. evaluate is async — Node has no synchronous
unix-socket client — so a synchronous JS tool call cannot be guarded. See
js/README.md for the limits and the round-trip number.
SensorEvent — one observed action as the guard sees it before the effect:
action (tool/function/MCP name), args (the arguments policy may inspect),
principal, span_id, ts, parent_principal (the delegator, when the action
runs on someone else's behalf), and phase — "call" for an action about to
run, "result" for the observation path.
Verdict / Decision — ALLOW (0), BLOCK (1), MODIFY (2). The integer
values match the native engine's out-param contract, so a Rust core can drop in
behind the same enum. A Decision carries reason, policy_id, attributed_to
(the argument that triggered the verdict — a name, never the value, so the
decision is explainable without echoing the payload), and a rewrite payload —
modified_args for a modified call, modified_result for a sanitized result,
each populated only for its own direction.
Blocked — the exception raised when an enforcing guard denies an action. It
carries the decision and the action, and is raised before the wrapped
function executes, so the side effect never happens.
Run identity — span(principal=...) scopes a span id (and optional
principal) to a block via contextvars, so it survives await boundaries and
every decision is attributable to one agent run. new_span(),
current_span(), current_principal(), and set_principal() are also public.
Fail-closed — deny_when treats a predicate that raises as a match and
denies, rather than silently allowing.
Every interceptor funnels into one transport-agnostic core,
enforce_tool_call(action, arguments), so a rule you write once enforces
identically whether the agent speaks MCP, LangChain, or a bespoke callable.
from interlock import guard, monitor
@guard(policy_id="payments.transfer.v1") # blocking (default): raises on deny
def transfer(amount: int, to: str) -> str: ...
@monitor() # observe-only: records, never blocks
def risky(x): ...Works on sync and async callables. The decorator binds the call signature
(defaults applied) into the event's args.
Wrap one session once on creation and every outbound call_tool is checked
before it reaches the server. No gateway, no proxy, no per-tool decorator.
from interlock import guard_mcp_session
session = guard_mcp_session(await make_client_session())
await session.call_tool("write_file", {"path": "/etc/passwd"}) # Blocked, never sentDependency-free on purpose: it duck-types on .call_tool, so it wraps any client
exposing that method and stays testable with a fake session.
from interlock import guard_langchain_tools
safe_tools = guard_langchain_tools(agent_tools, principal="agent-7")Returns drop-in replacements with the same name, description, and argument schema. The original tool runs untouched once its entry is vetted.
Use this directly from your own adapter:
from interlock import enforce_tool_call
args = enforce_tool_call("http_post", {"host": "evil.example", "body": "..."}) # raises BlockedReturns the arguments to actually use (rewritten on MODIFY), or raises
Blocked. Pass enforcement="monitor" to record the would-be decision instead.
Indirect prompt injection arrives in the tool result, not the call — the agent
reads the poison and acts on it. enforce_tool_result is the same core pointing
the other way: it takes what a tool returned, and returns the result the agent
should actually see (or raises Blocked).
from interlock import enforce_tool_result
result = enforce_tool_result("fetch_url", tool_output) # sanitized, not rawThe pair is deliberately not interchangeable. A result event carries
phase="result", so a rule written for a call never fires on a result and vice
versa — and a rewrite lands in Decision.modified_result, not
modified_args, so the two directions cannot be confused. guard_mcp_session
wraps both halves, so you cannot install only one by accident.
A Rule is any callable SensorEvent -> Optional[Decision]; the first rule
returning a non-ALLOW verdict wins, and the default is allow.
from interlock import PolicyEngine, deny_tool, deny_when
import interlock
interlock.install(rules=[
deny_tool("shell_exec"),
deny_when(lambda e: "rm -rf" in str(e.args.get("cmd", "")), reason="destructive"),
])
# or build an engine explicitly
engine = PolicyEngine(rules=[...]).add_rule(deny_tool("drop_table"))
interlock.install(engine=engine)install() sets the process-global engine the interceptors consult. With no
engine installed, get_engine() returns an empty allow-all engine.
When one small set of policy templates must apply to millions of distinct
subjects (agents, tenants, sessions) with per-subject overrides, PolicyEngine
alone is the wrong shape. Scopes split code from data:
- a template is code — a factory turning params into a
Rule, few and fixed; - a binding is data —
(template_id, params, enabled)on a scope node; - a scope is a path, e.g.
("root", "acme", "agent42").
Root→leaf walking merges bindings per template with three rules: override (a
more specific binding wins), union (union_params accumulate — allowlists
grow), and tombstone (enabled=False removes an inherited binding until a
deeper node re-adds it).
from interlock.policy import Binding, ScopedEngine, ScopeRegistry, ScopeTree, Template
from interlock import Decision, SensorEvent
def egress_allowlist(allowed_hosts):
hosts = frozenset(allowed_hosts)
def rule(e: SensorEvent):
if e.action == "fetch" and (e.args.get("host") or "") not in hosts:
return Decision.block("egress to " + str(e.args.get("host")), "egress_allowlist")
return None
return rule
TEMPLATES = {
# union_params: allowed_hosts accumulates up the chain instead of overriding
"egress_allowlist": Template("egress_allowlist", egress_allowlist, frozenset({"allowed_hosts"})),
"rate_limiter": Template("rate_limiter", lambda limit: ...),
}
tree = ScopeTree()
tree.bind(("root",), [Binding("egress_allowlist", {"allowed_hosts": ["api.internal"]})])
tree.bind(("root", "acme"), [Binding("rate_limiter", {"limit": 5})])
tree.bind(("root", "acme", "agent42"), [Binding("rate_limiter", {"limit": 2})]) # override
registry = ScopeRegistry(tree, TEMPLATES)
registry.resolve(("root", "acme", "agent42"))
# {"egress_allowlist": {"allowed_hosts": {"api.internal"}}, "rate_limiter": {"limit": 2}}
interlock.install(engine=ScopedEngine(registry)) # same evaluate(event) contractCost model. Resolution walks one root→leaf path, compiles a PolicyEngine,
and caches it under a per-node version tuple — so warm cost stays flat from ten
scopes to a million, while memory grows only with bound nodes. Editing an
ancestor bumps its version and lazily invalidates only the descendants that
inherit from it. ScopeRegistry(maxsize=...) bounds the LRU. The default scope
deriver, principal_scope, splits event.principal on /.
Delegation chains. When an agent hands work to a sub-agent, the sub-agent must not end up with more authority than its delegator. A binding says where its authority came from, and the check happens at resolve time rather than at call time:
# agent42 itself resolves egress_allowlist to {"api.internal"}, so:
tree.bind(
("root", "acme", "agent42", "subagent7"),
[Binding("egress_allowlist", {"allowed_hosts": ["api.internal"]},
delegated_from=("root", "acme", "agent42"))], # resolves — a subset
)
# widening the delegate past its delegator is refused at resolve time:
tree.bind(
("root", "acme", "agent42", "subagent7"),
[Binding("egress_allowlist", {"allowed_hosts": ["api.internal", "exfil.example"]},
delegated_from=("root", "acme", "agent42"))], # DelegationOverGrant
)Every binding whose delegated_from is set must resolve to a subset of what
that scope resolves to, for the same template — so a delegate cannot hand itself
a host, a budget, or a capability its delegator never held. Two escapes are
closed: a delegation loop or unbounded fan-out trips
ScopeRegistry(max_delegation_depth=...) (default 16) with
DelegationDepthExceeded, and both failures are raised, not logged, so a bad
grant cannot survive into a warm cache. Non-delegated bindings are untouched —
delegated_from=None is a no-op, so nothing that never delegates changes shape.
The chain is read from event.parent_principal, which the interceptors carry from
the ambient context; delegation_chain(event) exposes it as a scope tuple,
delegator first.
SCOPE_BACKEND reports which storage layer is live ("python" or "rust"). The
native crate reimplements exactly one class — the sparse ScopeTree — with
interned segment ids and packed values, materializing Python objects only for the
one resolved leaf on a cache miss. Everything above it is unchanged.
105 single-purpose detectors, one rule factory per module, each with a
POLICY_ID. Import the eight general ones from the package, or any detector by
its module path:
from interlock.detectors import (
jailbreak_detector, prompt_injection_detector, execution_guard,
tool_allowlist, tool_denylist, network_egress_guard,
sensitive_path_guard, pii_redaction_guard,
)
from interlock.detectors.sql_injection_guard import sql_injection_guard
interlock.install(rules=[sql_injection_guard(), network_egress_guard(allowed_hosts=["api.internal"])])Full catalog, grouped by attack surface (ls interlock/detectors/ for the live list):
Injection & jailbreak framing — prompt_injection · jailbreak ·
dan_persona_guard · crescendo_guard · many_shot_jailbreak_guard ·
leetspeak_jailbreak_guard · hypothetical_framing_guard ·
social_engineering_framing_guard · refusal_suppression_guard ·
goal_hijack_guard · system_prompt_extraction · indirect_injection_marker ·
translation_evasion_guard · delimiter_smuggling · homoglyph_injection ·
unicode_tag_injection · html_comment_injection_guard ·
markdown_link_injection_guard · payload_splitting_guard ·
base64_payload_scan · canary_leak_guard · pinned_context_guard ·
memory_write_injection_guard · tool_output_override_guard ·
tool_result_injection_guard · rag_source_allowlist ·
mcp_prompt_arg_injection
Classic argument-shaped vulns — sql_injection_guard ·
nosql_injection_guard · shell_injection_guard · ldap_injection_guard ·
xxe_guard · ssti_guard · path_traversal_guard · prototype_pollution_guard ·
open_redirect_guard · crlf_header_injection_guard · script_tag_output_guard ·
ansi_escape_output_guard · url_scheme_guard · mass_assignment_guard ·
pickle_deser_guard · download_extension_guard
Secret & PII redaction — bearer_token_redactor · basic_auth_url_redactor ·
aws_arn_redactor · github_pat_redactor · google_api_key_redactor ·
jwt_redactor · pem_private_key_redactor · db_uri_redactor ·
stripe_key_redactor · slack_webhook_redactor · s3_presigned_redactor ·
output_secret_redactor · secret_entropy_egress_guard · env_secret_read_guard ·
git_credential_guard · email_pii_redactor · output_email_redactor ·
iban_redactor · ssn_redactor · us_phone_redactor · ipv4_redactor
Egress & data movement — data_egress · egress_rate_limiter ·
private_ip_egress_guard · raw_ip_egress_guard · data_volume_egress_guard ·
markdown_image_exfil_guard · host_fanout_guard
Code execution & host access — execution_guard · code_eval_exec_guard ·
code_dunder_escape_guard · code_filesystem_guard · code_network_guard ·
dynamic_import_guard · subprocess_spawn_guard · destructive_command_guard ·
sensitive_file_write_guard · signal_kill_guard · memory_bomb_guard
MCP trust surface — mcp_server_allowlist · mcp_tool_pinning ·
mcp_tool_description_scan · mcp_trust_registry · mcp_resource_uri_guard ·
mcp_sampling_model_guard · mcp_consent_budget · mcp_capability_lease ·
mcp_server_attestation · mcp_surface_baseline
Budgets, rate & loops — call_rate_limiter · cost_budget_guard ·
tool_budget_limiter · write_action_limiter · duplicate_call_loop_guard ·
output_length_guard · memory_write_size_guard · failed_auth_lockout
Provenance & integrity stamps — event_hmac_stamp · content_digest_stamp ·
sequence_number_stamp · provenance_origin_tag
Tool policy — tool_policy
BRACE compiles the same policy into a confinement profile and runs the agent's shell-outs under kernel-enforced isolation. One policy, two enforcement points: detectors intercept the calls interlock can see, BRACE confines the processes it cannot.
from interlock.brace import Sandbox, compile_profile
profile = compile_profile(registry.resolve(("root", "acme", "agent42")))
result = Sandbox(profile).run(["python", "worker.py"])
if "net" not in result.enforced:
raise RuntimeError("refusing to run untrusted code with open network")Sandbox.run(argv) checks the launch through the policy engine as a
process_spawn event first, so a rule can block or rewrite it before any process
starts, then runs it under the backend. Sandbox.plan(argv) returns the wrapped
argv without running anything — useful for inspection and tests.
A SandboxProfile describes allowed reach and defaults closed: no network unless
an egress_allowlist binding grants hosts, no writable host path unless a
sandbox_fs binding grants one, and the child's environment is exactly what you
pass in. compile_profile(resolved) reads the {template_id: params} dict a
ScopeRegistry already produces, recognising egress_allowlist,
sandbox_fs (read_paths/write_paths), and sandbox_limits
(cpu_seconds/memory_mb).
Backends, strongest first, chosen for the host at runtime — or pinned with
Sandbox(profile, backend="bubblewrap"):
| Backend | Host | Enforces | Reports (enforced) |
|---|---|---|---|
bubblewrap |
Linux | namespaces, ro/rw binds, network unshare | pid, net, fs-ro, fs-rw, env |
sandbox-exec |
macOS | Seatbelt s-expression profile | net, fs-rw, env |
soft |
anywhere | rlimits + scrubbed env — not a security boundary | rlimit, env |
available_backends() lists what the host offers. SandboxResult.enforced
reports the guarantees that actually held as short tokens so a caller that needs
real isolation can fail closed instead of silently accepting a downgrade to
soft. Each token in that column is a tested claim, not a wish: every row is
an escape run for real plus a relaxed control (the same attack with only that one
confinement lifted, asserted to succeed, so the block is attributable to the
sandbox and not to the kernel refusing the syscall). The guarantee table in
interlock/brace/__init__.py is the human-readable form and
tests/integration/test_brace_token_matrix.py is the executable one; a drift test
pins each backend's reported set to that table in both directions, so a token
cannot be added or dropped without the table moving with it.
Known limitation — the missing rung. nftables is intentionally absent from
the selectable set: Sandbox(profile, backend="nftables") raises
ValueError: unknown backend: nftables, and backends/nftables.py is reachable
only by a caller who is already root and imports it directly. So there is no
unprivileged host-granular egress backend between bubblewrap's --unshare-net
(all network or none) and nftables (per-host, but root/CAP_NET_ADMIN and
Linux-only). That is a measured absence rather than an unimplemented TODO:
Seatbelt's network filters accept only * or localhost as the remote host, so
an allowlist of specific hosts cannot be expressed on macOS at all. When a
profile grants network, sandbox-exec and bubblewrap therefore open egress
entirely and report no net token, honestly. Confining that grant to the
listed hosts happens in-process, via the egress_allowlist detector on captured
calls; at the host level, only the privileged nftables path does it.
Detectors decide; receipts record the decision, in a form a third party can
check afterwards without trusting the process that wrote it. It is off by
default — with no sink installed, emit_receipt is a no-op and the guard
behaves byte-identically to a build without receipts. Turning it on is explicit:
from interlock import install
from interlock.sink import FileSink
install(sink=FileSink("audit.jsonl")) # or set_sink(my_sink)Every enforced decision then appends one Receipt — ts, action, verdict,
reason, policy_id, attributed_to, principal, span_id, phase,
args_digest, prev, mac — from the enforcement sites only (the decorator's
_emit and the MCP interceptor's call and result paths), never from a rule, so a
rule cannot forge a receipt or quietly drop one.
Two properties are worth naming, because they are the whole point:
- The arguments are digested, not stored.
args_digestis a SHA-256 over the arguments; the receipt records that a payload was seen and which one, without becoming a second copy of your customer data. An audit log that duplicates every payload is a retention liability, not an audit log. - The chain is tamper-evident.
previs the previous receipt'smac(GENESISfor the first) andmaccovers every other field, so editing a receipt breaks its own mac, and deleting or reordering one breaks the next receipt's linkage.verify()walks the chain and names the first bad index, because "the chain is broken" without an index leaves the auditor hunting:
sink = FileSink("audit.jsonl", key=KEY)
sink.verify() # None, or "receipt 7: MAC mismatch (edited, or verified with a different key)"What the default signer claims — and what it does not. The default is
hmac_sha256: a MAC, not a public-key signature. It proves the content has not
changed since signing and that whoever wrote it held the shared secret. It does
not prove which holder produced it, and it is not verifiable by a party
without the secret. When you need attribution rather than integrity, inject a
signer through the same Signer seam (an Ed25519 signer from cryptography);
pyproject.toml stays dependency-free either way. The default key is random per
process (process_secret()), so a durable audit trail passes its own long-lived
key.
One sink instance, one process, one chain. Concurrent threads share an
unbroken chain — the read-prev / sign / append step runs under a lock. Writers in
other processes are not serialized: two processes appending to one file each
start their own chain at GENESIS, and verify() reports the junction as a
break rather than silently accepting a spliced history. How you partition the
audit file across processes is a deployment decision, and the sink does not make
it for you.
interlock/
enforce.py Verdict, Decision, Blocked
event.py SensorEvent
context.py span / principal via contextvars
_runtime.py the process-global engine
receipt.py Receipt, compute_mac, verify_chain, emit_receipt
sink.py ChainedSink, InMemorySink, FileSink
interceptors/ decorator.py, mcp.py, langchain.py (all funnel to enforce_tool_call)
policy/ engine.py (rules), scopes.py (templates, registry, ScopedEngine)
detectors/ 105 detector modules
brace/ profile.py, sandbox.py, result.py, backends/, trace.py
rust/ optional PyO3 accelerator for the scope store
js/ the Node.js client: wire, RemoteEngine, MCP + fs/exec/http + model-SDK interceptors
benchmarks/ scope-store memory and warm-resolve benchmark
tests/ 117 unit-test modules + tests/integration/
examples/ block_shell_tool.py
The public repo is the interlock/ package plus tests/. js/ is a separate
build system (npm) that never enters the wheel — pyproject.toml packages only
interlock*. rsi/ is internal, git-ignored build apparatus (see below).
git clone <repo> && cd interlock
python -m pytest # full suite from the repo root (conftest handles sys.path)
python -m pytest tests/test_scopes.py -q
python examples/block_shell_tool.py # the demo
python benchmarks/scope_memory.py # scope-store memory / warm-resolve numbersThere is no build step and no lint config to satisfy: the package is pure Python
with no dependencies, and conftest.py pins the repo root on sys.path and
resets the global engine between tests.
One module in interlock/detectors/, exposing a rule factory and a POLICY_ID:
"""What the payload class is, why it is visible at this point, and what is matched."""
POLICY_ID = "my_guard"
def my_guard(*, extra_patterns=None):
def rule(event):
# return None for "no opinion"; Decision.block(...) to deny
return None
return ruleConventions the suite relies on, and which keep the catalog uniform:
- Never raise on odd input. A non-dict
args, anint, orNoneyields no opinion. A detector is a guard, not a parser; raising is a fail-closed denial in production and a test failure here. - Precision over recall. Each detector matches a shape an ordinary value will not contain, and says so in its module docstring — the reasoning is the documentation.
- One rule, one concern, no imports beyond stdlib. Detectors must stay dependency-free, like the core.
- Add a test module
tests/test_<name>.pycovering both the match and the non-match (the non-match is what stops the detector from being noise). - Detectors are deliberately not auto-registered. Import and
installthe ones you want; nothing is enabled behind your back.
The Rust crate reimplements only ScopeTree. Build it, then confirm the swap:
cd rust && maturin develop --release
python -c "from interlock.policy import SCOPE_BACKEND; print(SCOPE_BACKEND)"
python -m pytest tests/test_scope_backend.py -q # checks native/python paritytests/test_scope_backend.py uses the pure-Python tree as the parity oracle, so
the two backends must agree on every resolution. *.so and rust/target/ are
git-ignored; the canonical build is maturin develop.
This README drifted once already: it advertised a milestone roadmap while the tree had moved well past it, and its detector count was stale by more than a dozen. Before editing, check the claim against the code:
ls interlock/detectors/*.py | grep -v __init__ | wc -l # detector count
git ls-files interlock/ | wc -l # tracked package files
python -m pytest -q # suite must be greenrsi/ is a git-ignored recursive-self-improvement loop that grows the detector
catalog by mining recent arXiv guardrail research and adding one verified
technique at a time. The discipline is that every change is judged by what the
code actually did — interlock's own pytest — not by a model's opinion of its own
output; a vacuous pass is downgraded to a failure. rsi/backlog.md holds the
ranked technique families, rsi/verify.py is the reward gate, and the paid
worker lanes spend real money, so run it deliberately:
PYTHONPATH=. RSI_TARGET=100 python rsi/run_100.pyShipped in this tree:
- Tri-state verdicts, fail-closed rule engine,
@guard/monitor, and run identity. - In-process MCP interception (
guard_mcp_session) and the LangChain adapter — explicit wrap, one sharedenforce_tool_callcore. - Hierarchical policy scopes: override / union / tombstone merge, version-keyed
cache,
ScopedEngine, delegation chains with resolve-time attenuation (delegated_from/max_delegation_depth), and the optional nativeScopeTree. - 105 detectors across injection, redaction, egress, code execution, MCP trust, budgets, and provenance.
- Verdict provenance: every non-ALLOW
Decisioncarriesattributed_to, so a block names the input that triggered it, not merely the rule that fired. - Signed execution receipts and an append-only, tamper-evident sink
(
interlock/receipt.py,interlock/sink.py) — off by default, opted into withinstall(sink=...). - Tool-result sanitization:
enforce_tool_resultpoints the same tri-state decision at the observation path, so injection in a tool result is remediated rather than only observed. - MCP capability leases, server attestation, and tool-surface drift detection,
extending the
mcp_*detectors from static allowlisting to time-boxed grants. - BRACE confinement: profile IR,
compile_profilefrom policy,Sandbox, the bubblewrap / Seatbelt / soft backends with reported guarantees, and an escape-resistance matrix that runs each claimed token's attack for real. - A scope-store benchmark under
benchmarks/.
Upcoming, roughly in expected-value order — the first three come from a survey of
445 arXiv papers from 2026
(research/agent-guardrails-2026.md),
the rest are open items from earlier milestones:
Verdict.ESCALATE— a fourth outcome: pause for a human, with the evidence attached, instead of approximating it with a block.- Guardrail robustness harness — false-positive rate on benign traffic and resistance to induced blocking. A fail-closed library with an unmeasured FP rate is a product risk; forcing blocks is a denial-of-service primitive.
- Flow-typed policy dimension — express "output of tool A must not reach argument B", so the policy layer does not dead-end as flat rules.
- Declarative YAML/JSON policies with a safe condition evaluator and hot reload.
- Automatic interception — import hooks for MCP clients and model SDKs, plus dependency tamper detection.
- An overhead benchmark suite with a stated methodology.
Each item above has a full entry — problem, why now, what changes, impact on
callers, a falsifiable acceptance criterion, and its risks — in
ROADMAP.md. Logit / decoding-layer intervention is recorded there
as explicitly out of scope.
Apache-2.0. See LICENSE.