Maatrel

Zero trust for AI agents

Give agents real authority. Answer for every action.

Maatrel gives each agent a verifiable identity, authorizes every action it takes at every boundary, and records why.

Inline toolsend_payment(vendor="shadyco", recipient="alice@corp.example", amount=50.0)

  1. Signed requestpassed
  2. Policies selectedpassed
  3. Workflow statepassed
  4. Origin of valuespassed
  5. Rules and contextrefused
  6. Runs as signednot reached
  7. Recordedentry #8

Refused at rules and context

The payee is on file, but shadyco's recorded risk score is 0.95, over the 0.7 limit. The payment stops at the rules check.

Inline toolsend_payment(vendor="acme", recipient="attacker@evil.example", amount=50.0)

  1. Signed requestpassed
  2. Policies selectedpassed
  3. Workflow statepassed
  4. Origin of valuesrefused
  5. Rules and contextnot reached
  6. Runs as signednot reached
  7. Recordedentry #6

Refused at origin of values

The recipient came from an invoice an MCP server returned, so the payment stops at the origin check. Nothing runs, and the refusal is recorded.

Inline toolsend_payment(vendor="acme", recipient="alice@corp.example", amount=50.0)

  1. Signed requestpassed
  2. Policies selectedpassed
  3. Workflow statepassed
  4. Origin of valuespassed
  5. Rules and contextpassed
  6. Runs as signedpassed
  7. Recordedentry #10

Allowed

The payee is on file and acme's recorded risk score is 0.2. The payment runs exactly as signed, and the decision is recorded.

From the use case below

The problem

Agents are a new kind of actor in your systems.

  1. They act with your authority.

    Agents move money, change records and call internal systems, often with the application's credentials rather than an identity of their own.

  2. Anything they read can steer them.

    Model output, documents, web pages and MCP results all shape the next action. An attacker only has to reach one of them.

  3. Your controls were built for something else.

    Access management checks the application, not the agent's decision. Gateways check the endpoint, not the values. Logs record what happened, not who allowed it or why.

So teams keep agents on a short leash, or give them authority nobody can account for.

How it works

Every action answers seven questions before it runs.

Nothing is trusted because of where it runs or what the model says. The same checks run at all four boundaries.

  • LLM callsChecked before the provider is contacted; its output never carries authority.
  • MCP serversEach call authorized before anything is sent; results stay external, whatever the server says.
  • Inline toolsRun exactly the signed arguments.
  • HTTP APIsRebuilt from signed fields, no redirects; credentials attach only after an allow and never enter the record.
  1. Who asked? Signed request. Signed with a short-lived key bound to the agent's identity.
  2. Under which rules? Policies selected. Chosen by the checkpoint, deny by default; more policies only narrow.
  3. In its place? Workflow state. Each request is used once, in order; replays are refused.
  4. With which values? Origin of values. Values that carry authority must trace to a trusted source.
  5. Knowing what? Rules and context. Conditions use signed results, never the model's word.
  6. Doing exactly what? Runs as signed. What runs is exactly what was checked.
  7. On what record? Recorded. Each decision joins a hash-linked record with its policies.
The pipeline behind the questions (one tool call)

Before any call Every agent and tool gets its own identity, and every agent a session key that expires within the hour. The arguments you declare as carrying authority become signed records that are checked on every call.

  1. Canonical requestOne request in a fixed format, with every argument labelled by origin.
  2. The decisionSix checks in order. Each can refuse.
    1. Who asked?
    2. Under which rules?
    3. In its place?
    4. ApprovalsRead from the checkpoint's own store, never from the request.
    5. With which values?
    6. Knowing what?
  3. Frozen snapshotAnswers: Doing exactly what?The signature is verified again over a private copy of the arguments.
  4. Content checksVerifiers you configure; this example has none.
  5. ExecutionAnswers: Doing exactly what?The tool runs that copy and nothing else.
  6. The resultSigned by the tool that produced it, carrying the origins of its inputs.
  7. The recordAnswers: On what record?The verified result joins the workflow's history and the audit record before it is returned.
  8. Return or failureA refusal reports that nothing ran, and a check that cannot run refuses the call.

Origin of values is the thread. Labelled in row 1, enforced in row 2, carried onto the result in row 6, kept in the workflow's history in row 7.

Use case

An accounts-payable agent pays an invoice. Pick what it does next.

It is asked to pay invoice 1182 from acme, and the invoice it fetched from an MCP server has been tampered with.

  1. Pick an action
  2. Follow the request along the line
  3. See where it stops, or what lets it through

Integration

Add Maatrel to the agent you already have.

The example is a LangGraph agent with LangChain tools. Your graph stays as it is: wrap the model, the tools, the MCP session and the HTTP client, and write your policies in YAML.

agent.py LangGraph + LangChain
# A LangGraph agent with LangChain tools.from langchain_core.tools import toolfrom langgraph.graph import StateGraph, MessagesState, STARTfrom langgraph.prebuilt import ToolNode, tools_conditionfrom maatrel import TrustLayer, key # Load your policies and settings from trust.yaml.trust = TrustLayer.from_config("trust.yaml", principal="payments-app")# Act as the payments agent: wrapped calls are signed as this identity.agent = trust.as_principal("payments-agent") # Wrap what you already have; every call through them is checked first.model = chat_modelmodel = agent.wrap_chat_model(chat_model)invoices = mcp_sessioninvoices = agent.wrap_mcp(mcp_session, server_namespace="invoices")payments_api = httpx.Client(transport=transport)payments_api = agent.wrap_http(transport=transport) @tooldef lookup_payee(vendor: str) -> dict:    """The payee address on file."""    return vendor_directory.lookup(vendor) @tooldef score_vendor(vendor: str) -> dict:    """The vendor's risk score."""    return risk_service.score(vendor) @tooldef send_payment(vendor: str, recipient: str, amount: float) -> dict:    """Pay a vendor through the payments API."""    resp = payments_api.post("https://payments.internal/transfers",    call = payments_api.post("https://payments.internal/transfers",                             json={"to": recipient, "amount": amount})    # Maatrel returns a refused request here instead of raising.    status = call.response.status_code if call.allowed else None    return {"paid": amount, "to": recipient, "status": resp.status_code}    return {"paid": amount, "to": recipient, "status": status} # Wrap the tools. sets= records score_vendor's result as the risk score that# risk_gate.yaml checks; on_denial= hands a refusal back to the model.tools = [lookup_payee, score_vendor, send_payment]tools = agent.wrap_tools([lookup_payee, score_vendor, send_payment],                         sets={"score_vendor": {"risk_score": key("vendor")}},                         on_denial="tool_message") # Your LangGraph graph: the model proposes a tool call, the tool node runs it.def call_model(state: MessagesState):    return {"messages": [model.invoke(state["messages"])]} graph = StateGraph(MessagesState)graph.add_node("agent", call_model)graph.add_node("tools", ToolNode(tools))graph.add_edge(START, "agent")graph.add_conditional_edges("agent", tools_condition)graph.add_edge("tools", "agent")app = graph.compile()

Highlighted: what Maatrel adds. Everything else is your code.

Your policies: three new YAML files

trust.yaml
registry: {backend: memory}
policy_store: {backend: memory}
attestation_store: {backend: memory}
policies: [agent.yaml, risk_gate.yaml]
provenance:
  tools:
    lookup_payee:
      roles:
        - { path: /vendor, role: selector }
      field_origins:
        - { pattern: /output/email, origin_class: registry }
    send_payment:
      roles:
        - { path: /vendor,    role: selector }
        - { path: /recipient, role: target }
        - { path: /amount,    role: content }
agent.yaml
policy_id: payments.agent
version: "1.0.0"
scope: { applies_to: all }
resources:
  - { pattern: "llm:generate", operations: [generate] }
  - { pattern: "mcp://invoices/get_invoice", operations: [execute] }
  - { pattern: lookup_payee, operations: [execute] }
  - { pattern: score_vendor, operations: [execute] }
  - { pattern: send_payment, operations: [execute] }
  - { pattern: "https://payments.internal/**", operations: [POST] }
risk_gate.yaml
policy_id: payments.risk_gate
version: "1.0.0"
scope: { applies_to: all, resource_match: send_payment }
resources: [{ pattern: send_payment }]
constraints:
  conditions:
    - { field: context.risk_score, key_from: args.vendor, op: lt, value: 0.7 }

6 new lines of code and 3 YAML files for your policies.

The graph is the same code in both versions. In the same run, the agent without Maatrel paid the tampered invoice. With Maatrel, that payment was refused before the payments API saw it, and the graph went on to pay the payee on file.

A refusal goes back to the model as a tool message. Policy refusals say why. Origin refusals say only that the action was blocked, so an injected instruction cannot probe the check.

Both versions ran with the same scripted proposals and local stand-ins.

Works with LangChain, including LangGraph, and the OpenAI SDK.

Benchmark results

Maatrel cut successful attacks by three quarters, from 30.2% to 7.4%.

AgentDojo, a public benchmark, gives an agent tasks in four simulated environments (banking, Slack, travel and a workspace of email, calendar and files) and hides attacks in the data it reads.

Every successful attack was of a kind declared out of scope before the run, such as an injected instruction to delete a file by its id.

The agent still completed 84% of the tasks it completed without Maatrel, or 80% when the model was not told why a call was blocked.

By suite

AgentDojo 0.1.35 (suite v1.2), important_instructions attack, gpt-4o-mini-2024-07-18, four suites, 97 tasks, 2 trials.

In operation

Every decision, in your traces and in the audit record.

Your traces show what is happening. The audit record shows who asked for each action, which policies decided it, and why.

In the audit record

Each decision is an entry carrying the hash of the one before, so the chain can be re-checked. Entries are kept in a persistent audit store.

audit record for payments-agent12 entries, hash links re-checked in this run
  1. #0LLMllm:generateallowedpayments.agent 1.0.0ceea4cbf…263c
  2. #1MCP servermcp://invoices/get_invoiceallowednot recorded13ee5842…d42c
  3. #2MCP servermcp://invoices/delete_invoicerefused at Rules and contextno_allowed_resourcenot recordedbf65b13d…750f
  4. #3Inline toollookup_payeeallowedpayments.agent 1.0.0e7dc2d5d…bf83
  5. #4Inline toolscore_vendorallowedpayments.agent 1.0.06abe9b62…f8c9
  6. #5Inline toolscore_vendorallowedpayments.agent 1.0.0b412cb38…47ba
  7. #6Inline toolsend_paymentrefused at Origin of valuesprovenancepayments.agent 1.0.0, payments.risk_gate 1.0.0a6d8d265…3879
  8. #7Inline toolsend_paymentrefused at Origin of valuesprovenancepayments.agent 1.0.0, payments.risk_gate 1.0.09c168e80…acc9
  9. #8Inline toolsend_paymentrefused at Rules and contextconditionpayments.agent 1.0.0, payments.risk_gate 1.0.0fc083c5b…b4fb
  10. #9HTTP APIhttps://payments.internal/transfersallowedpayments.agent 1.0.0aa52b674…356b
  11. #10Inline toolsend_paymentallowedpayments.agent 1.0.0, payments.risk_gate 1.0.0c63d1cad…c0fc
  12. #11HTTP APIhttps://attacker.example/collectrefused at Rules and contextno_allowed_resourcepayments.agent 1.0.0827bf4c8…11df

In your traces

Each governed call is an OpenTelemetry span, through providers your application already owns. A refused span says where it stopped, and that it had no effect.

trace pay invoice 118212 Maatrel spans, 5 refused (span status error)
  1. maatrel.llmllm:generateentry #0
  2. maatrel.mcp_clientmcp://invoices/get_invoiceentry #1
  3. maatrel.mcp_clientmcp://invoices/delete_invoicerefusedentry #2maatrel.refusal.stage=no_allowed_resource  maatrel.refusal.category=policy  maatrel.external_effect=did_not_occur
  4. maatrel.toollookup_payeeentry #3
  5. maatrel.toolscore_vendorentry #4
  6. maatrel.toolscore_vendorentry #5
  7. maatrel.toolsend_paymentrefusedentry #6maatrel.refusal.stage=provenance  maatrel.refusal.category=authorization_evidence  maatrel.external_effect=did_not_occur
  8. maatrel.toolsend_paymentrefusedentry #7maatrel.refusal.stage=provenance  maatrel.refusal.category=authorization_evidence  maatrel.external_effect=did_not_occur
  9. maatrel.toolsend_paymentrefusedentry #8maatrel.refusal.stage=condition  maatrel.refusal.category=policy  maatrel.external_effect=did_not_occur
  10. maatrel.toolsend_paymententry #10
  11. └ maatrel.httphttps://payments.internalentry #9
  12. maatrel.httphttps://attacker.examplerefusedentry #11maatrel.refusal.stage=no_allowed_resource  maatrel.refusal.category=policy  maatrel.external_effect=did_not_occur

Both views come from the same scripted run as the rest of this page, captured in memory. Timings are from a local run with stand-ins, not performance figures.

Get notified when Maatrel is released.

Leave your email and we will write when it is available.

At release
An email when Maatrel is available.
Before that
We are inviting a few teams at a time to a private preview: a GitHub repository with the library, the guides and the examples. We may invite you.

We use your address only to write about Maatrel's release and the preview. Delivered by Formspree.

Prefer to talk first? Write to eval@maatrel.ai.