Self-serve
Run it in your boundary.
One command installs the ArtzAIn decision engine and its dashboard on your own machine or VM. No account. No card. No vendor callouts. Your data never leaves your boundary, and the evidence it seals verifies offline. You need Docker and Python 3.10+; the installer checks everything else and tells you the one-line fix for anything missing. On Debian and Ubuntu it uses pipx when present, then pip --user, then a dedicated virtualenv: distro pip blocks pip install --user (PEP 668).
The command ends with your browser on the first-run page: set an email and password (that account is the platform admin) and land on the operator checklist. First sealed decision inside 15 minutes, without editing a single file. Both scripts are short and readable on purpose: install.sh · install.ps1.
Then wire it into your agent: where decide() goes when your model calls a tool, and how to cover every tool.
Requirements
A laptop is enough.
What you get
The real engine, the real dashboard, your keys.
This is the same decision core the hosted platform runs, not a demo build. Every agent action is checked against policy before it runs and sealed into an Ed25519-signed, hash-chained audit ledger before the response returns. The operator dashboard, Roger, the policy catalog, the agent registry, and the full operator manual are served locally beside it. Signing keys are generated inside your boundary on first boot and never leave it, and evidence exports verify offline with artzain audit verify. Every intact export verifies as SELF-ATTESTED, a pass that never blocks, with or without a license certificate: ATTESTED also requires the CogNEXUS Evidence Root to be pinned, and as of artzain 0.6.19 it is not.
The install runs the GPU-free core profile: local LLM chat, the coder fixer, and LightGBM propensity stay idle (they mark themselves degraded rather than pretending), while the decision path, sealing, policy enforcement, connectors, and evidence exports are all fully live. GPU inference in your boundary is an engagement conversation: hello@cognexuslabs.ai.
Tool calls
Put decide() between the model and the side effect.
ArtzAIn does not hook your LLM client. Nothing intercepts tool calls automatically, and there is no
middleware to install in front of the model. (The governance
envelope is a base_url swap that screens prompts, reply text, and each tool call a reply
proposes. It clears proposals, not executions: it never sees a tool run, or what your code does with a
call once the reply has arrived. Its tool-call decisions are filed under
chat_completion_output, but the checks keyed on action read the tool's name
from the call, so they classify a proposal as they classify the same call sent to decide()
below. Two things differ: the envelope puts no lawful_basis or transaction in
context, so under failure_policy: "closed" a proposed tool that needs a lawful
basis is refused before your code sees it; and the agent registry judges the envelope's agent as it
judges the envelope's own traffic.)
You place decide() yourself, in the gap between the model proposing a tool call and your
code executing it. That gap is the control point.
user input → screen_user_input() → model → tool call → decide() → execute
Input screening runs before the model. decide() runs after the model and before the side effect.
Mapping a tool call onto a decision
| Argument | Comes from |
|---|---|
action | The tool name the model chose |
target | The resource the call touches, read out of the arguments |
payload | The whole call as JSON: {"tool": name, "arguments": {…}} |
kind | "tool_call" |
On the engine, the checks keyed on action (a lawful basis for personal-data actions, the
compliance bundles' action classes, fraud scoring for transactions, your agent's capability scope in the
registry) also read the tools the payload names: every name a call gives its tool (tool,
name, tool_name, function or its name), and both
values of a key given twice, for each of the first 20 calls. Name the tool as action and
they see one action, as before. If action names something else, the decision is classified
under both: the payload can add a finding, and never clears one that action draws. A
mismatch is not a finding in itself. The contract check reviews a payload whose tools cannot all be read
that way: more than 20 calls, a list holding something other than calls, either of those in the value a
first-wins parser keeps for a repeated key, or nesting too deep to read.
kind="tool_call" makes the payload a structured call rather than prose. On the engine, the
call's shape and any per-tool contracts in your policy bundle are checked. The destructive-action,
injection, policy and PII screens read the serialized call, and every string in it as the tool receives
them: JSON-decoded, including JSON inside a string, such as a stringified arguments, up to
three levels deep. The destructive-action screen reads each string on its own, and each list of two or
more strings joined with spaces, the way an argv list runs (a stray number in the list is included), so a
keyword list such as ["drop", "table"] reads as the command it spells. The injection, policy
and PII screens read the strings together, so one argument can change the result for another: a customer
named in one argument makes abusive language in another a conduct finding. The conduct rules count a
client word in any value of the call; an argument or tool name counts only when it holds the
profanity as well. So an argument called customer, or a tool called
customer.notify, does not by itself make profanity in a message a client finding
(offline, since artzain 0.6.24), as the same message sent as model_output would not be
one; a sentence used as a key in a translation catalog can be. A call that is not JSON, a value
holding JSON that a strict parser does not read (a document cut short, say) and JSON nested in
strings past the third level are read as text, names included. The PII screen runs on the
engine only.
Send the whole call, not only the arguments: bare arguments fail the shape check and go to
review, or, when one of them happens to be called name or tool, are
read as a call to a tool of that name. model_output is for text your model wrote that is
headed for a side effect, such as an email body. When that text is JSON, the destructive-action,
injection, privacy, policy and EU-overlay screens read it decoded as they read a call, and its
strings draw the findings listed below. A reply has no tool name, so a client word in a JSON key
still names a client for the conduct rules.
Parse the model's arguments with a strict JSON parser and send the parsed call, as the examples below do.
If your dispatcher repairs malformed JSON before a tool runs, what runs can differ from what was
screened. A command a tool assembles from separate fields, such as a cmd beside its
args, is not seen whole either: screen the assembled command inside the step as well, as
described under Coverage.
An argument's text draws the same injection findings it would draw as
user_input; an argument that is itself valid JSON is screened through its decoded strings
instead. For example:
- A line of only three or more
-or#, or of only three backticks (such as the fence that closes a Markdown code block), is amediumdelimiter finding, and the call comes backreview. A policy bundle can relax that with apermissiveinjection preset fortool_call, at the cost of every othermediuminjection finding. - A run of four or more
\xNNor\uNNNNescapes written out as text, the way source code holds them, can be ahighencoding finding, and sodenyunder any preset. - Base64 is decoded and checked for words such as
root,adminorpasswordwhen the decoded bytes are text, including base64 wrapped across lines. A match is ahighencoding finding. A certificate, key, image or compressed file decodes to bytes that are not text, and a word inside it, such as a root CA's name, is not a match (offline, since artzain 0.6.18). Base64 of text is searched whatever it is for, so a JWT whose claims name anadminrole can deny, and so can a file that is mostly text, such as a ZIP archive of text files stored without compression.
Serialize with ensure_ascii=False, as the examples below do. Without it, each non-ASCII
character travels as a six-character \uXXXX escape (twelve for a character above U+FFFF,
such as most emoji), which counts against the payload limit. Offline screens before artzain 0.6.16, and
engines without this change, also read a run of those escapes as an encoding attack. Current screens still
do when the escaped characters are invisible, except the England, Scotland and Wales flag emoji (offline,
since artzain 0.6.18). Invisible characters that can carry hidden text, such as tag characters or a long
string of variation selectors, are a finding whether they are escaped or not.
The result carries outcome, which is allow, deny, or review.
Only allow runs the tool. review means do not act yet; on the engine, a human
resolves it in the review queue. decide() raises DecisionError on any non-2xx
response, when the engine cannot be reached, and when the request cannot be sent. A 503 means the
engine refused to decide. A payload holding an unpaired surrogate is not sent at all:
json.loads accepts the escape of half a UTF-16 pair, which is what a JavaScript string
cut in the middle of an emoji serializes to, and UTF-8 cannot carry it. Offline, the same payload
comes back deny, and a call that carries the surrogate as a JSON escape comes back
deny online and offline alike. Treat every DecisionError as
deny, as the examples below do.
OpenAI-style tool calls
Tool calls arrive as a tool_calls array, and function.arguments is a JSON string.
The assistant message goes back into the history ahead of the tool results that answer it.
pythonimport json
import artzain
resp = client.chat.completions.create(model=..., messages=messages, tools=tools)
msg = resp.choices[0].message
messages.append(msg)
for tc in msg.tool_calls or []:
name = tc.function.name
try:
args = json.loads(tc.function.arguments)
except json.JSONDecodeError:
args = None
if not isinstance(args, dict):
d = {"outcome": "deny"} # unparseable arguments: never run them
else:
try:
d = artzain.decide(
action=name,
target=str(args.get("id") or args.get("to") or "unknown")[:300],
payload=json.dumps({"tool": name, "arguments": args}, ensure_ascii=False),
kind="tool_call",
)
except artzain.DecisionError:
d = {"outcome": "deny"} # the engine did not decide: fail closed
if d["outcome"] == "allow":
result = dispatch(name, args)
elif d["outcome"] == "review":
result = queue_for_human(tc, d)
else:
result = "Blocked by policy."
messages.append({
"role": "tool",
"tool_call_id": tc.id,
"content": str(result),
})
Anthropic tool calls
Same idea, different shape. Tool calls arrive as tool_use content blocks rather than a
tool_calls array, and block.input is already a dict, so serialize it yourself for
the payload. Results go back as tool_result blocks inside a user message.
pythonimport json
import artzain
resp = client.messages.create(model=..., messages=messages, tools=tools, max_tokens=1024)
results = []
for block in resp.content:
if block.type != "tool_use":
continue
try:
d = artzain.decide(
action=block.name,
target=str(block.input.get("id") or block.input.get("to") or "unknown")[:300],
payload=json.dumps({"tool": block.name, "arguments": block.input}, ensure_ascii=False),
kind="tool_call",
)
except artzain.DecisionError:
d = {"outcome": "deny"} # the engine did not decide: fail closed
if d["outcome"] == "allow":
content = str(dispatch(block.name, block.input))
elif d["outcome"] == "review":
content = str(queue_for_human(block, d))
else:
content = "Blocked by policy."
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": content,
"is_error": d["outcome"] == "deny",
})
if resp.content:
messages.append({"role": "assistant", "content": resp.content})
if results:
messages.append({"role": "user", "content": results})
On a deny, feed the block back as atool_resultrather than dropping it. Everytool_useblock requires a matching result, and telling the model it was blocked usually produces a sensible next step instead of a retry of the same call.
Coverage
Cover every tool, not every call site.
Wrapping each tool call site individually does not scale. Coverage ends up resting on whoever adds the next tool remembering to gate it, and that is the first thing to go wrong. Two patterns make coverage structural instead.
Gate the dispatcher, not the call sites
Nearly every agent funnels tool execution through one place: a registry lookup, a
dispatch(name, args) function, or an MCP client. That is one chokepoint, not one per tool.
Put decide() there and every tool is covered by construction, including tools added later by
people who never read this page.
pythonimport json
import artzain
TOOLS = {} # your existing name -> callable registry
def guarded_dispatch(name, args):
try:
d = artzain.decide(
action=name,
target=resolve_target(name, args),
payload=json.dumps({"tool": name, "arguments": args}, ensure_ascii=False),
kind="tool_call",
)
except artzain.DecisionError:
d = {"outcome": "deny"} # the engine did not decide: fail closed
if d["outcome"] != "allow" or name not in TOOLS:
return blocked(d)
return TOOLS[name](**args)
If your tools are declared with a decorator, move the guard into registration. When @tool
only ever registers a function wrapped in guarded_dispatch, an ungated tool cannot be
registered at all.
Default-deny unknown tools
The dispatcher pattern handles tools that go through your loop. Default-deny handles the ones you forgot.
Declare the tools you know in the guard_config of your team's policy bundle, and let
anything else fall through to an outcome that is not allow:
json"guard_config": {
"tool_contracts": {
"send_email": { "required_args": ["to"] },
"create_ticket": {},
"*": { "deny_unknown_tools": true }
},
"resolution": { "medium": "review" }
}
A call naming a tool the bundle does not declare draws a medium finding on the
tool-call-contract vote:
call[0] 'drop_table': tool not declared in bundle tool_contracts
On its own a medium finding is advisory, and the call is allowed. The resolution
line escalates it to review (or to deny, if you prefer), and it escalates every
medium vote in the decision, not only this one. Because the dispatcher runs a tool only on
allow, review stops the call as surely as deny does.
A tool nobody thought to declare then stops the first time it runs against your engine.
reasons names the tool and says the bundle escalated the call, and the review queue entry
carries the same line:
tool-call-contract (medium, escalated to review by bundle): call[0] 'drop_table': tool not declared in bundle tool_contracts
The vote itself still reads allow in contributing_agents, because the escalation
belongs to the bundle, not the enforcer. reasons quotes only the vote's first finding; for a
batch of calls, read the rest from that vote. Omission stops being a silent hole and becomes a visible
error, which is the behavior you want from a control plane. It needs kind="tool_call" and an
engine with the bundle active, and the install above is enough.
Worth knowing where the boundary sits. decide() governs intent, so it sees what the
agent routes through it. A library making its own HTTP call, or a subprocess spawned outside your tool
layer, is not visible to it. Complete coverage at that level needs an effect-layer control underneath:
egress proxy, network policy, or syscall filter. Those catch everything and understand nothing, so the
usual arrangement is both, semantics at the dispatcher and a backstop at egress.
decide() or screen_agent_action()
decide() is the sealed path. With an API key configured, the call is policy-evaluated and
written to the hash-chained, Ed25519-signed audit ledger before the outcome comes back, so it is what
belongs on anything consequential. It is also what the evidence bundle is built from:
bashartzain audit export --profile eu-ai-act # needs COGNEXUS_API_KEY
artzain audit verify <bundle> # offline; signature checks need artzain[verify]
screen_agent_action() is the lighter destructive-action guard, and it runs locally. Use it
when you only want pattern screening on SQL, shell, git, or cloud payloads and do not need a signed
decision on the ledger. A critical match trips the kill switch for that run and raises
AgentKilledError. By default, five critical trips within 60 seconds in one process trip a
global panic that stops every run until clear_global_panic().
pythondef execute_sql(sql):
artzain.screen_agent_action(sql, run_id=run_id, agent_id="billing-bot", source="execute_sql")
return db.execute(sql)
The two compose cleanly: screen inside a step, decide at the boundary.
Offline mode
With no API key configured, decide() still runs. The same three guards (prompt injection,
destructive action, and policy rules) evaluate in-process, and the result comes back with
offline: true. Nothing is sealed and no team bundle applies, so the shape check, tool
contracts and default-deny do not run offline.
The call site does not change when you connect. Create an API key on the dashboard, set
COGNEXUS_API_KEY to it and COGNEXUS_API_BASE_URL to your install
(http://localhost:8080 by default), and the same code produces sealed decisions. There is no
rewrite between trying it out and running it for real. The outcomes can change, though: online, the
shape and contract checks, your bundle and the rest of the engine's enforcers run, so a call that was
allow offline can come back review or deny. Run the loop against
your install before you rely on it.
Choosing what to gate
Gating every tool call is the strictest setting and the noisiest. The common pattern is to leave reads
ungated and put decide() on anything that changes state or leaves the building: writes,
sends, payments, deploys, deletes, and outbound messages. If you skip reads, keep an explicit list of
read-only tools in the dispatcher, so a new tool is gated until someone decides it is a read.
Reversibility is a useful sorting rule. If undoing the action is cheap, screening is usually enough. If undoing it means an apology, it belongs on the ledger.
The examples use the Python SDK (pip install artzain). The TypeScript client takes the same
fields but is remote-only. The same guide, condensed, is in the
SDK repository README.
The trial, plainly
30 days. Then expansion pauses. Governing never stops.
A fresh install runs a 30-day trial, counted from first boot and visible on the dashboard from day one. When it ends, connecting new sources, registering new agents, and minting new API keys pause until a license certificate is imported. Everything already running keeps running: decisions keep sealing, existing connectors keep syncing, and your evidence keeps verifying. An engine whose job is governance does not hold your agents hostage over paperwork. Importing the certificate lifts the pause live, no restart.
Supply chain
Digests you can verify, not tags you have to trust.
The installer never pulls a floating tag. Releases publish to a public registry as multi-arch images whose digests are recorded in the stable-channel manifest and signed with cosign by the release workflow's own identity, verifiable by anyone:
Leaving is honest too: artzain local down stops the stack and keeps your data; down --purge destroys it only after you type the install's own id back, and reminds you to export your evidence first. Nothing phones home either way.