Brainy Brainy
Docs Brainy

Guardrails

In this section

A subagent does not go rogue because it is malicious. It goes rogue because it lost the context it was handed: the rules were prose in a prompt, the window flooded, the prompt was compacted, and what is left is a capable model with no restrictions and a tool belt. Everything downstream of that — the deleted directory, the publish nobody approved, the public repo — is a consequence of one design mistake: the rules lived in the agent.

Guardrails move them into the store. A rule is a ROW: durable, generation-pinned, bound to an actor, a scope and a list of doors, carrying the authority that let it exist and the one sentence a refused caller is shown. Every door evaluates the rows before it decides anything. An agent that never read a rule still meets it. An agent whose window flooded loses nothing that binds it.

A guardrail is stored under the row kind policy-rule, and the mirror job keeps its name, policy-mirror-rebuild: the stored names did not change with the word.

The shape of a rule

{
  "id": "only-the-source-registry",
  "subject": "*",
  "scope": { "pathPrefix": "/projects/brainy" },
  "doors": ["Bash", "add", "github_*"],
  "effect": "refuse",
  "reason": "Packages publish only to The Source, never npmjs.",
  "authority": "record:PLATFORM-RENAME-TRAIN-FLEET",
  "locked": true,
  "since": "2026-09-15",
  "until": null,
  "match": [
    { "field": "command", "matches": "npm publish" },
    {
      "field": "command",
      "notMatches": "--registry[= ]https://source\\.soulcraft\\.com"
    }
  ]
}JSON

field

what it is

why it is a field and not a convention

id

a slug

it is quoted verbatim in the refusal, so it must read as a reason

subject

an actor slug, role:<name>, either negated (!role:steward), or *

an actor is whoever the door AUTHENTICATED — never a name the request supplied about itself

scope

pathPrefix · world · store

absent axes are unscoped; a present axis must match or the rule stands aside

doors

door names, tool names, globs (github_*), or *

a rule may govern tools brainy does not serve — see below. Any name for a door names that door: the door's own (add, memory_rebuildProjections, accord_postRound, vfs_writeFile), its MCP tool (brainy_add, accord_file_ask for file, memory_list_rows for list, memory_search for context), or the snake-case spelling of a memory_, accord_, vfs_ or guardrail_ door (memory_rebuild_projections, accord_post_round, vfs_write_file). Rules are stored and checked under the door's own name. The tool names come from the MCP catalog itself, so a new tool needs no edit here. A glob is matched as written, and a tool that is not one door (memory_write, the session tools) is checked at the doors its writes reach

effect

allow · refuse · ask

ask files a decision and refuses until it is ruled

reason

one sentence

a rule with no reason is a wall; a rule with a reason is an instruction

authority

mandate:<id> · decision:<id> · record:<id>

no agent grants itself a rule — a rule with no authority cannot be declared

locked

whether it moves under a lesser authority

a ratified rule is locked by default

since / until

calendar days

until is how a rule expires with nobody having to rule on it again

match

clauses over the call's own arguments (matches · notMatches · equals · prefix · anyOf · absent · isActor), ALL of which must hold

it is what lets a rule say "not gh repo create --public" rather than only "not the shell"

There is no lookaround in a pattern: the engine's regular expressions are linear-time by construction, which is a property worth keeping in a gate that runs on every write. "This, but not that" is therefore two clauses — matches and notMatches — which is also the form a person reading the rule can check. A pattern carrying (?!…) is refused by name, saying so.

Rules are immutable rows. A rule that must change is retired (a retirement row that names it) and declared again — two audited writes — so the question "what bound whom on the day of the incident" stays answerable.

The verdict, and the precedence law

guardrails.check({ actor, door, args, scope }) answers { verdict: 'allow' | 'refuse' | 'ask', rule?, reason, skipped[], digest }. It is pure, deterministic, and implemented ONCE — in Rust, in native/src/guardrails/evaluate.rs. The in-process API, brainy serve, the MCP door and the harness hook all reach that one function through the crossing profileWriteDoor already is, which is why they cannot drift.

  1. The most restrictive matching rule wins — refuse beats ask beats allow. Not "the most specific wins": specificity decides which rule is NAMED, never whether the call proceeds. A narrow allow must not punch a hole in a wide refuse, because the actor most motivated to declare that narrow rule is the one the wide rule binds.

  2. No matching rule is allow — stated loudly, because it is a decision. This engine is not a whitelist: a store with no rules behaves exactly as it did before this profile existed. The refusal a fresh agent meets comes from a rule that binds *.

  3. A rule that needs a name the door does not have STANDS ASIDE, counted. That is the isActor clause — "the author is not the caller" — and the subject forms that name an actor. Learned from the brainy host's own 0.6.366 incident (2026-09-14, rolled back in seven minutes): a rule whose subject names a specific actor cannot be evaluated at a door that knows only an account. It is skipped, NAMED in skipped[], and never silently dropped. The failure mode being avoided is not "too permissive" — it is "every write refused, for every party, because one substitution was unknowable at that door".

  4. A refusal on a READ door is a refusal, never an empty answer. An empty answer reads as "this brain knows nothing", which is a different fact.

Where it is enforced: inside the door

guardrails.check is called INSIDE add, relate, update, remove, the memory writes, the accords writes and every brain.vfs write verb — before the plan is decided, so a caller cannot skip it by not calling it. The actor is whoever the door already knows: the minted key's subject on serve/MCP, the caller the in-process API carries (the same field memory.turn() refuses without). On brainy host, every governed write reached over its HTTP routes or their MCP twins — the raw engine writes, import, the memory, accords and file-system writes, and the session doors' own writes — is enforced as the AUTHENTICATED credential that opened that door — never the pool-opened instance's own identity, which on a shared, multi-tenant host speaks for nobody; an in-process caller with no host at all is unaffected, and the service key is never the caller of a raw write or an import.

One check per call

A governed call is checked ONCE, at the door its caller named, as that caller. The writes that door makes inside itself — a file write's own add, a remembered memory's own rows, an accord round's own transact, an import's own rows — are part of the same call: they carry its verdict and its actor and are not checked again. So a rule is met at the door it names: a refuse on transact refuses a transact someone calls, never the commit beneath an accord round; a refuse on add refuses an add, never the row beneath a file write (name vfs_writeFile for that). import is checked under its own name, with its options (never the data being imported) as the call a rule reads and its vfsPath as the scope.

The verdict covers the call's own duration only. Work the call starts and does not wait for is a call of its own, checked as whoever makes it. A write with no caller at all, outside any checked call, is still refused loudly (GuardrailDoorUnreachableError, "context.caller is required") once a store holds a rule. With no rule on a store none of this crosses anything: an unruled store answers allow without a check.

The engine's own writes: system:ripple

The name is the sharp-wave ripple: the burst in which a sleeping brain replays the day and files what matters — the engine's own upkeep in miniature. The system: half says it is the engine, never a person or an agent.

The writes the engine makes on its own account — the memory jobs (decay, episodes, projections, consolidation, citation ranks, retirement), the language classify job's labels, the import deduplicator's merges and removals, recall's strengthen pass and co-recall edges (a read that writes), and the file system's root, created or repaired at open — are checked, and audited when refused, as the reserved actor system:ripple. A tier-1 rule binds it like anyone its subject matches. An owner's (tier-2) rule binds it only when its subject names system:ripple exactly: a rule written about people and keys — *, a role, one agent's key — was not written about the store's upkeep, so a fence on an agent's key never stops the store's own jobs. To bind the engine, name it. No key or token may carry the actor; one that does is refused by name (ReservedActorError). memory_rebuild_projections is checked as the caller who asked, at that name; the rows the rebuild writes are the engine's.

A brain's guardrails cannot forbid an operator ceremony (the orphan-id retraction and the backup, persist, today) in 13.0.0: a ceremony is reached under the service key through the engine-mechanics arm, which runs no guardrail check, so a rule never sees it.

The VFS write verbs — all ten, and what a rule sees

Every verb that mutates the file system is governed, under its own door name (vfs_<verb>) — the same names brainy guardrail check --tool vfs_<verb> takes over the harness hook below. The call a rule's match clauses and scope read is composed from the verb's own arguments, never its content:

verb

door

what a rule reads

writeFile

vfs_writeFile

{ path, options? } — the bytes being written are never part of the call

appendFile

vfs_appendFile

{ path, options? }

unlink

vfs_unlink

{ path }

mkdir

vfs_mkdir

{ path, options? }

rmdir

vfs_rmdir

{ path, options? }

rename

vfs_rename

{ path: source, paths: [source, destination] }

move

vfs_move

{ path: source, paths: [source, destination] }

copy

vfs_copy

{ path: source, paths: [source, destination], options? }

setxattr

vfs_setxattr

{ path, name } — the attribute's VALUE is never part of the call

removexattr

vfs_removexattr

{ path, name }

For rename/move/copy, scope.pathPrefix matches when any of the verb's paths is under the prefix — never the source alone — so a rule fencing /projects/anima refuses a move that lands a file inside it exactly as it refuses one that carries a file out. Each path is checked in the verb's own order (source, then destination): the first one a rule matches is the one whose refusal (or ask) the caller sees, and the call never proceeds either way. A move calls rename internally rather than going back through this door a second time, so it enforces once, as vfs_move — a rule naming only vfs_rename does not also bind a move.

Every refuse and every ask writes an audit row — actor, door, rule, generation, and the call's DIGEST rather than its arguments, because an audit trail read by whoever is investigating an agent must not become a second copy of what that agent was writing. guardrails.audit({ since, actor }) reads them. Audit rows, rule rows, retirement rows and ratification rows all carry visibility: internal and are fenced from ordinary recall exactly as board rows are — hidden from a default find(), visible with includeInternal. (Rows written before this fence's own fix carried the tier as an ordinary metadata key rather than the entity's own scalar; the fence recognizes both spellings, so such a row is hidden too, with no rewrite needed. brain.health() reports how many such rows a store still holds, in its guardrails-legacy-visibility check, counted when asked and never at open.)

The memory write doors — nine, and what a rule sees

Every brain.memory door that writes is governed, under its own door name (memory_<name>), the same #governNamespace wrapper the VFS table above describes. Two of the nine carry a full rule ceremony of their own — turn (M1-M5) and abstract (M6-M8), src/memory/rules.ts — and are declared in the contract; the other seven write just as really but enforce no profile-specific law, so they are named directly, not in the contract (src/memory/names.ts's MEMORY_WRITES_WITHOUT_CEREMONY; thread ONE-MEMORY-WRITE-LIST explains why a CONTRACT_PROFILE_DOORS entry is not possible for them). A rule fencing any of the nine by name refuses it before the door's own write runs.

Stated plainly because it is not obvious from the door's own shape: the call a rule's match clauses and scope read is the door's FIRST argument, exactly as shaped by the caller — never a merge of every argument. turn, abstract, remember and declare take one params object, so a rule sees every field of it. retire(id, options) and reinforce(id, options) take the id FIRST and the options bag second — a rule sees the bare id string, never reason, supersededBy, outcome or author, because those live in the argument a rule is not shaped from. cite(ids, options) is the same shape one step wider: a rule sees the bare ids array, never outcome or author. A rule that needs to read past an id or a list of ids is not something this leg's shaping supports — brain.vfs's own per-verb shapeArgs is the precedent for building one, a separate, wider change than declaring the doors.

door

mcp tool

what a rule reads

declare

memory_declare

{ shape? }

remember

memory_remember

{ content, type, subtype?, source, confidence?, id?, visibility?, author?, metadata?, relate?, vfsPath?, chunk?, boundaries? }

turn

memory_turn

{ conversationId, turn, speaker, gist, measuredTokens, by?, asked?, done?, open?, refs?, metadata? }

retire

memory_retire

the bare id string — reason/supersededBy/author are the second argument, never read

reinforce

memory_reinforce

the bare id string — recallId/outcome/by/author are the second argument, never read

cite

memory_cite

the bare ids array — outcome/by/author are the second argument, never read

abstract

memory_abstract

{ content, absorbs, source?, vfsPath?, tags?, confidence? }

rebuildProjections

memory_rebuild_projections

{} — no arguments; also see the note below

rank

memory_rank

{ pageSize?, writeEpsilon?, select? } or {}

rebuildProjections is checked at its own door, as the caller who asked, under memory_rebuildProjections — a rule may name it so or by its tool name, memory_rebuild_projections (the same door). The rows the rebuild writes are the engine's own housekeeping (system:ripple, above). consolidate, jobs and runJob are in-process doors with no wire twin; a consolidation's writes are housekeeping too.

When a service speaks for a person: the actor, and via

A multi-tenant service holds one key and serves many people. From 12.12.0 its requests may say who each call is for — X-Brainy-Caller, and the optional X-Brainy-Owner / X-Brainy-Roles beside it (docs/host.md's "The caller on the call"). The rule engine sees no header: it sees an ACTOR, and that actor is the asserted person. A rule written against person:ana@example.com binds a write a service made on Ana's behalf exactly as it binds one Ana's own key made, and role:<name> subjects resolve from the roles the service asserted for her rather than from the key's own (a service key has none).

Beside the actor travels via — service:<id>, the service that spoke. Where it is carried, and where it is not, stated plainly:

  • Carried on the credential through every write path, and on the refusal a rule raises: GuardrailRefusedError.via, which crosses the wire like every other field that class carries. So a refused write names both the person it was attributed to (actor) and the service that spoke for them (via).

  • Not carried on the stored AUDIT ROW, or on any stored accords or memory row. Those rows are written by the core, whose write context crosses into the native profile write door with a shape fixed in the engine binary; the row records actor and nothing about who asserted it. That is a real limit, not an oversight, and it is written here rather than left for a reader to assume a field that does not exist. Closing it is an engine change — a field on the profile write context — and belongs to whoever rules on the row shape.

A call that asserts nobody carries no via at all. An absent via and a via reading "nobody" are different facts, and only the first one is true of a credential that speaks for itself.

role:<name> subjects resolve from the credential

A role:<name> subject binds whoever's own credential carries that role — a minted bk1 token's optional roles claim (an array of non-empty strings), never a name the request supplied about itself. A token carrying no such claim, and every keys-file principal (the file format has no roles column), both hold [], which no role:<name> subject ever matches — absence changes nothing.

guardrail_check's and guardrail_binding's own roles ARGUMENT is a different, narrower thing: it is honoured only for the read or the simulation that specific call performs, and it never substitutes for the credential in that caller's own real enforcement. A visitor who calls guardrail_check({ roles: ["share:x"] }, ...) sees the verdict a holder of share:x would see, but their next write is still checked against their OWN token's roles — asking the question never answers it.

A role:<name> subject is still a NAMED subject — precedence law item 3 above ("a rule that needs a name the door does not have stands aside, counted") applies to it exactly as it applies to a plain actor slug. So an actor is required for it to be evaluable at all: guardrail_check supplies its own authenticated caller as the actor whenever a call omits one (the same default guardrail_binding has always used), which is what makes calling guardrail_check({ door, ... }) with no actor field still test a role:<name> rule against the credential's own roles, rather than skipping the rule for want of a name.

Rules over tools brainy does not serve

A rule's doors may name ANY tool the agent might call — Bash, web_search, github_*, an MCP tool from another server. Brainy enforces its own doors inside the engine. For foreign tools there are two adapters, and the evaluation is the same function in both.

(a) The harness hook — available now

brainy guardrail check --actor <slug> --tool <name> --args-json -   # stdinBASH

Exit 0 = allowed. Exit 2 = refused, with refused by rule <id>: <reason> on stderr. A coding-agent harness with a pre-tool hook blocks the tool call on exit 2 and shows the agent stderr — which is the whole mechanism: the agent is told the rule at the moment it matters, by something it cannot forget to consult.

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "*",
        "hooks": [
          {
            "type": "command",
            "command": "brainy guardrail check --brain $BRAINY_STORE --actor $AGENT_SLUG --tool \"$TOOL_NAME\" --args-json -"
          }
        ]
      }
    ]
  }
}JSON

Other harnesses have the same shape: a pre-tool command, stdin carrying the arguments, a non-zero exit that blocks — name the actor/tool environment variables however your own harness's hook convention supplies them.

(b) The gateway — designed here, built later

brainy serve fronts the client's OTHER MCP servers: the client connects to brainy, brainy holds the upstream connections, every tools/call is checked against the same rules, an allowed call is forwarded verbatim and a refused one comes back as a named refusal (never a protocol error). The hook governs one machine's harness; the gateway governs every client of one store, including harnesses with no hook mechanism at all. The design constraint that decides it: the gateway must forward BYTE-IDENTICAL arguments and must never see a credential it does not need — the upstream server's own auth stays between the client and that server wherever the transport allows it.

(c) The mandate slice — always

memory.context() returns a mandate slice FIRST, never budget-trimmed: the rules that bind this caller, by actor and scope. The door enforces them anyway. The slice exists so the agent is TOLD, every turn, what it is bound by — an agent that knows the rule does not spend a turn discovering it by refusal.

Who may make a rule: two tiers

TIER 1 — tighten-only, no approval. Any agent may declare rules that only RESTRICT (refuse or ask) and only bind ITSELF and the subagents it filed an action for. Effective immediately, audited, retireable by the declarer or the owner. A tier-1 rule can never loosen, override or lock anything — declare refuses each of those by name, pointing at propose.

TIER 2 — everything else needs the owner. An allow rule, a rule on another agent, a lock, an unlock, retiring a rule the agent did not declare, a rule outside its own scope: guardrails.propose writes an INERT pending row and files a decision on the owner's board (adopt · adopt with edits · refuse — a picker with exactly one recommendation, rung by the doorbell). The rule activates only when guardrails.ratify is called under the owner's own credential — a key whose scope includes owner, minted only by a human sign-in. Never by author slug: every session today shares one tenant key, and a slug can be typed by anyone.

Rules the owner ratifies are locked by default. Unlock is a further owner ruling that names the rule (guardrails.unlock, a new row, audited). until expires a rule with no ruling at all. And there is one kill switch: guardrails.freeze — an owner ruling that suspends every tier-1 rule at once, its audit row naming each rule it suspended.

Confining an agent's key to board rows: the participant fence

A participant key is an agent's key on a person's brain. participantFenceRules({ actor }) returns the three rows that say what that key may write, and declareParticipantFence(brain.guardrailsAs({ caller, owner: true }), { actor }) declares them in one call under the owner's own credential (over the wire the owner declares each row through guardrail_declare). Declared by the owner they are tier 2 and locked: the key they bind cannot retire, unlock or loosen them.

The key may write ONLY board rows: add, update or relate whose own metadata.collection is exactly board. Reads are untouched, and a person's own key is untouched, because every row binds one actor, never *.

row

doors

clause

effect

<actor-slug>-writes-board-rows-only-field-missing

add update relate

metadata.collection is absent

refuse

<actor-slug>-writes-board-rows-only-field-differs

add update relate

metadata.collection does not match ^board$

refuse

<actor-slug>-no-writes-beyond-board-rows

remove addMany updateMany removeMany unrelate updateRelation relateMany transact import clear memory_* vfs_*

none

refuse

There is no allow row, by the precedence law above: nothing is a whitelist, so an allow could not change a verdict. "Absent or different" is two rows because a clause over a field that is not there never holds. The three never match one call together; where two could, both refuse and the smaller id is named.

transact is refused whole. It is governed once, with the list of operations as its argument, never once per operation, so no clause can say "every operation carries the board field". A participant key writes one call at a time.

The accord_* write doors are not fenced: they write board rows by construction, and their inner commit carries the round's own verdict ("One check per call", above), so a participant key posts rounds. import is refused by the third row at its own door. The session doors' capture and save write a file, so they are refused at vfs_writeFile. Every other key on the store — the person's own included — writes through all of these doors as before: each row binds one actor, and none of them binds the engine's own housekeeping.

What happens while the metadata index rebuilds

A rule is a row, and the rules in force are composed from rows — so the rule check inside every write door needs to READ. A brain that opens with a derived index to rebuild answers reads from the generation it already holds and keeps serving writes throughout; but a brain that has nothing to answer a filtered read from — a store whose index must be built from scratch, one whose prior generation cannot be proven complete, one whose rebuild failed — refuses filtered reads by name for the length of that window. A rule check made of filtered reads would then refuse with them, and because the check fails closed, every write would be refused too. That is a write outage caused by an index that writes do not depend on, and this is how it is closed.

The rules are mirrored into one row, read by id. A read by id resolves straight from storage and consults no derived index, so it answers exactly when a filtered read cannot — and only then. The rows decide whenever they can be read: a mirror is derived, and a derived answer believed over its own source is how a rule set quietly loses a rule. So the mirror answers the window in which the rows refuse and not a moment more, and a rule check that reaches the rows also checks the mirror against them and has it rebuilt if it has fallen behind. Every rule write refreshes that row inside its own transaction — the rule and the mirror land together or neither lands, and a concurrent writer moves the store past the transaction's generation pin and refuses the whole plan. The mirror keeps the ROWS, and the composition that turns rows into "the rules in force" is the same function whichever source they came from: a mirror can never become a second, drifting evaluator. The rows stay the source of truth and the audit's record; the mirror is derived, and its staleness is measured against the last Guardrail write rather than the store's generation, so ordinary writes never make it stale.

So, while filtered reads are refusing:

  • A write is admitted and CHECKED. The rule check answers from the mirror; a rule that refuses the call still refuses it, with its audit row, exactly as it would at any other moment.

  • A store that has never declared a rule is checked against an empty rule set — "no declared rule binds this call", a verdict the brain verified against the rule rows when it opened. Such a store is given no mirror row at all: there is nothing to mirror, and a row would be a write on the first open of every store that exists.

  • A store whose rule rows were written before this release has its mirror composed from those rows by a background job, behind the doors, at the first open that can read them. A read-only open writes nothing — a reader has no write doors to gate.

  • If neither the mirror nor the rows can be read, the write is refused by name (GuardrailDoorUnreachableError), and the refusal names both halves and the cure. A write is never admitted unchecked, whatever the state of the store.

  • The READ doors still refuse while the rows they read cannot be served — guardrail_audit among them. A refusal on a read door is a refusal, never an empty answer (see the precedence law's fourth rule).

The doors

door

what it does

who

guardrail_check

evaluate one call; audit a refusal or an ask

anyone

guardrail_declare

declare a tier-1 rule — tighten-only, binds the declarer and its own subagents

any agent

guardrail_propose

write an inert rule and file the owner's decision

any agent

guardrail_ratify

activate a pending rule under the owner's ruling

owner credential

guardrail_retire

retire a rule (a retirement row)

the declarer, or the owner

guardrail_unlock

lift a lock, naming the ruling

owner credential

guardrail_freeze

suspend every tier-1 rule, naming each

owner credential

guardrail_audit

read the refusals and asks

anyone

guardrail_binding

read the rules that bind one actor, matched on subject exactly as guardrail_check resolves it

anyone (match only for the owner or the service key)

guardrail_rules

read every rule in force, optionally narrowed to a scope

anyone (match only for the owner or the service key)

What a rule shows over the network

guardrail_rules and guardrail_binding answer each rule as it is stored: its id, subject, scope, doors, effect, reason, tier, authority, declaredBy, declaredAt, status, locked, suspended, and since, until and decision when the rule has them (absent when it has none, never null). Every caller who may read a rule sees all of that. The one field they do not all see is match, the argument predicate: the exact patterns, regular expressions included, that decide whether a call is caught. It is carried only when the caller holds an owner-scope token or presents the service key; a service asserting a caller sees what that caller may see, so a service that asserts a non-owner reads the row without it, and a read with no credential at all (the stdio bridge) withholds it too. The in-process brain.guardrails.rules() / .binding() are unchanged and answer the whole row.

What this is not

It is not a sandbox and it does not contain a process. An agent with a shell and no hook can run anything the operating system lets it run; what guardrails guarantees is that every door brainy serves, and every tool routed through an adapter, is checked against rules the agent cannot forget, edit or outrun — and that every refusal leaves a row saying what was stopped, for whom, and under whose authority.