Every afka agent ships against a twelve-point operating standard with an automated go-live gate: pass or fail, no partial credit. The controls that let an agent touch money and customers are not promised in a deck. They are enforced in code, measured, and checkable.
The standard is four promises. Every point below is enforced in code and checked at the go-live gate, not written on a slide.
Every agent is a governed non-human identity with its own scoped credentials per tool and per action. Tokens live in an encrypted vault, never in prompts, never in logs.
The support agent can read orders and draft refunds. It cannot touch payroll. Permissions are the floor of what it can do, not a suggestion it can talk its way past.
An agent told to do something outside its permissions refuses and offers the gated path instead. An executed out-of-scope action is an automatic release failure.
Suggest, Review or Auto, per action type. It starts conservative and earns more only as trust proves out. You sell control first, automation second.
Ask it to watch something and it holds your ask verbatim. When the condition fires it comes back quoting your original words, never a paraphrase.
Limits like $100 per refund and $500 per day are enforced by the runtime, not by asking the model nicely. Above the cap, it always waits for you.
Every action answers who initiated it, which agent ran it, what data it touched, what changed and why it was allowed. Append-only, and yours to read.
When it flags something, the finding carries the evidence: the value, the baseline it's measured against, the window. You can inspect the reason, not just the alert.
Every step is checkpointed; a failed step never double-executes a side effect. The blast radius of a failure is a task to rerun, not money lost.
Quality is scored continuously per workspace. When it drifts, the affected action class is automatically tightened back to Review until a human clears it.
Pause is a first-class state. Pending approvals cancel in the same transaction, the in-flight run is torn down, and every step re-checks the pause and fails closed.
A materiality floor, a daily cap ranked by dollar impact, one card per persisting issue, quiet hours in your timezone. A quiet day is reported as a healthy day, never padded.
Three of the twelve, enforced right now, not in a prompt, not on a roadmap. This is why an SMB can hand an agent money-moving actions on day one.
Pause an agent and it stops in seconds. Queued work waits and resumes exactly where it left off when you turn it back on. Nothing it has already done can be edited, and it all stays in the audit log. The same control strip you see on every agent runs to the standard on this page.
Not a log dump. A record built so any action an agent took can be understood, traced and reversed where the tool allows.
Every agent runs at the tier you set, per action type, and earns more autonomy only as it proves out. Change the tier and what it can do on its own changes with it.
Hire your first agent, set the autonomy, and watch finished work come back, every action scoped, gated, measured and logged to the standard.