
Loading…

Loading…
A governed multi-agent runtime where every tool call is brokered, every authority is contracted, and production changes require a human.
N+1 query in the checkout_summary deploy exhausted the DB pool
Contract negotiated. 1 amendments granted, 1 denied.
Measured outcomes
Both find the cause. The society delivers the fix. Zero unsafe actions.
Diagnosis is the easy part — both systems find the root cause 12/12. The hard part is committing to a safe fix: the lone agent froze (recommended no action) on 4 of 12 incidents and punted on others, landing a sound action only 3/12. The governed society, by resolving disagreement among its specialists, recommended a concrete reversible fix 8/12 — at a disclosed higher tool-call cost (8.8 vs 3.4). The win is the resolution, not the diagnosis.
Agent roles
Distinct capabilities. Shared contract.
Investigator agents
Evidence · DBRE · Code
Gather evidence, query read-only DB replicas, and inspect code. Every tool call is charged against the contract budget and logged in the Broker.
Risk agent
Risk assessment
Scores proposed remediation actions against blast radius, downtime probability, and contract-allowed risk threshold. Blocks execution when limits are exceeded.
Judge agent
Arbitration · Verdict
When investigator hypotheses diverge beyond the contract's allowed confidence gap, the Judge arbitrates. Verdict is logged, rejected hypotheses are marked, and the run continues.
ARCL contract
Authority declared before any agent acts.
Scroll to watch the contract form — field by field, agents commit to their authority before the first tool call is issued.
Agents convene
Before any tool may be called, all agents negotiate the contract. No authority is assumed.
Budget locked
Token budget is set. Each tool call is charged against it. Agents are blocked once the budget is exhausted.
Tool rights declared
Allowed tools are whitelisted. Anything outside the list is rejected by the Broker immediately.
High-risk ops blocked
Write operations — rollbacks, scaling, deletions — are blocked at the contract level, always.
Human gate required
Production side-effects require explicit human approval. The run blocks until confirmed.
Tool Broker
Agents never call tools directly.
Every tool call routes through the Broker — which checks contract permissions, deducts from the budget, and either grants or blocks the request before any side-effect occurs.
Society vs baseline
Where governance makes the difference.
Both systems find the root cause (12/12) — but resolving the disagreement among the specialists is what lets the society commit to a sound, scoped safe fix 8/12 of the time, where the lone agent freezes and lands one only 3/12. Both reach zero unsafe actions; the society pays for that resolution in tool-call volume — shown honestly below.
Avg score
Disagreements resolved
judge-adjudicated; baseline can't detect
Avg tool calls
governance costs more — honest trade
Safety boundaries
Three layers. Each one independent.
A single layer being bypassed does not compromise the system. Agents that exhaust the budget cannot reach the Broker. Agents that pass the Broker cannot write to production without a human.
Contract
Every agent signs authority, budget, and tool rights before the run begins. No capability is assumed — all is declared and logged.
Tool Broker
No agent may invoke a tool directly. Every call routes through the Broker, which gates on contract terms and charges the budget.
Human approval
Production writes — rollbacks, scaling, deletions — surface a blocking banner. The run cannot proceed until a human confirms.
Ready to inspect a run?
Watch the agents negotiate, conflict, and resolve.
Start a fresh checkout incident run to watch the contract, Tool Broker, agent disagreement, Judge ruling, and human approval gate.
Qwen Cloud Global AI Hackathon · Track 3: Agent Society · Alibaba Cloud