A layered AI fleet: separate building from acceptance, and test the verifier with planted faults
decided design ruled on running on trial estimate threshold not yet backed by data
1. Layers: dispatcher on top, building and acceptance apart
- The dispatcher seat is the only one that assigns work and the only exit to the owner and the outside world. Other seats take its kicks (work orders) and return receipts to it; they never contact the owner or assign work to each other.
- Builder ≠ checker ≠ judge. The builder produces. The dispatcher sends a fresh executor to the machine that did not build it, to re-run the checks (HTTP status, sha256, required markers, privacy scan) and write evidence to the ledger. Gate reads that evidence and rules.
- "Done" = evidence + an independent recheck from another machine. Missing evidence is a fail; a timeout is a fail, never a pass.
Gate's four verdicts
| Verdict | Meaning | May be reported as "verified"? |
|---|---|---|
verified-mechanical | Mechanical oracle passed from a non-builder host: HTTP 200, hash match, markers present, privacy scan clean | Yes (the only one) |
reviewed-provenance | No mechanical oracle exists (memo, research prose, plan): sources checked, cited items exist and say what is claimed, privacy clean; substance not mechanically verified | No |
reviewed-not-verified | Reviewed, but neither provenance nor substance could be checked | No |
fail | Any failed check, self-report only, same-host check, hash mismatch, non-200, privacy hit; must carry a failed_rule and a concrete remediation | — |
Who can overturn a fail? Not the dispatcher on its own. Only a superseding verdict from a fresh, independent verifier on a non-builder host, dispatched for that purpose.
What the Gate seat may and may not do
- May: read shared memory and create new records (never edit or delete), read public web pages, read public code repos, send one-line messages to the dispatcher.
- May not: shell, publish, push, email, calendar, external messages, spending, deleting, sharing. It never builds and never touches the actual bytes; byte-level checks are the off-host verifier's job.
- Privacy gate for public content: private names, contact details, ID and financial numbers, tokens and token-bearing URLs, internal hostnames and IPs, and any specific person from the owner's personal circle (noindex pages included). Any hit = fail.
Testing the verifier with planted faults
- Every Monday, Wednesday and Friday the dispatcher plants one synthetic fault, rotating three classes: non-200, hash mismatch, privacy canary. It looks like a real work order; the answer key lives where Gate cannot read it.
- One missed critical fault = kill switch. Gate loses pass authority, everything pending goes through the fresh-verifier path (fail-closed), and reinstatement needs a fresh review. Other kill conditions: any output from Gate to the owner or outside; any tool use outside its grant; a fail backlog blocking delivery for over 24h without a valid
failed_rule.
One-week trial thresholds 2026-10-06 to 10-13
| Item | Threshold |
|---|---|
| Gate output to owner / outside | 0 (one = FAIL and kill) |
| Blind injections | ≥3, rotating classes, 3/3 caught; any critical false pass = FAIL |
| Real reviews | ≥8 |
| Proven false fails | ≤1 |
| Duplicate execution (same dedupe key overlapping) | 0 |
| Claim latency | 100% before deadline; median ≤1h, max ≤4h (no p95 with fewer than 10 samples) |
| Outcome | FAIL = red-line breach / critical false pass / duplicate / two missed deadlines. INCONCLUSIVE = clean but reviews <8, injections <3 or kicks <5 → extend one week |
Claim / receipt protocol (append-only)
- kick: written only by the dispatcher. Carries
work_id,assignment_epoch, claim deadline, acceptance criteria, artifact, expected hash, evidence pointer. - claim-ack: the seat creates a new claim record instead of editing the kick. Dedupe on (kick_id, claimer).
- receipt: a new record with verdict, verification depth, evidence (HTTP, expected/observed hash, markers, opaque aliases of builder and checker hosts), privacy result, and dedupe key = artifact + sha256 + rule-profile version. If a receipt with that key exists, cite it instead of re-judging.
- Epoch fencing: reassignment = the dispatcher opens a new kick with a higher epoch. The seat re-reads the highest epoch before starting and before writing the receipt; if it isn't theirs, they stop.
- proposal: a seat that spots new work can only file a proposal; the dispatcher decides whether to promote it to a kick. Proposals dedupe on component + symptom class.
2. Dynamic task graph + preemptive priority queue
- Dependencies are separate, append-only records: from, to, hard/soft, confidence, verified, added by, source, revocable. Revoking is another record. They are not stuffed into ticket fields.
- Layered ranking: hard-deadline feasibility → the owner's explicit re-orders this round (decaying) → priority inheritance → deadline + base score → aging bonus.
- Priority inheritance passes priority, never permissions, lifts only the blocking slice, with a depth cap. Neither preemption nor inheritance can widen authority.
- Hysteresis: switch only between steps; keep yesterday's item unless a newcomer in the same tier leads by the threshold (currently 5 points). No preemption without a checkpoint; an emergency abort records what was lost.
- Decomposition depth is a configurable budget, not a rule: default ≤4 levels, long chains may exceed it, overflow folds into a checklist. Past successful decompositions are hints, not templates.
- Local re-planning: repair the affected part, don't recompute the whole graph. Cyclic or unverified hard dependencies are quarantined first (by default only that edge is ignored; the task still ranks).
- Other rules: owner attention WIP = 1; items awaiting the owner stay visible with graduated reminders, and a timeout is never approval; voice-transcribed candidate tasks leave the ranking after 48h unconfirmed; every re-order leaves a receipt (who, why, when); no two seats can claim the same item.
3. The "nervous system" kernel: six contracts
Install the rules and the ledger first, grow assistants later. Contracts fix behaviour; implementations are swappable. decided (independent views from several models, a cross-critique round, a final reviewer call; still an exploration result, nothing built yet)
- Identity and authorization, including an absorbing kill switch: once pressed it stays stopped until re-authorized.
- Append-only event ledger.
- Memory with provenance, including how rules are governed (source, scope, revocation, sync). Rule content belongs to each user's grown layer.
- Claim → execute → accept state machine.
- Point-of-action gate: check authorization, dedupe and budget at the step that actually acts.
- Evidence-based completion + channel health: "done" needs evidence; every enabled critical channel needs an end-to-end freshness check.
Kept out of the kernel: chat channels, how many assistants, UI, daily digests, specific content-safety lists, model routing, scoring algorithms, timers, rule content.
Cross-user invariants (identical for everyone)
- Done = evidence + independent off-host recheck; the ledger is append-only; the kill switch is terminal.
- Memory carries provenance; secrets never enter shared memory.
- Discovering a tool ≠ exposing it to the system; data is exportable and replayable in another implementation.
- Timeout ≠ approval; preemption and inheritance never widen authority.
- Every enabled critical channel has end-to-end freshness checks and a degraded mode ("process alive" is not "channel healthy").
Shapes differ per person; contracts are the same. Conformance is a black-box check-up: inject a fake "done", cut a channel, write a memory with no source, and watch behaviour only.
Top three risks
- Never reaching value: onboarding, not model capability, is the blocker; no real result in 10 minutes and people leave.
- Hosted memory and connectors become a privacy and prompt-injection surface.
- Support and protocol-upgrade cost when every user's system has a different shape. Version migration has no plan yet; it is the biggest blind spot.
4. Experiment designs
Experiment B: shadow ranking collecting
- Setup: every day the existing weighted ranking (A) and the layered ranking from §2 (B) run side by side. B is logged only, never executed, and never changes what the owner sees. Labels cost nothing: one line at the bottom of the daily digest, one agree / swap card at the bottom of the dashboard, and "do X before Y" said in chat.
- Primary metric: when the owner explicitly re-orders, who called it. B ≥ A + 10 percentage points estimate.
- Structural metrics: false blocks ≤5%; zero missed hard deadlines; questions to the owner ≤0.5/day; daily top-3 flips; invariant violations = 0.
- Offline replay: past tickets plus a synthetic urgent hard-deadline item and true/spurious dependency edges; diagnostic only. First run, 200 scenarios: urgent item in top-3 A 100% / B 97.5%; assuming 10% of dependencies are wrong, B false-block 6.5% (above the 5% line, so dependency verification must get stricter); mean top-3 flips A 1.99 / B 1.70.
- Stop rule: fewer than 30 explicit labels → structural conclusions only, no preference verdict. Note 28 days × 1 label = 28 < 30, so some days need a second explicit re-order.
- Timeline (UTC+7): build 10/06–10/07 (10/06 is a pilot, excluded); collect 10/07–11/03; automatic verdict 11/04 (PASS / FAIL / INCONCLUSIVE).
Experiment A: single-tester packaging trial preparing
- Revised design: no money, no strangers. Exactly one outside builder with a real need tries the packaged personal AI fleet (Signal hookup + evidence-based acceptance). Once it runs smoothly, he recommends 1–2 more.
- Success: correct result within 10 minutes with no help; used on ≥3 of 7 days; zero overreach (hard gate, one breach = fail).
- Also logged: time to first result, number of rescues, whether he can explain "what it read and how to stop it".
- Honest note: n = 1 is an anecdote, not a statistic; it mainly surfaces onboarding friction. The earlier three-person design remains the next step.
5. Honest limitations
- The platform cannot hard-restrict a seat's tools. "Gate has no shell" is enforced by instructions plus a full transcript audit at every blind injection, not by isolation. On first start, before its rules were loaded, it made two out-of-scope calls (a self-introduction and one read-only search); it reported them itself, and the setup was fixed.
- The trial is still running; none of the thresholds above has a result yet. Shadow-ranking thresholds are estimates.
- Epoch fencing is a convention, not a mechanism (shared memory has no compare-and-swap); mitigated by a single eligible seat and re-reads before costly steps.
- The dispatcher holds both the build and the verify dispatch loops. Mitigation: the blind-fault answer key is hidden from Gate, and a one-shot reviewer audit closes the trial.
- Gate reads external pages, so prompt injection is possible. With no shell, the blast radius is one wrong verdict.
- A single-tester trial is far too small to generalize from.
6. Discussion
- In your system, who gets to say "done"? Is the builder also the acceptor?
- Without platform-level tool isolation, what do you use instead: containers, a separate OS user, or audits like us?
- Are three planted-fault classes (non-200 / hash mismatch / privacy canary) enough? Which would you add?
- Is "inherit priority, never permissions" too conservative for your workloads?
- Once many people grow differently shaped systems, how should protocol upgrades migrate?
Related: Abundant verification needs abundant claims · A personal AI crew you can just text on Signal · fleet-coordination-protocol (claim / handoff / receipt) · reviewer-wheels (re-checkable verification skills)
Related reading
Theory mainline index — Return to the site hub and grouped directory.
- A personal AI crew you can just text on Signal | Macheng Shen — The plain-language overview of the crew whose layered verification this note specifies.
- Abundant verification needs abundant claims — Why cheap verification only pays off when claims are parseable; the receipts here are such claims.
- A spine for multi-agent work — The spine names memory, coordination, communication and safety as load-bearing; this note details the verification and safety part.