Owner Command Deck – Planetary Orchestrator Fabric
The owner wields absolute authority over the fabric. This manual demonstrates how non-technical operators execute pause/resume, reroute workloads, and adjust shard policies without modifying code.
Golden Rules
- Single source of truth: Only the owner multisig defined in
config/fabric.example.jsoncan push changes. - Declarative commands: All interventions are JSON payloads replayed via
scripts/v2/or the orchestrator owner console. - Auditability first: Every command writes to
reports/<label>/owner-script.jsonandevents.ndjson.
Control Surface
View original Mermaid source
stateDiagram-v2
[*] --> Live
Live --> Paused: owner:pause-all
Paused --> Live: owner:resume
Live --> Throttled: owner:set-latency-budget
Throttled --> Live: owner:restore-latency
Live --> SpilloverEscalated: owner:reroute-shard
SpilloverEscalated --> Live: owner:normalize-shardCore Commands
Pause Everything
npm run owner:system-pause -- \
--network sepolia \
--pause-contract 0xSystemPauseAddress \
--reason "Drill: Helios spillover protection"The demo generates an equivalent payload inside reports/<label>/owner-script.json under "pauseAll" so you can replay it verbatim.
Update Latency Budgets
node scripts/v2/ownerControlSurface.ts \
--action set-latency-budget \
--shard mars \
--latency-ms 350 \
--multisig 0xFABR1C00000000000000000000000000000000MSThe orchestrator instantly applies the change and persists the update inside summary.json.
Force Spillover to Helios
node scripts/v2/ownerControlSurface.ts \
--action reroute-shard \
--source earth \
--target helios \
--reason "Earth backlog > 90%"Register Surge Shards
{
"type": "shard.register",
"reason": "Provision edge spillway during surge",
"shard": {
"id": "edge-surge",
"displayName": "Edge Surge Lattice",
"latencyBudgetMs": 140,
"spilloverTargets": ["earth", "luna"],
"maxQueue": 1800,
"router": {
"queueAlertThreshold": 1200,
"spilloverPolicies": [
{ "target": "earth", "threshold": 1350, "maxDrainPerTick": 60 },
{ "target": "luna", "threshold": 1500, "maxDrainPerTick": 40 }
]
}
}
}- The shard comes online immediately. Jobs can target it in subsequent ticks, and checkpoints persist the new topology.
- Routers in neighbouring regions honour the new spillover routes instantly—no restart required.
Retire Shards Without Downtime
{
"type": "shard.deregister",
"shard": "edge-surge",
"reason": "Consolidate capacity after surge",
"redistribution": { "mode": "spillover", "targetShard": "earth" }
}- Spillover mode re-queues pending work into the chosen target shard. The ledger records new
job.spilloverevents, and dashboards illuminate the transfer. - Switch
modeto"cancel"with an optional"cancelReason"to evaporate outstanding jobs instead of reassigning them—handy for decommission drills. - Summary metrics (
jobsCancelled,spillovers, ledger totals) and mission atlases update the moment the command executes so auditors can verify the retirement instantly. - Retirement removes incoming spillover targets/policies and retires nodes assigned to that shard, keeping checkpoints restorable. Re-register those nodes in an existing shard to reuse them. The final shard cannot be removed; pause the fabric instead. Invalid spillover destinations are rejected before changing any jobs.
Command the Checkpoint Engine
Trigger an immediate checkpoint when you want a golden snapshot before governance actions:
node scripts/v2/ownerControlSurface.ts \ --action checkpoint-save \ --reason "Archive state before Helios redeployment"Retarget checkpoint cadence & storage without restarting the orchestrator:
node scripts/v2/ownerControlSurface.ts \ --action checkpoint-configure \ --interval 20 \ --path demo/Planetary-Orchestrator-Fabric-v0/storage/checkpoints/mainnet-governance.json \ --reason "Mainnet launch hardening"The orchestrator immediately persists the new settings and records them under
ownerState.checkpointand the nextcheckpoint.jsonartifact.
Retarget Reporting Archives
Direct reports to fresh destinations mid-run without redeploying tooling:
{
"type": "reporting.configure",
"reason": "Route artifacts to governance bucket",
"update": {
"directory": "demo/Planetary-Orchestrator-Fabric-v0/reports/governance-archive",
"defaultLabel": "governance-drill"
}
}summary.jsonupdatesownerState.reportingso dashboards and auditors immediately see the new directory/label.- Checkpoints persist the new settings; subsequent resumes keep writing to the chosen archive without manual edits.
Reroute Specific Jobs
node scripts/v2/ownerControlSurface.ts \
--action job-reroute \
--job-id job-00123 \
--target helios \
--reason "Redirect precision workload to Helios GPU array"The orchestrator removes the job from its existing queue (or the active node), rewrites its spillover history, and pushes it into the target shard router.
summary.jsoncaptures the reroute underowner.job.rerouteevents alongside updated spillover metrics.Prefer the declarative job locator when you want the fabric to pick a live target automatically:
{ "type": "job.reroute", "locator": { "kind": "tail", "shard": "mars", "offset": 8, "includeInFlight": true }, "targetShard": "helios", "reason": "Owner escalated to Helios precision array" }The locator walks the shard queue from the tail so the newest Mars jobs (or, if needed, their in-flight siblings) are selected deterministically—even when the total job count changes between runs.
Cancel Redundant Jobs
node scripts/v2/ownerControlSurface.ts \
--action job-cancel \
--job-id job-00124 \
--reason "Owner resolved ticket manually"The job is removed immediately, marked as cancelled, and surfaced in metrics (
jobsCancelled) for transparent auditing.Event logs emit both
owner.job.cancelledandjob.failedentries so replay systems retain determinism.To cancel without memorising identifiers, pass a locator payload such as:
{ "type": "job.cancel", "locator": { "kind": "tail", "shard": "earth", "offset": 4, "includeInFlight": true }, "reason": "Owner resolved ticket manually" }This grabs the freshest Earth job (or a still-running one) so operators stay hands-off even during massive job floods.
Resume From Checkpoint
- Stop the orchestrator (or simulate a crash with
Ctrl+C). - Run
demo/Planetary-Orchestrator-Fabric-v0/bin/run-demo.sh --resume --checkpoint demo/Planetary-Orchestrator-Fabric-v0/storage/checkpoint.json. - Confirm the resume log prints
"checkpointRestored": true.
Automated Restart Drill
Launch the full stop/resume rehearsal in one command:
demo/Planetary-Orchestrator-Fabric-v0/bin/run-restart-drill.sh \ --label "owner-drill" \ --stop-after 180 \ --owner-commands demo/Planetary-Orchestrator-Fabric-v0/config/owner-commands.example.jsonThe script halts the fabric at tick
--stop-after, extracts the new checkpoint path fromsummary.json, and resumes automatically.Inspect
reports/owner-drill/summary.json.runto verify{ "checkpointRestored": true, "stoppedEarly": false }after completion.
Customizing for Production
| Control | Demo Implementation | Production Hook |
|---|---|---|
| Pause / Resume | owner:system-pause script |
Gnosis Safe transaction or multisig contract call |
| Latency Budgets | JSON patch in orchestrator runtime | On-chain parameter update via SystemPause.executeGovernanceCall |
| Spillover | Deterministic router rule update | Emission of governance event consumed by routers |
| Thermostat | JSON payload in owner-script.json |
owner:update-thermostat script hitting RewardEngineMB |
| Node Quotas | Config patch + heartbeat filter | Kubernetes/HashiCorp Nomad autopilot APIs |
Scheduling Commands Ahead of Time
Load declarative schedules by passing
--owner-commands demo/Planetary-Orchestrator-Fabric-v0/config/owner-commands.example.jsontobin/run-demo.sh. The orchestrator will automatically execute each entry at the specified tick.The schedule format mirrors
OwnerCommandpayloads. Example excerpt:{ "tick": 155, "note": "Boost Earth queue budget to handle backlog", "command": { "type": "shard.update", "shard": "earth", "update": { "maxQueue": 6400, "router": { "queueAlertThreshold": 3600 } } } }After each run,
owner-commands-executed.jsonreports which commands executed, which were skipped (e.g., resumed from checkpoint), and which remain pending.Schedule payloads may include
checkpoint.save(instant snapshot) andcheckpoint.configure(update interval/path) entries; CI validates both commands execute cleanly.
Direct Owner Command Payloads
owner-script.json now emits two complementary blocks:
Legacy payloads (
pauseAll,rerouteMarsToHelios,adjustLatency,resumeAll) for existing on-chain command surfaces.Direct owner commands aligned with the orchestrator’s runtime (
system.pause,shard.update,node.register, etc.) so operators can patch configuration mid-flight. Example:{ "directOwnerCommands": { "registerHeliosBackup": { "type": "node.register", "reason": "Spin up backup GPU helion", "node": { "id": "helios.solaris-backup", "region": "helios", "capacity": 20, "specialties": ["gpu", "astronomy"], "heartbeatIntervalSec": 9, "maxConcurrency": 12, "endpoint": "https://helios-backup.fabric.ops/api", "deployment": { "orchestration": "kubernetes", "runtime": "cuda", "image": "registry.agi/helios-gpu-worker:1.4.0", "version": "1.4.0", "entrypoint": "/opt/agi/run-helion.sh", "resources": { "cpuCores": 24, "memoryGb": 96, "gpuClass": "A100", "storageGb": 320 } }, "availabilityZones": ["helios-gateway-a", "helios-gateway-b"], "pricing": { "amount": 0.00094, "currency": "USDC", "unit": "job" }, "tags": ["gpu", "backup"], "compliance": ["SOC2-Type-II", "Solaris-Safety"] } } }
}
Use these payloads with `run-demo.sh --owner-commands`, or apply them interactively by piping JSON into your governance tooling.
## Safety Net Checklist
- ✅ `owner-script.json` contains the exact command payloads for replay.
- ✅ `summary.json` echoes the updated policies so auditors can verify intent vs effect.
- ✅ `events.ndjson` logs the change with timestamp, owner signature, and hash.
- ✅ `checkpoint.json` records the new state so restarts never revert policies.
Use this manual to keep the owner in full command across planetary distances without touching a code editor.