Restart Drill – Planetary Orchestrator Fabric
The restart drill demonstrates that AGI Jobs v0 (v2) lets a non-technical mission director halt the orchestrator mid-flight, resume from the latest checkpoint, and continue without losing jobs or telemetry.
Two-Stage Flow
- Stage One – Controlled Halt
bin/run-restart-drill.shcallssrc/index.tswith--stop-after-ticks <N>. Jobs are seeded across shards, owner commands execute as scheduled, and at tickNthe orchestrator writes a checkpoint and stops.events.ndjsongains asimulation.stoppedentry capturing tick, directive, and outstanding queues.summary.json.runrecordsstoppedEarly: trueandstopReason: "stop-after-ticks=<N>".
- Stage Two – Resume
The script parsessummary.jsonto discover the active checkpoint path (even if the owner retargeted it during stage one) and restarts the orchestrator with--resume.- The resume run appends events to the existing stream.
summary.json.runnow shows{ checkpointRestored: true, stoppedEarly: false }.ownerCommands.skippedBeforeResumelists commands executed before the halt.
Command Reference
# Full drill with defaults
./demo/Planetary-Orchestrator-Fabric-v0/bin/run-restart-drill.sh
# Custom label, faster halt, alternate schedule
./demo/Planetary-Orchestrator-Fabric-v0/bin/run-restart-drill.sh \
--label "edge-drill" \
--stop-after 120 \
--jobs 8000 \
--owner-commands demo/Planetary-Orchestrator-Fabric-v0/config/owner-commands.example.jsonBehind the scenes the drill forwards the following flags to the TypeScript entry point:
--stop-after-ticks– positive integer; determines when the orchestrator halts.--preserve-report-on-resume– ensures reports persist andevents.ndjsonis appended rather than replaced.--checkpoint– supplied only during the resume phase with the path extracted fromsummary.json.
Verifying Success
- Open
reports/<label>/summary.jsonand confirm:run.stoppedEarlyistrueafter stage one andfalseafter stage two.run.stopTickmatches the tick from the drill.metrics.jobsCompletedequalsmetrics.jobsSubmittedafter the resume.
- Inspect
reports/<label>/events.ndjson:- A
simulation.stoppedevent appears exactly once. - Subsequent events show resumed processing and checkpoint saves.
- A
- Open
demo/Planetary-Orchestrator-Fabric-v0/ui/dashboard.html, drop thereports/<label>folder, and inspect the auto-rendered shard tables, owner command ledger, and spillover map to visualise pre/post drill topology. - Review
reports/<label>/owner-commands-executed.jsonto audit which commands ran before and after the restart.
Production Hardening Tips
- Rotate labels per drill (
--label my-drill-$(date -u +%Y%m%d%H%M)) so historical runs remain available for auditors. - Store checkpoints on durable, access-controlled storage; the demo defaults to
storage/checkpoint.jsonbut owner commands can retarget to any path. - Couple the drill with alerting: trigger notifications on the
simulation.stoppedevent so SRE teams know the halt was intentional.
The restart drill proves that AGI Jobs v0 (v2) behaves like the superintelligent orchestrator operators expect—capable of pausing an entire planetary workload and resuming without missing a beat.
Current recovery boundary
The shell launcher resolves the repository from its own location and requires jq. Stage two uses --finish to clear an inherited mission-plan stop limit. Missing checkpoints fail instead of starting a new mission. TypeScript snapshots now require format version 1 and a valid SHA-256 digest; earlier unsigned snapshots must be preserved separately and are not silently migrated. Atomic replacement protects against partial file writes, not every storage/power-loss failure. Use one process per checkpoint. These are controlled simulation shutdowns, not evidence of live-provider or paid-settlement recovery.