META-AGENTIC TREE SEARCH / SEARCH LAB
Explore the paths.
Make the choice.
Let specialist policies compete. Follow each rewrite. Challenge the best workflow before it becomes a decision.
Editable models · Reproducible search · Independent review
Explore a search01SearchCompeting rewrites share a finite budget.
02ChallengeFreeze the design. Use separate workloads.
03ReviewExport the complete reasoning trail.
THE RESULT
A design worth examining.
The active workflow remains the baseline. Every candidate is unapproved.
THE SEARCH TREE
See how the choice emerged.
The map shows actual expanded nodes. The highlighted route received this iteration’s rollout score; dots show competing workflow designs.
Selection → expansion → rollout → one backpropagation per visited node.
Inspect node statistics and the rollout
| Stage | Selected model choice |
|---|
INDEPENDENT CHALLENGE
Improvement has conditions.
Paired, held-out workloads compare the frozen candidate with the baseline. These checks test supplied assumptions; they do not authorize deployment.
| Measure | Baseline | Candidate |
|---|
THE PROPOSED WORKFLOW
One choice per specialist.
| Stage / rewriting specialist | Baseline | Candidate |
|---|
Inspect a held-out workload and its schedule
| Stage | Pool / lane | Start → finish | Rework |
|---|
TAKE THE EVIDENCE WITH YOU
A result you can reproduce.
Download the scenario, full search tree and trace, unapproved policy proposal, review brief, unsubmitted $AGIALPHA job, and checksums.
A SHA-256 fingerprint identifies the exact run.
Import recomputes every result. A matching checksum alone is not accepted as evidence or validator approval.
Edit the workflow model and review gates
Change the resource pools, ordered dependency graph, model choices, objective or review limits. At most 1,024 complete designs are supported. All durations, defects and costs are synthetic inputs.
WHAT THIS DEMONSTRATES
Search with an accountable trail.
Each meta-agent is a named, bounded rewrite operator over another stage’s policy. A fixed-point UCT rule balances exploration and observed utility. The simulator schedules dependencies and finite resource pools, propagates faults, and charges for detected rework.
This is a synthetic planning laboratory. It does not execute customer work, train language models, prove real-world reliability, or submit transactions. A passing candidate still needs independent, authenticated review and representative real measurements.