MuZero × MCTS × LLM
An idea is a proposal.
A rollout is evidence.
Connect grounded model advice to learned-model search. Train a small neural planner, inspect its decisions, and test both first actions in a real simulator.
No API key needed Optional local Ollama Review before action
Retrieve & propose
Inspect ranked task facts with source hashes. Optionally ask an installed local model for an action with exact quotations.
Train & search
Learn reward, value and policy from simulator transitions. See where tree search spends its visits and what it predicts.
Measure & review
Compare the proposal with search, actual action traces and four baselines. Download a JSON report with source hashes and measurements.
START LOCALLY · PYTHON 3.11–3.13
Your first experiment
From a repository checkout on Linux x86_64, install the hash-locked CPU profile and start the dashboard:
bash alpha_factory_v1/demos/muzeromctsllmagent_v0/install_and_launch.shThen open http://127.0.0.1:7862 and choose Retrieve, train & compare. First installation downloads dependencies. After a run, select Download JSON report. Failed runs clear previous results so you can correct inputs and retry. Stop the server with Ctrl+C.
Setup, commands & troubleshooting →A tiny task with an inspectable answer
MiniChoice offers an immediate reward or a two-step path. These are the environment rules, not claimed training results:
INITIAL ACTION 0
0.3 nowThe episode ends immediately.
INITIAL ACTION 1
1.0 laterReach the second state, then choose action 1. Choosing action 0 there instead yields −0.2.
The local app measures what the trained policy actually does. A model can be wrong. Repeated deterministic episodes do not demonstrate general intelligence or generalization.
Research preserved. Scope made explicit.
This page is a launch guide; neural training runs in the local app. No funds, wallet authorization, token settlements or external actions are involved.
Preserved browser illustration · Original research archive · Explore harder planning environments