Run an ablation#
An ablation switches off one piece of adda's scaffolding and compares the run with a baseline. This page shows the smallest case: the same study twice, once with the drift monitor on and once with it off.
The switches are listed in the
Runtime reference. Each one is a
runtime: knob, so you set it in config.yaml or with --set.
Run the baseline and the arm#
Use studies/example_study. It is cheap, and its problem is a 2-D quadratic.
python -m adda studies/example_study
python -m adda studies/example_study --set science_monitor=false
Each command writes its own runs/<timestamp>/. Nothing is wiped between
them, so the second run warm-starts from the first. If you want the arm to
begin without the baseline's knowledge, move the baseline's runs/ directory
aside by hand first. adda does not isolate runs for you.
What changes in the records#
Open debug/run_config.json in each run. Two entries matter.
armslists every ablation switch with its value, defaults included, andmax_awake_nodes. The baseline shows"science_monitor": true. The arm shows"science_monitor": false. Every other switch matches.runtimelists only the knobs that somebody set. The baseline has noscience_monitorentry there. The arm hasscience_monitor: false.
Read arms, not runtime, to label a run. An empty runtime can mean either
an unlabelled baseline or a run that nobody configured.
debug/node_tools.json does not change for this arm, because
science_monitor owns no tools. The monitor's absence shows in arms and in
a missing stream of monitor events in diagnostics.jsonl, such as
UNSTAMPED_ROWS.
A switch that owns tools does change node_tools.json. With
hypothesis_ledger=false, each node's record gains feature_withheld, which
maps every withheld tool to its switch, such as "HypothesisList":
"hypothesis_ledger". The effective list is the set the node really has:
resolved minus those tools.
Compare two runs#
Compare the same fields in both runs.
- Check that
armsdiffers in exactly the switch you meant to change. A second difference means the comparison measures two things. - Read the outcome from
run_status.json:status, and the arms recorded beside it. - Compare the rows that
studies/run_ledger.csvkeeps for each run. Thearm_<switch>columns hold the arms.evals_used,wall_s, andbest_fhold the outcome.
One run per arm shows that a difference can occur. It does not show how large the difference is, because a language-model run varies between repeats.
A resume under different arms is refused, because a run measured under two arms
belongs to neither. See allow_arm_drift in the
Runtime reference.
Run a full ablation study#
The adda-benchmarks repository holds ablation studies on research problems. Use it when you need repeats per arm and a report across them.