Skip to content

Runtime reference#

This page lists every key of the runtime: block in a study's config.yaml. For the other keys, see the Configuration reference.

The runtime: block holds everything that tunes how the run executes, as opposed to what it is asked to do. Every knob has a working default, so the block is optional. Set a knob only when you have a reason to.

runtime:
  debug: true            # write transcripts and diagnostics under runs/<ts>/debug/
  recursion_limit: 200   # LangGraph step ceiling for one run

config.yaml is the source of truth for these knobs. The environment sets no knob. An F3DASM_<KEY> variable left exported in a shell stops the run at start with an error that names it, so no arm runs under a label it does not have. A key that this page does not list is ignored with a warning at startup, so a typo tells you instead of silently reverting to the default.

key meaning default
context_window Tokens the backend accepts. 0 asks the server (Ollama's served num_ctx, vLLM's max_model_len). Set it to pin the value or to sweep it. The run records the resolved number and its source. 0
context_policy How the run keeps one agent turn inside the served context window. It applies to OpenAI-compatible backends only, because the Claude SDK manages its own context. compact summarizes the middle of the conversation, so a delegation's result survives even when its prose does not. trim drops those messages instead: it is free and deterministic, but blind to what it discards. You cannot turn it off, because an unchecked context is the crash this setting exists to stop. compact
max_output_tokens Tokens that one model reply may generate. 0 derives the cap from the context window (a quarter of it, capped at 65536), the same share the trim reserves for the reply. -1 removes the cap. The cap bounds a looping turn that would otherwise generate for hours on a large-window server. 0
debug Capture full transcripts, diagnostics, and per-delegation logs under runs/<ts>/debug/. The run-analysis workflow requires it. false
recursion_limit LangGraph step ceiling for one run. 2000
max_awake_nodes How many nodes may be awake at once, the strategizer included. The strategizer always holds one reserved slot, and workers share the rest. A delegation over the cap is reported as QUEUED: too many nodes working (N/N) and starts, first come first served, when a slot frees. A node that blocks in Wait on its own queued child, or whose report awaits review, hands its slot back meanwhile. A critic call runs inside its caller's slot. Memory scales with this knob: 5 awake nodes need --mem >= 32G on Slurm (run 20260928T225501 peaked at 16.76 of 16 GB with three awake). 5
max_consecutive_errors Consecutive failures to one target before the run halts. 12
run_backstop_multiple Multiple of the wall budget after which the run is asked to wind down and closes as backstop_time. The run is force-closed only if the wind-down overruns. 2.0
delegate_cutoff_multiple Multiple of the wall budget past which the run refuses new delegations. In-flight delegations are never touched. It must stay below run_backstop_multiple, or it can never fire. 0 disables it. 1.5
followup_wait_s How long a FollowUp waits for a human answer. 600
stop_grace_s Only with the watchdog launcher: the seconds before its deadline at which it asks the run to wind down (a stop request in debug/) instead of waiting to stop it, so agents hand over real retrospectives. It is on unless you change it: unset means a tenth of the deadline, at most 900 s. The deadline does not move. It must be shorter than the deadline. 0 disables it. min(900, deadline/10)
resume_close_with_retrospectives On a resume whose process was lost (crash, SIGKILL, or out-of-memory), wind the run down at once instead of continuing it, so the entry node writes the retrospective that the crash cost it, including the time bullet. The run closes as crashed. The run logs workers that died in flight as missing (RETROSPECTIVES_MISSING, reason process lost). adda invents nothing for them. false
allow_arm_drift A resume under different ablation arms than the run started with is refused, because a run measured under two arms belongs to neither. Set this to accept the mix. The arms that the run started with stay recorded as arms_initial in run_config.json. false
peer_message_wait_s How long SendMessage(wait_for_reply=True) waits for a peer's reply before it returns. It applies only while peer_interaction is on. 300
llm_retry_max Retry attempts for a failed model call. 5
llm_retry_base Base seconds for the delay between retries. 2.0
llm_stream_idle_timeout Seconds of stream silence before a call is abandoned. 0 disables it. 600.0
llm_tool_idle_timeout Seconds that a single tool call may stall. 0 disables it. 0.0
bash_timeout_s Seconds that a shell command may run before the run moves it to the background instead of blocking the agent. The agent can still pass its own timeout, capped at 600 s. The default is the same on every backend. 120.0
llm_max_buffer_mb Cap on a single response buffered in memory. 30.0
thinking_display summarized or omitted: whether Claude's thinking comes back as a readable summary (visible in transcripts and the viewer) or empty. Newer models default to omitted. Billing is identical either way. It applies only to models that support adaptive thinking (Opus and Sonnet 4.6 and later, and the 5.x families). Other models are untouched. summarized
llm_metadata_fetch Look up model metadata (context window, pricing) at startup. true
llm_metadata_timeout_s Seconds to wait for that lookup. 8.0
llm_quantization Quantization hint for a locally served model. none

Literature retrieval#

These knobs matter only when the graph has a literature reviewer.

key meaning default
retrieval_mode Corpus ranking strategy: auto (RRF when BM25 and dense retrieval are available, else BM25, else substring), hybrid, bm25, or substring. Only auto degrades. An explicitly requested mode that cannot be satisfied raises an error instead of silently falling back to a different one. auto
citation_weighting Multiply BM25 scores by 1 + log10(citations+1) before rank fusion. This is a popularity prior on the lexical side only, and it is untested. true
launch How the viewer starts a run for a study that has its own launcher, such as an sbatch wrapper. It is a mapping with command (argument list, run in the study directory), id_pattern (a regular expression with one group, matched against the launcher's output to capture the job id), stop_command (argument list that contains {id}, which needs id_pattern), and timeout_s (default 120). The viewer runs only these commands and stops only an id it captured. When absent, the viewer offers no launcher control. The viewer reads this key, not the run. none

Ablation switches#

Leave these alone unless you are running an experiment.

These are the switchable parts of the scaffolding. They exist to be measured, not tuned. Every one defaults to on, and turning one off makes a normal run worse by construction. They are the arms of an ablation, which asks whether a piece of machinery earns its cost. An experiment sweeps them, and a study author leaves them alone.

To run an arm without editing the committed study, pass it on the command line:

python -m adda studies/example_study --set hypothesis_ledger=false --set verdict_validator=false

You can repeat --set. python -m adda.watchdog takes the same flag and forwards it. A --set outranks config.yaml, an unknown knob is an error, and the value lands in run_config.json like any explicit knob. For a replicate sweep, launch each replicate from a clean copy of the study: no runs/, no workspace/, and no archived pipeline_*.ipynb. The literature notes under runs/lit_reviewer_notes and the archives are study-scoped, and would otherwise carry one arm's work into the next.

Each switch withholds everything it owns at once: its runtime object, the tools that exist only because of it, and the prompt section that tells the agent to use them. An agent in an arm is never left calling a tool that is gone. For how that is wired, and how to declare a new switch, see Features.

pipeline_deliverable is the exception that is also an ordinary study choice. A study with no notebook deliverable can legitimately turn it off.

key meaning default
hypothesis_ledger The run's falsifiable-hypothesis record. Off withholds its three tools and its prompt section, so the agent is never told to use a tool that is gone. The ablation is partial: the strategizer's method argues the Popperian workflow throughout, and that stays. true
milestones_enabled Run the process-milestone gate. true
science_monitor The runtime drift monitor. It flags evaluations that are missing from the ledger and rows without a stamp, and escalates repeats to the critic. true
f3dasm_api Let the implementer and the data generator look up the installed f3dasm's API (ConsultF3dasm): signatures, docstrings, and source, read off the package that the run executes against, so it cannot go stale. Off withholds the tool and the one prompt section that instructs its use. The CI-verified <f3dasm_api> excerpt stays, so the arm is the excerpt only, which is the state before the tool existed. true
verdict_validator An independent judge reviews each HypothesisUpdate verdict against the falsification charter and appends a concern when it disagrees. It is advisory: it annotates and never blocks. Off, HypothesisUpdate behaves as it did before the judge existed, with no judge call, no note, and no diagnostics. It judges nothing without the ledger, so hypothesis_ledger: false requires verdict_validator: false too. A run that leaves it on refuses to start. true
doe_playbook The implementer's design-of-experiments method prior: the space-filling recipe, the evaluation-budget arithmetic, and the surrogate-guided exploit loop. Off leaves the f3dasm API and the oracle contract intact, and the agent chooses its own method. true
pipeline_deliverable Require pipeline.ipynb as the deliverable. Turn it off for a study with no notebook. true
reproduction_gate Enforce the reproduction gate in Done(). Before a run can close as GATED, the deliverable must reproduce lazily against the canonical store. It must add zero evaluations and modify no rows. It requires pipeline_deliverable: that knob decides whether a notebook is required at all, and this one decides whether an authored notebook must also prove that it reproduces. A study with pipeline_deliverable: false must set reproduction_gate: false too, or the run refuses to start. true
peer_interaction The SendMessage peer and human messaging tool. Every delegation report opens for the delegator's review instead of finalizing on delivery, which replaces the old Confer, Reply, worker FollowUp, and ReportProgress surface. Off is the no-peer-messaging arm: there is no SendMessage, reports finalize on delivery, and only the entry node keeps a human channel (FollowUp). true
budget_notes The in-band budget text that an agent reads: the per-turn constraint snapshot, the budget warnings and wrap-up ladder, and the snapshot on a delegation report and on a worker's task message. Budgets stay soft, and the cost backstop and the critic's constraints are unchanged. true
delegation_contract The <delegation_contract> rules in every worker's preamble: numbers come from tool output, report failures, and do not extend the task. true
model_verification The data generator's correctness check: before delivery it tests the model against expectations that come from outside its own code and records each check in validate_{name}.json, and the handbook chapter verify-before-you-trust explains why. Off, the data generator checks only the interface, with one sample through .call(). true
reprompt_unfinished The bounded re-prompt (up to 3) after a turn ends without an accepted Done(), and the UNGATED banner on the run summary. Off, the run ends at the first such turn. true

Every run records the arms it ran under, defaults included, in run_config.json (arms), in run_status.json, and in the arm_* columns of studies/run_ledger.csv. The runtime entry there lists only the knobs that somebody set, so read arms to tell an all-defaults baseline from a run that nobody labelled.