Authoring a study#
A study is the single input to a run. It's a folder you prepare; adda reads it, does the work, and writes its results back into the same folder:
from adda import AgenticRun
AgenticRun(study_dir="my_study").execute()
This page shows what to put in that folder.
What the folder contains#
my_study/
PROBLEM_STATEMENT.md # required: the brief the agents work from
config.yaml # optional: the model, budgets, and how to evaluate a design
workspace/ # optional: your evaluator, if you ship one
evaluator.py
Everything else you'll see later (pipeline.ipynb, a runs/ directory) is
produced by the run: you don't write it.
PROBLEM_STATEMENT.md (required)#
This is the whole task definition. The agents work from it and the critic checks the result against it, so write it precisely: a vague brief produces a vague, unverifiable answer. Cover:
- Objective and success criteria: the headline number or claim the run must deliver (for example, "maximise the normalised buckling load; report the design and its value").
- Design space: every input variable, with its bounds, type (continuous, integer, or categorical), and units.
- Deliverables: anything the run must produce beyond the notebook (a plot, a mechanism explanation).
- What "valid" means: feasibility limits, regimes of validity, noise thresholds. State these explicitly; they're what keep the agents honest.
config.yaml (optional)#
Every setting has a default, so you can omit this file entirely. The config reference lists every key, its meaning, and its default. The keys you set most often are:
model: claude-haiku-4-5-20251001 # which language model to use
eval_budget: 200 # soft cap on real evaluations
budget: "01:00:00" # soft wall-clock limit
Put settings for your own scripts under study:, which adda never reads.
How designs get evaluated (the evaluator)#
The evaluator is your ground truth: the function that scores a design. It's the
one metered path: the only calls counted against eval_budget. (Surrogates and
optimisers the agents build on top are their own business and aren't metered.)
Declare it one of these ways:
A function you ship, one argument per input variable:
evaluator:
entrypoint: "workspace/evaluator.py:evaluate" # path within the study folder
output_names: [y] # names of what it returns
A precomputed table, each query resolves to the nearest row (results are approximate, since a query may land between rows):
evaluator:
lookup:
pool: "experiment_data" # path within the study folder
input_columns: [x1, x2]
output_columns: [y]
Described in the brief: omit the evaluator block and describe the scoring
oracle (a binary, a dataset, a physics model) in PROBLEM_STATEMENT.md; the
agents write the evaluator themselves during the run.
Nothing: no evaluator and none described falls back to the honor system, and the agents self-report. Fine for exploring, not for a result you want verified.
Declaring the objective#
Say what "best" and "feasible" mean, once, and the run ledger and the viewer's figure of merit both use it. Nothing is inferred from column names.
objective:
column: score
direction: max # max | min
feasible: feasible # optional: a 0/1 output column
lines: # optional: labelled reference lines on the Data chart
- {value: 2.0, label: "reference"}
- {value: 20.0, label: "goal"}
unit_label: {divide_by: 2.0, label: "x reference"} # optional: display scaling
A design counts only if its column value is finite and, when feasible is
given, that column is 1; a finite value on an infeasible design does not count.
column and feasible must be outputs the evaluator declares (output_names
or the output_columns of the lookup); an unknown name refuses the run at start.
When the evaluator is written during the run, the start check is skipped; each
oracle later registered without a declared column is reported to the delegating
agent and recorded as an OBJECTIVE_COLUMN_MISSING diagnostics event.
lines are drawn on the viewer's best-so-far chart as dashed, labelled rules;
value is in the objective column's own units, label is any non-empty text,
and no line is drawn unless you declare it. unit_label only changes how the
viewer displays the axis, lines and values (raw value divided by divide_by,
shown with label); the store and the ledger stay in raw units.
The objective is scored on every store that records both declared columns: the
canonical store and each namespace. Results are therefore comparable across
oracle families, and the best row names the store it came from. A store missing
a declared column is listed as "not scored" rather than dropped.
Without the block the ledger records objective: undeclared, judges rows by
the finite rule alone, and reports the running min and max instead of a best.
The viewer's stage funnel counts how many designs survive each 0/1 output column in turn. Declare the stages, in order, at the top level of config.yaml:
funnel: [valid, simulated, converged]
Each must be an output the evaluator declares. With no funnel: the viewer
draws no funnel: the 0/1 columns are listed as independent flags (count of ones
per column) in a compact table, because they are not ordered stages unless you say so.
What the run produces#
At the study root you get pipeline.ipynb, the deliverable. Its opening
cells are the write-up; its code cells reproduce the headline result. Each run
also writes a timestamped folder under runs/ with the evaluation record, logs,
and a status file. See Understanding a run's output for what's
in there and how to read it.
Everything on this page so far is config.yaml and PROBLEM_STATEMENT.md; the graph itself
(which agents exist, how they delegate), each agent's system prompt, and each
agent's backend/model are Python-level extension points instead. See
Customize agents and tools if the built-in strategizer and
four specialists aren't the shape your problem needs.
Before a long run, check#
PROBLEM_STATEMENT.mdstates explicit success criteria, the design space (bounds, types, units), and any deliverables.config.yamlparses, and itsevaluatorpoints at a real file/attribute or a real lookup pool.- Your evaluator imports and runs on one sample without error.
- Your backend is reachable (for the default Claude backend, you're logged in; see Installation).
A minimal worked example#
A runnable copy of this folder ships in the repository at studies/example_study
(it's exercised by the test suite, so it can't drift from what the runtime
actually expects).
config.yaml:
model: claude-haiku-4-5-20251001
backend: claude
eval_budget: 200 # soft cap on real calls to evaluate(), not a hard stop
evaluator:
entrypoint: "workspace/evaluator.py:evaluate" # module:function, relative to the study
output_names: [y] # names the one value evaluate() returns
workspace/evaluator.py:
def evaluate(x1: float, x2: float) -> float:
"""One argument per input; returns the output named in output_names."""
return (x1 - 1.0) ** 2 + (x2 + 2.0) ** 2 # the ground truth being scored
PROBLEM_STATEMENT.md:
# Minimise a 2-D quadratic
Objective: minimise y = (x1-1)^2 + (x2+2)^2.
Success: report the argmin (x1*, x2*) and the value y*, reproduced in
pipeline.ipynb from the run's evaluation record.
Design space: x1, x2, continuous, in [-5, 5], dimensionless.
Then run it:
from adda import AgenticRun
AgenticRun(study_dir="studies/example_study").execute()
Launching under a watchdog#
budget and run_backstop_multiple (described earlier) are both checked from inside
the run, so neither can help if the run genuinely wedges—a hung model call,
a stuck simulation—and never reaches its own next check. For that, launch
the study as a child process under an external watchdog instead: it owns the
wall-clock deadline from outside, and force-kills the whole run (every
process it spawned, not just the top one) if the deadline passes.
python -m adda.watchdog studies/example_study --budget 00:45:00
The deadline is twice whatever budget you give it (or config.yaml's own
budget: if you don't pass --budget)—plenty of headroom, since this is a
last-resort cutoff for a hang, not a way to police a slow run. python -m
adda <study-dir> on its own still works exactly as before; this is an
additional, safer way to launch the same run when you want a hard outer
backstop.
Stopping the watchdog (Ctrl-C, or kill on its PID) takes the run down with
it: the whole process tree is reaped and the watchdog exits 130 (SIGINT) or
143 (SIGTERM). A run is never left running with nothing watching it.