Core concepts#
If you've just run the Quickstart, you've already seen adda work end to end: one problem statement in, one reproducible notebook out. This page names the pieces that made that happen, so you can reason about what a run is doing, and write a better problem statement for your own, rather than treating it as a black box.
The graph and the open loop#
The agents are nodes in a graph. One node, the strategizer, is the hub: it reads the problem, decides what to do next, and hands work to the specialists. It runs an open loop, meaning it is not a fixed pipeline of numbered steps. After every piece of work comes back, the strategizer looks at the current state and chooses the next move. The loop ends when the strategizer declares the work done and that decision survives review.
The specialists it delegates to:
- literature reviewer: finds and reads relevant papers.
- data generator: turns a way of evaluating a design into a metered oracle the rest of the system can call.
- implementer: writes and runs the actual code (sampling, surrogates, optimisation) against that oracle.
- critic: an adversarial reviewer that tries to find holes in a claimed result before it is accepted.
Delegation#
When the strategizer hands work to a specialist, that is a delegation. Each
delegation has an id (D001, D002, …), a task description, and a report that
comes back. Delegations are the unit of work and the unit of accounting: every
real evaluation is attributed to the delegation that produced it, which is what
lets adda tell you exactly where each number came from.
The hypothesis ledger and the falsification charter#
adda does science, so it tracks hypotheses explicitly. A hypothesis is a claim with a testable criterion, a prediction, a prior, and (as evidence comes in) a verdict. These live in the hypothesis ledger.
The rules for what counts as real evidence live in the falsification charter. The charter is Popperian: a hypothesis cannot be marked SUPPORTED without a recorded attempt to falsify it. This is the epistemic backbone. It exists so the system cannot quietly talk itself into a conclusion the evidence does not carry.
The canonical evaluation ledger#
Every real oracle evaluation is written once, under a lock, to a shared
canonical ledger (an ExperimentData store), and stamped with the delegation
that produced it. This store is the single source of truth for what was actually
measured. It is protected: a stray write that would shrink it, or reset a
completed evaluation, is refused. The headline number in the final deliverable
must trace back to rows in this ledger, or the deliverable cannot reproduce it.
The deliverable and the reproduction gate#
The output of a run is a Jupyter notebook, pipeline.ipynb. It is not a summary
written after the fact; it is the work. Its markdown cells hold the writeup and
its code cells rederive the headline result from the canonical ledger.
Before a run is allowed to close, the notebook goes through the reproduction gate: it is executed end to end in a clean sandbox, and the number it produces is checked against the number the run claims. A run that cannot reproduce its own headline does not pass. This is why the notebook you get back runs as-is.
The reproduction gate is one of several checks between a decision and its counting—some of which refuse outright, and some of which only speak up. How a run is kept honest is the full inventory, including which ones you can switch off.
Model backends#
The agents are driven by a language model through a backend. adda ships several: the Claude command-line tool (default), any OpenAI-compatible endpoint, Ollama, OpenRouter, and vLLM (including a mode that serves a model on a Slurm GPU node the framework owns for the run). The backend is a configuration choice; the graph and the science do not change with it.
Resource governance#
Long autonomous runs need guardrails. adda separates two kinds:
- Soft budgets (
eval_budget, the evaluation count, andbudget, the wall-clock time) nudge the strategizer when it is spending heavily, but never hard-stop the science. - Hard caps stop the run outright: per-delegation memory (host safety,
enforced by a watchdog that also reaps runaway processes and force-exits a
stalled run), and an optional
budget_usdcost ceiling (resumable—raise it and resume; inactive on a backend with no per-call cost data, for example Ollama).
What you provide, what you get#
You provide one file: PROBLEM_STATEMENT.md in a study directory (plus a
config.yaml if you want to set the backend, budgets, or the evaluator). You get
back pipeline.ipynb (the reproducible answer) alongside the run's evaluation
record and logs. See Authoring a study to set one up,
Watching a run to follow it live (and answer it, if it
asks), Understanding a run's output for what comes back, and
the Quickstart to run one.