Skip to content

Quickstart#

One file in, one notebook back. This runs it once, end to end, so you can see the shape of it before writing your own.

Before you start#

You need the package and a model to drive the agents, two commands, see Installation. If claude already runs in your terminal, you're set.

1. Describe the problem#

A study is just a directory with one required file. Save this as studies/quickstart/PROBLEM_STATEMENT.md — any well-defined objective works the same way; this one is a standard 2D test function (Branin):

Minimise the 2D Branin function over its standard domain.
Report the best design found and the objective value there.
from pathlib import Path

study_dir = Path("studies/quickstart")
study_dir.mkdir(parents=True, exist_ok=True)  # save PROBLEM_STATEMENT.md here, per above

2. Run it#

AgenticRun reads the problem statement and runs to a gated deliverable, nothing else to configure. model="haiku" is an alias for the latest Claude Haiku release — no dated tag to get wrong or watch go stale. For a different provider, or to pin an exact model tag, see Use a different model or backend. execute() returns the final report text.

from adda import AgenticRun

report = AgenticRun(
    study_dir=study_dir,
    model="haiku",  # alias for the latest Claude Haiku; swap for any backend/model
    eval_budget=80,  # a soft cap: nudges the strategizer, never hard-stops it
).execute()

print(report)  # the same write-up that lands in pipeline.ipynb
## Minimization of 2D Branin Function — Final Report

### Best Design Found
- Location: (x = 9.4184, y = 2.3951)
- Objective value: f = 0.403634
- Global minimum: f ≈ 0.397887 (0.14% above optimum)

### Campaign Overview
Two strategies were run and compared: a surrogate-guided Bayesian
optimization (80 evaluations, best f = 0.4263) and a dense LHS exploration
(100 evaluations, best f = 0.4036 — the overall best). Total: 180
evaluations, all recorded in the canonical ledger with full provenance.

### Hypothesis Verdicts
- H1 (surrogate-guided optimization beats a threshold): SUPPORTED.
- H2 (dense exploration beats the surrogate at equal budget): INCONCLUSIVE
  — the two campaigns used different budgets (100 vs. 80), so the
  comparison is confounded and isn't claimed as a clean result.

(Full report continues with deliverable structure and reproducibility
checks — truncated here for the docs page.)

3. What you got back#

Two things, in the study directory:

  • pipeline.ipynb: the deliverable. Its opening cells are the write-up above; its code cells re-derive the headline number from the run's own evaluation record, so it reproduces on re-run rather than asking you to take it on faith.
  • runs/<timestamp>/run_status.json: GATED means it passed an adversarial review and reproduces. If you check exactly one thing, check this.

See Understanding a run's output for the rest: logs, the evaluation record, and what to open when a result looks off.

4. Now run your own problem#

Everything above works the same for any problem; only step 1 changes. Replace the PROBLEM_STATEMENT.md text with your own objective, design space (bounds, types, units), and what counts as a valid answer, then call AgenticRun on that directory. That's the only required change.

If your design needs to be scored by your own code (a simulator, a dataset, a physics model) rather than something the agents can reason out on their own, see Authoring a study: it covers the evaluator options and a fully worked example, file by file.