Skip to content

Oracle and ledger#

The metered evaluator and the canonical evaluation ledger.

adda.get_evaluator(namespace: str | None = None) -> InstrumentedDataGenerator #

The ONE door to a registered ground-truth oracle.

Locates run_config.json by walking up from Path.cwd(), reads all configuration from it, and derives the delegation ID from the current working directory name (expected pattern D###).

The oracle is the source registered for this run, and its evaluations are written to the canonical store with provenance. There is no way to substitute an arbitrary inner generator or redirect the store — that is what makes ground-truth metering airtight. (Surrogates, stubs, and analysis are the agent's own DataGenerators, run freely off-ledger — never through here.)

Parameters:

Name Type Description Default
namespace str or None

The design namespace whose oracle + ledger to resolve. None (the default) uses the flat single-study config and the canonical store — the behavior every existing study relies on. A non-None namespace resolves run_config["oracles"][namespace] (its own oracle + its own isolated store). When omitted, the namespace falls back to the F3DASM_NAMESPACE environment variable, so a delegation scoped to a namespace keeps the agent's call site a plain get_evaluator().

None

Returns:

Type Description
InstrumentedDataGenerator

Raises:

Type Description
ValueError

If the cwd is not a D### directory and F3DASM_DELEGATION_ID is not set, if no evaluator source is registered, or if namespace is not a registered namespace.

FileNotFoundError

If run_config.json cannot be found by walking up from cwd.

Source code in src/adda/_src/evaluation/oracle_resolution.py
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
def get_evaluator(namespace: str | None = None) -> InstrumentedDataGenerator:
    """The ONE door to a registered ground-truth oracle.

    Locates ``run_config.json`` by walking up from ``Path.cwd()``, reads
    all configuration from it, and derives the delegation ID from the
    current working directory name (expected pattern ``D###``).

    The oracle is the source registered for this run, and its evaluations are
    written to the canonical store with provenance. There is no way to substitute
    an arbitrary inner generator or redirect the store — that is what makes
    ground-truth metering airtight. (Surrogates, stubs, and analysis are the
    agent's own DataGenerators, run freely off-ledger — never through here.)

    Parameters
    ----------
    namespace : str or None, optional
        The design namespace whose oracle + ledger to resolve. ``None`` (the
        default) uses the flat single-study config and the canonical store — the
        behavior every existing study relies on. A non-``None`` namespace resolves
        ``run_config["oracles"][namespace]`` (its own oracle + its own isolated
        store). When omitted, the namespace falls back to the ``F3DASM_NAMESPACE``
        environment variable, so a delegation scoped to a namespace keeps the
        agent's call site a plain ``get_evaluator()``.

    Returns
    -------
    InstrumentedDataGenerator

    Raises
    ------
    ValueError
        If the cwd is not a ``D###`` directory and ``F3DASM_DELEGATION_ID`` is
        not set, if no evaluator source is registered, or if ``namespace`` is not
        a registered namespace.
    FileNotFoundError
        If ``run_config.json`` cannot be found by walking up from cwd.
    """
    delegation_id = _resolve_delegation_id()
    run_config = _load_run_config()

    namespace_from_env = False
    if namespace is None:
        env_ns = os.environ.get("F3DASM_NAMESPACE", "")
        if env_ns:
            namespace, namespace_from_env = env_ns, True
    try:
        cfg = _effective_oracle_config(run_config, namespace)
    except ValueError as exc:
        if namespace_from_env:
            # De-footgun: get_evaluator() SILENTLY inherits F3DASM_NAMESPACE, so a
            # namespace-scoped delegation asking for the DEFAULT/baseline oracle
            # gets a confusing "unknown namespace". Name the env source + the fix.
            raise ValueError(
                f"{exc} NOTE: namespace {namespace!r} was inherited from the "
                "F3DASM_NAMESPACE environment variable, not passed explicitly. "
                "For the default/baseline oracle, clear it first "
                "(`env -u F3DASM_NAMESPACE ...`) or call from a process where it "
                "is unset."
            ) from exc
        raise

    store_dir = Path(cfg["store_dir"])
    lock_path_str = cfg.get("lock_path")
    lock_path = (
        Path(lock_path_str)
        if lock_path_str
        else store_dir / "experiment_data" / ".lock"
    )
    source = cfg.get(
        "source",
        cfg.get("evaluator_name", ""),
    )
    fidelity_column = cfg.get("fidelity_column")
    # Extensible provenance declared for this run (open schema; stamped on
    # every row by the wrapper, never by the agent). Tolerate a non-dict.
    _prov = cfg.get("provenance")
    extra_provenance = _prov if isinstance(_prov, dict) else {}

    study_dir_str = cfg.get("study_dir")
    if study_dir_str is None:
        raise ValueError(
            "run_config.json is missing 'study_dir' key; "
            "re-run your study to regenerate it."
        )
    inner = load_inner_evaluator(cfg, Path(study_dir_str))
    if inner is None:
        raise ValueError(
            "No ground-truth oracle is registered for this run. A source "
            "must be authored and registered by the datagenerator agent "
            "(it writes a registration.json the runtime reads) before "
            "evaluations can be ledgered. If this study genuinely has no "
            "registerable oracle, report evaluation counts manually via "
            "ReportEvals (honour-system, off-ledger)."
        )

    # Hard memory cap + PID registration for THIS campaign process (the one
    # hard resource boundary). Done here because every campaign reaches the
    # oracle through get_evaluator, regardless of how it was launched.
    _apply_process_governor(cfg, store_dir, delegation_id)

    # F3DASM_DEDUP_SCOPE=all: set by notebook_exec.sandbox_env() for
    # reproduction-gate/deliverable execution, where the process is stamped
    # with a synthetic delegation id that never matches the real
    # delegation(s) that actually generated the ledger's rows — the default
    # per-delegation dedup scope would silently fail to recognize ANY
    # existing row as already-seen there. Real campaign delegations never
    # set this, so they keep the default "delegation" scope.
    dedup_scope = os.environ.get("F3DASM_DEDUP_SCOPE", "delegation")

    oracle_rev = oracle_revision(cfg, Path(study_dir_str)) or None
    if oracle_rev and dedup_scope == "delegation":
        from .oracle_edits import check_oracle_edit
        check_oracle_edit(
            cfg, Path(study_dir_str), namespace, oracle_rev, delegation_id)

    return InstrumentedDataGenerator(
        inner=inner,
        store_dir=store_dir,
        delegation_id=delegation_id,
        source=source,
        fidelity_column=fidelity_column,
        lock_path=lock_path,
        extra_provenance=extra_provenance,
        eval_budget=cfg.get("eval_budget"),
        dedup_scope=dedup_scope,
        oracle_rev=oracle_rev,
    )

adda.LookupDataGenerator #

Evaluate candidates by nearest-neighbour lookup against a fixed pool.

For every execute call the generator finds the pool row whose normalised input vector is closest (L2 distance over min-max-scaled coordinates) to the candidate and copies that row's output columns onto the sample. The mechanism is the canonical Level-1 evaluator for agentic-f3dasm — the agent never proposes coordinates that the pool knows about ahead of time, but every proposal can be mapped to a pool row deterministically.

Parameters:

Name Type Description Default
pool ExperimentData

The pre-computed dataset to look up against. Must already carry finished output columns for every row.

required
input_columns list of str

Names of the input columns (already present on every pool row) that contribute to the L2 distance.

required
output_columns list of str

Names of the output columns to copy from the matched pool row. If None, every key in the matched row's _output_data dict is copied.

None

Attributes:

Name Type Description
pool ExperimentData

The pool the generator was constructed with (kept by reference for downstream introspection; never mutated).

seen_indices set of int

Pool indices that have already been returned during this run; the set is reset by :meth:reset_seen.

Notes

The deep-copy of pool output data on every call guarantees that the pool is never mutated by downstream code that edits the returned sample. The generator deliberately does NOT raise on a repeated pool hit — exhaustion of one region is not exhaustion of the search. Strategies can poll :meth:consume_repeats after a batch to surface a warning to the agent.

Source code in src/adda/_src/evaluation/lookup.py
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
class LookupDataGenerator(DataGenerator):
    """Evaluate candidates by nearest-neighbour lookup against a fixed pool.

    For every ``execute`` call the generator finds the pool row whose
    normalised input vector is closest (L2 distance over min-max-scaled
    coordinates) to the candidate and copies that row's output columns
    onto the sample. The mechanism is the canonical Level-1 evaluator for
    agentic-f3dasm — the agent never proposes coordinates that the pool
    knows about ahead of time, but every proposal can be mapped to a pool
    row deterministically.

    Parameters
    ----------
    pool : ExperimentData
        The pre-computed dataset to look up against. Must already carry
        finished output columns for every row.
    input_columns : list of str
        Names of the input columns (already present on every pool row)
        that contribute to the L2 distance.
    output_columns : list of str, optional
        Names of the output columns to copy from the matched pool row.
        If ``None``, every key in the matched row's ``_output_data`` dict
        is copied.

    Attributes
    ----------
    pool : ExperimentData
        The pool the generator was constructed with (kept by reference for
        downstream introspection; never mutated).
    seen_indices : set of int
        Pool indices that have already been returned during this run; the
        set is reset by :meth:`reset_seen`.

    Notes
    -----
    The deep-copy of pool output data on every call guarantees that the
    pool is never mutated by downstream code that edits the returned
    sample. The generator deliberately does NOT raise on a repeated pool
    hit — exhaustion of one region is not exhaustion of the search.
    Strategies can poll :meth:`consume_repeats` after a batch to surface a
    warning to the agent.
    """

    def __init__(
        self,
        pool: ExperimentData,
        input_columns: list[str],
        output_columns: Optional[list[str]] = None,
    ) -> None:
        self.pool = pool
        self.input_columns = list(input_columns)
        self.output_columns = (
            None if output_columns is None else list(output_columns)
        )

        # Materialise the pool's input vectors and bounds once so every
        # execute() is a constant-time normalisation + linear scan.
        pool_inputs = np.asarray(
            [
                [row._input_data[col] for col in self.input_columns]
                for row in pool.data.values()
            ],
            dtype=float,
        )
        self._pool_indices: list[int] = list(pool.data.keys())
        self._pool_inputs = pool_inputs

        # Per-column min/max for L2 normalisation. A column with zero
        # spread (all values equal) gets a unit denominator so the
        # contribution to L2 collapses to 0 rather than NaN.
        col_min = pool_inputs.min(axis=0)
        col_max = pool_inputs.max(axis=0)
        spread = col_max - col_min
        spread[spread == 0.0] = 1.0
        self._col_min = col_min
        self._col_spread = spread
        self._pool_normalised = (pool_inputs - col_min) / spread

        # Provenance for warning surfaces consumed by strategy adapters.
        self.seen_indices: set[int] = set()
        self._repeats: int = 0

    # --------------------------------------------------------- public API

    def execute(
        self, experiment_sample: ExperimentSample, **kwargs
    ) -> ExperimentSample:
        """Look up the nearest pool row and copy its outputs onto the sample.

        Parameters
        ----------
        experiment_sample : ExperimentSample
            Sample whose input data carries values for every column in
            :attr:`input_columns`.
        **kwargs : dict
            Unused; present to match the ABC signature.

        Returns
        -------
        ExperimentSample
            The same sample object, with output data filled and the job
            status marked as finished.

        Raises
        ------
        KeyError
            If the sample is missing one of the configured input columns.
        """

        # Vectorised L2 over the normalised pool.
        candidate = np.asarray(
            [experiment_sample._input_data[col] for col in self.input_columns],
            dtype=float,
        )
        normalised_candidate = (candidate - self._col_min) / self._col_spread
        distances = np.linalg.norm(
            self._pool_normalised - normalised_candidate, axis=1
        )
        nearest_local_idx = int(np.argmin(distances))
        nearest_pool_idx = self._pool_indices[nearest_local_idx]

        matched_row = self.pool.data[nearest_pool_idx]
        # Deep-copy guards against the agent mutating the pool via the
        # returned sample.
        source_outputs = deepcopy(matched_row._output_data)
        if self.output_columns is None:
            outputs_to_copy = source_outputs
        else:
            outputs_to_copy = {
                key: source_outputs[key]
                for key in self.output_columns
                if key in source_outputs
            }

        experiment_sample._output_data.update(outputs_to_copy)
        experiment_sample.job_status = JobStatus.FINISHED

        if nearest_pool_idx in self.seen_indices:
            self._repeats += 1
        else:
            self.seen_indices.add(nearest_pool_idx)

        return experiment_sample

    def consume_repeats(self) -> int:
        """Return and reset the repeat counter.

        Strategy adapters call this after a batch to surface a pool
        exhaustion warning to the agent.

        Returns
        -------
        int
            Number of times ``execute`` returned a previously-seen pool
            row since the last call to this method.
        """

        repeats = self._repeats
        self._repeats = 0
        return repeats

    def reset_seen(self) -> None:
        """Forget which pool indices have already been returned.

        Useful for fresh runs that reuse the same pool instance.
        """

        self.seen_indices.clear()
        self._repeats = 0
pool = pool instance-attribute #
input_columns = list(input_columns) instance-attribute #
output_columns = None if output_columns is None else list(output_columns) instance-attribute #
_pool_indices: list[int] = list(pool.data.keys()) instance-attribute #
_pool_inputs = pool_inputs instance-attribute #
_col_min = col_min instance-attribute #
_col_spread = spread instance-attribute #
_pool_normalised = (pool_inputs - col_min) / spread instance-attribute #
seen_indices: set[int] = set() instance-attribute #
_repeats: int = 0 instance-attribute #
execute(experiment_sample: ExperimentSample, **kwargs) -> ExperimentSample #

Look up the nearest pool row and copy its outputs onto the sample.

Parameters:

Name Type Description Default
experiment_sample ExperimentSample

Sample whose input data carries values for every column in :attr:input_columns.

required
**kwargs dict

Unused; present to match the ABC signature.

{}

Returns:

Type Description
ExperimentSample

The same sample object, with output data filled and the job status marked as finished.

Raises:

Type Description
KeyError

If the sample is missing one of the configured input columns.

Source code in src/adda/_src/evaluation/lookup.py
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
def execute(
    self, experiment_sample: ExperimentSample, **kwargs
) -> ExperimentSample:
    """Look up the nearest pool row and copy its outputs onto the sample.

    Parameters
    ----------
    experiment_sample : ExperimentSample
        Sample whose input data carries values for every column in
        :attr:`input_columns`.
    **kwargs : dict
        Unused; present to match the ABC signature.

    Returns
    -------
    ExperimentSample
        The same sample object, with output data filled and the job
        status marked as finished.

    Raises
    ------
    KeyError
        If the sample is missing one of the configured input columns.
    """

    # Vectorised L2 over the normalised pool.
    candidate = np.asarray(
        [experiment_sample._input_data[col] for col in self.input_columns],
        dtype=float,
    )
    normalised_candidate = (candidate - self._col_min) / self._col_spread
    distances = np.linalg.norm(
        self._pool_normalised - normalised_candidate, axis=1
    )
    nearest_local_idx = int(np.argmin(distances))
    nearest_pool_idx = self._pool_indices[nearest_local_idx]

    matched_row = self.pool.data[nearest_pool_idx]
    # Deep-copy guards against the agent mutating the pool via the
    # returned sample.
    source_outputs = deepcopy(matched_row._output_data)
    if self.output_columns is None:
        outputs_to_copy = source_outputs
    else:
        outputs_to_copy = {
            key: source_outputs[key]
            for key in self.output_columns
            if key in source_outputs
        }

    experiment_sample._output_data.update(outputs_to_copy)
    experiment_sample.job_status = JobStatus.FINISHED

    if nearest_pool_idx in self.seen_indices:
        self._repeats += 1
    else:
        self.seen_indices.add(nearest_pool_idx)

    return experiment_sample
consume_repeats() -> int #

Return and reset the repeat counter.

Strategy adapters call this after a batch to surface a pool exhaustion warning to the agent.

Returns:

Type Description
int

Number of times execute returned a previously-seen pool row since the last call to this method.

Source code in src/adda/_src/evaluation/lookup.py
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
def consume_repeats(self) -> int:
    """Return and reset the repeat counter.

    Strategy adapters call this after a batch to surface a pool
    exhaustion warning to the agent.

    Returns
    -------
    int
        Number of times ``execute`` returned a previously-seen pool
        row since the last call to this method.
    """

    repeats = self._repeats
    self._repeats = 0
    return repeats
reset_seen() -> None #

Forget which pool indices have already been returned.

Useful for fresh runs that reuse the same pool instance.

Source code in src/adda/_src/evaluation/lookup.py
199
200
201
202
203
204
205
206
def reset_seen(self) -> None:
    """Forget which pool indices have already been returned.

    Useful for fresh runs that reuse the same pool instance.
    """

    self.seen_indices.clear()
    self._repeats = 0

adda.load_experiments(store_root: Path | str | None = None) -> dict[str, ExperimentData] #

Load EVERY experiment store of a run as a dict {name: ExperimentData}.

A namespaced run holds one clean ExperimentData per experiment at nested paths — the default store at <root>/experiment_data/ and each design experiment at <root>/<name>/experiment_data/ — so a single ExperimentData.from_file (the single-study idiom) loads only the default store and silently misses the rest. This is the multi-namespace load idiom for pipeline.ipynb: one call returns them all, keyed by experiment name (the default/baseline store is "default").

store_root defaults to $F3DASM_CANONICAL_STORE (set in the notebook's execution env), so the notebook body is just experiments = load_experiments(). Empty/absent stores are skipped; an empty run yields {}. Never raises on a missing store.

Source code in src/adda/_src/evaluation/ledger_summary.py
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
def load_experiments(
    store_root: Path | str | None = None,
) -> dict[str, ExperimentData]:
    """Load EVERY experiment store of a run as a dict ``{name: ExperimentData}``.

    A namespaced run holds one clean ``ExperimentData`` per experiment at nested
    paths — the default store at ``<root>/experiment_data/`` and each design
    experiment at ``<root>/<name>/experiment_data/`` — so a single
    ``ExperimentData.from_file`` (the single-study idiom) loads only the default
    store and silently misses the rest. This is the multi-namespace load idiom
    for pipeline.ipynb: one call returns them all, keyed by experiment name
    (the default/baseline store is ``"default"``).

    ``store_root`` defaults to ``$F3DASM_CANONICAL_STORE`` (set in the notebook's
    execution env), so the notebook body is just
    ``experiments = load_experiments()``. Empty/absent stores are skipped; an
    empty run yields ``{}``. Never raises on a missing store.
    """
    if store_root is None:
        store_root = os.environ.get("F3DASM_CANONICAL_STORE", "")
    store_root = Path(store_root)
    out: dict[str, ExperimentData] = {}
    for store in experiment_stores(store_root):
        name = "default" if store == store_root else store.name
        try:
            data = ExperimentData.from_file(project_dir=store)
        except Exception:  # noqa: BLE001 — empty/absent store
            continue
        if len(data) > 0:
            out[name] = data
    return out