Skip to content

Use a different model or backend#

By default every agent in a run uses the same backend and model. This page shows how to change that for the whole run or for one agent, and how to set up each backend.

Change the backend for the whole run#

Say the run should go through Ollama instead. Drop a config.yaml next to PROBLEM_STATEMENT.md:

backend: ollama
model: qwen2.5:7b

Nothing else changes. Same graph, same prompts, same AgenticRun(study_dir=study_dir).execute() call (model/backend now come from config.yaml, so the constructor doesn't need them). Every agent in the built-in graph now runs on Ollama, because none of the shipped agents (StrategizerAgent, LiteratureReviewAgent, DataGeneratorAgent, F3dasmImplementerAgent, AdversarialCritiqueAgent) sets its own backend, so each one falls back to the run's default: agent.backend or self._backend.

Change the backend or model for one agent#

backend and model in config.yaml apply to the whole run. To set them for one agent, name the agent in a nodes: block:

backend: claude
nodes:
  implementer:
    backend: ollama
    model: qwen2.5:7b
    base_url: http://localhost:11434/v1

The strategizer and every other agent keep the run's backend. The other keys of a nodes: entry, and the rules they follow, are in Customize agents and tools.

To set a backend or model on an agent that you define in Python, see Customize agents and tools.

The available backends#

Whichever backend a run defaults to, whether set for the whole run or per agent, all backends report token usage through the same telemetry, so cost and throughput stay comparable.

Claude command-line tool (default)#

Uses the local claude command-line tool: see Installation if you haven't set it up yet. Once claude runs on its own in your terminal, there's nothing else to configure beyond the model:

backend: claude
model: claude-haiku-4-5-20251001

Ollama#

A local Ollama server. It listens on http://localhost:11434/v1 unless you set base_url.

backend: ollama
model: qwen2.5:7b
base_url: http://localhost:11434/v1

Inside the container runner, the host's server is at http://host.docker.internal:11434/v1. Put that in base_url.

OpenAI-compatible endpoints (OpenRouter, vLLM, others)#

Any server that speaks the OpenAI API. The base URL comes from, in order: nodes.<node>.base_url, the top-level base_url, the endpoint of a Slurm-served model (llm_slurm), then the backend's default. The environment never sets the endpoint: an exported VLLM_BASE_URL, OLLAMA_BASE_URL or OPENROUTER_BASE_URL stops the run at start-up. Setting both base_url and llm_slurm.enabled is an error.

backend: openrouter        # or: vllm
model: meta-llama/llama-3.1-70b-instruct
base_url: http://host:8000/v1
export OPENROUTER_API_KEY=...        # the key stays in the environment

A local model on a Slurm GPU node (vLLM)#

adda can own a model served on a separate Slurm GPU allocation for the whole run: it submits the vllm serve job, waits for the node and a ready server, points the backend at it over the cluster network, and cancels the job on every exit path. Enable it with an llm_slurm block; a config-time throughput estimate warns if the model/GPU choice is likely to be painfully slow.

backend: vllm
llm_slurm:
  enabled: true
  model: gemma-4-27b-it
  gpu_model: a100            # enables the config-time speed check
  cluster:
    partition: gpu
    account: my_acct
    runner: "uv run python"
    env_setup: ["module load cuda", "module load vllm"]

The llm_slurm block also accepts resource overrides (gres, mem, time), queue and serve timeouts, and a tensor_parallel size for splitting a large model across GPUs.