API harnesses
API harnesses call an LLM API directly instead of running a coding agent. A step on an API harness sends one request (system prompt + your rendered prompt) and uses the response text as the step’s output.
Two API harnesses are built in: Claude API and OpenAI-compatible API. To use several OpenAI-compatible providers at once (say OpenAI and a local Ollama), add more with Add harness and set Type to LLM API.
Editor fields
Section titled “Editor fields”| Field | Meaning |
|---|---|
| Provider | Anthropic (Claude API) or OpenAI-compatible (OpenAI, OpenRouter, Ollama, LM Studio…) |
| Base URL | Anthropic: leave empty for api.anthropic.com. OpenAI-compatible: the URL that /chat/completions is appended to, e.g. http://localhost:11434/v1. Empty means https://api.openai.com/v1. |
| API key | Stored encrypted with your OS keychain. Takes precedence over the environment variable. |
| …or environment variable | Name of the env var holding the key. Empty means ANTHROPIC_API_KEY (Anthropic) or OPENAI_API_KEY (OpenAI-compatible). |
| Max output tokens | Sent as max_tokens (Anthropic) or max_completion_tokens (OpenAI-compatible). |
| Models | The one-click choices in the step editor, with a ★ default. Fetch latest pulls the provider’s list. |
The key lookup order is: API key on the harness → the named env var in the step’s environment (which includes Settings → Execution → Environment variables, project env and step env) → the same env var in the app’s own environment.
Claude API
Section titled “Claude API”| Setting | Default |
|---|---|
| Id | anthropic-api |
| Provider | Anthropic |
| Key env var | ANTHROPIC_API_KEY |
| Max output tokens | 32000 |
| Models | claude-opus-5, claude-opus-5-5, claude-fable-5-1, claude-sonnet-5, claude-haiku-4-5 |
| Default model | claude-opus-5 |
How a step runs:
- Uses the official Anthropic SDK and streams the Messages API response. Text appears in the log line by line as it arrives.
- Sends the step’s system prompt as
systemand the rendered prompt as a single user message. - Enables adaptive thinking (
thinking: { type: "adaptive" }). - If no key is found at all, the SDK falls back to its own credential lookup.
- Records input, output, cache-read and cache-write tokens. The API doesn’t return a price, so the step’s cost stays $0.
- Fails with The model declined this request (refusal) on a refusal stop reason, and with Response hit max_tokens if the answer was cut off. Raise Max output tokens in that case.
The Bug fix (fast lane) template uses the Claude API for its triage step.
Fetch latest lists models from the Anthropic Models API.
OpenAI-compatible API
Section titled “OpenAI-compatible API”| Setting | Default |
|---|---|
| Id | openai-compatible |
| Base URL | https://api.openai.com/v1 |
| Key env var | OPENAI_API_KEY |
| Max output tokens | 16000 |
| Models | gpt-5, gpt-5-mini |
| Default model | gpt-5 |
How a step runs:
POST {base URL}/chat/completionswithmodel,messages(asystemmessage if the step has a system prompt, then theuserprompt) andmax_completion_tokens.- Sends
Authorization: Bearer <key>only if a key was found, so keyless local servers work. - Not streamed: the whole answer is logged when it arrives.
- Uses
choices[0].message.contentas the output and recordsprompt_tokens/completion_tokensif the server returnsusage. No cost. - The request timeout is the step’s timeout, or 10 minutes if the step has none.
Setting Max output tokens to 0 leaves max_completion_tokens out of the request, which helps with servers that reject that field.
Fetch latest calls GET {base URL}/models (with the bearer key if any) and reads the data array.
Provider settings
Section titled “Provider settings”| Field | Value |
|---|---|
| Base URL | empty, or https://api.openai.com/v1 |
| Key | an OpenAI API key, or OPENAI_API_KEY |
| Models | e.g. gpt-5, gpt-5-mini — or click Fetch latest |
| Field | Value |
|---|---|
| Base URL | https://openrouter.ai/api/v1 |
| Key | an OpenRouter key (set …or environment variable to e.g. OPENROUTER_API_KEY if you keep it in the environment) |
| Models | provider-prefixed ids such as anthropic/claude-sonnet-5; Fetch latest returns OpenRouter’s full catalog |
| Field | Value |
|---|---|
| Base URL | http://localhost:11434/v1 |
| Key | none needed |
| Models | the names you pulled, e.g. qwen3-coder:30b; Fetch latest lists local models |
A localhost base URL counts as available without a key. See Local models with Ollama.
| Field | Value |
|---|---|
| Base URL | http://localhost:1234/v1 (LM Studio’s local server default) |
| Key | none needed |
| Models | the loaded model’s id; Fetch latest lists them |
| Field | Value |
|---|---|
| Base URL | http://<host>:8000/v1 (vLLM’s OpenAI-compatible server default port) |
| Key | only if you started vLLM with --api-key |
| Models | the served model name; Fetch latest lists it |
A vLLM server on another machine isn’t localhost, so detection reports it unavailable until a key or key variable is set. Runs still work without one if the server doesn’t require it.
Adding a second API harness
Section titled “Adding a second API harness”- Harnesses → Add harness.
- Set Name (e.g.
Ollama) and switch Type to LLM API. - Choose Provider, fill in Base URL, and either API key or …or environment variable.
- Add model ids under Models (press Enter after each; click ★ to set the default), or click Fetch latest.
- Save. The harness appears under LLM APIs and in every step’s Harness menu, marked
(API).
Timeouts and cancellation
Section titled “Timeouts and cancellation”The step’s Timeout (minutes) applies to API calls too; a timed-out call fails with Timed out after N min. Cancelling a run aborts the in-flight request.