Skip to content

API harnesses

API harnesses call an LLM API directly instead of running a coding agent. A step on an API harness sends one request (system prompt + your rendered prompt) and uses the response text as the step’s output.

Two API harnesses are built in: Claude API and OpenAI-compatible API. To use several OpenAI-compatible providers at once (say OpenAI and a local Ollama), add more with Add harness and set Type to LLM API.

Field Meaning
Provider Anthropic (Claude API) or OpenAI-compatible (OpenAI, OpenRouter, Ollama, LM Studio…)
Base URL Anthropic: leave empty for api.anthropic.com. OpenAI-compatible: the URL that /chat/completions is appended to, e.g. http://localhost:11434/v1. Empty means https://api.openai.com/v1.
API key Stored encrypted with your OS keychain. Takes precedence over the environment variable.
…or environment variable Name of the env var holding the key. Empty means ANTHROPIC_API_KEY (Anthropic) or OPENAI_API_KEY (OpenAI-compatible).
Max output tokens Sent as max_tokens (Anthropic) or max_completion_tokens (OpenAI-compatible).
Models The one-click choices in the step editor, with a ★ default. Fetch latest pulls the provider’s list.

The key lookup order is: API key on the harness → the named env var in the step’s environment (which includes Settings → Execution → Environment variables, project env and step env) → the same env var in the app’s own environment.

Setting Default
Id anthropic-api
Provider Anthropic
Key env var ANTHROPIC_API_KEY
Max output tokens 32000
Models claude-opus-5, claude-opus-5-5, claude-fable-5-1, claude-sonnet-5, claude-haiku-4-5
Default model claude-opus-5

How a step runs:

  • Uses the official Anthropic SDK and streams the Messages API response. Text appears in the log line by line as it arrives.
  • Sends the step’s system prompt as system and the rendered prompt as a single user message.
  • Enables adaptive thinking (thinking: { type: "adaptive" }).
  • If no key is found at all, the SDK falls back to its own credential lookup.
  • Records input, output, cache-read and cache-write tokens. The API doesn’t return a price, so the step’s cost stays $0.
  • Fails with The model declined this request (refusal) on a refusal stop reason, and with Response hit max_tokens if the answer was cut off. Raise Max output tokens in that case.

The Bug fix (fast lane) template uses the Claude API for its triage step.

Fetch latest lists models from the Anthropic Models API.

Setting Default
Id openai-compatible
Base URL https://api.openai.com/v1
Key env var OPENAI_API_KEY
Max output tokens 16000
Models gpt-5, gpt-5-mini
Default model gpt-5

How a step runs:

  • POST {base URL}/chat/completions with model, messages (a system message if the step has a system prompt, then the user prompt) and max_completion_tokens.
  • Sends Authorization: Bearer <key> only if a key was found, so keyless local servers work.
  • Not streamed: the whole answer is logged when it arrives.
  • Uses choices[0].message.content as the output and records prompt_tokens / completion_tokens if the server returns usage. No cost.
  • The request timeout is the step’s timeout, or 10 minutes if the step has none.

Setting Max output tokens to 0 leaves max_completion_tokens out of the request, which helps with servers that reject that field.

Fetch latest calls GET {base URL}/models (with the bearer key if any) and reads the data array.

Field Value
Base URL empty, or https://api.openai.com/v1
Key an OpenAI API key, or OPENAI_API_KEY
Models e.g. gpt-5, gpt-5-mini — or click Fetch latest
  1. Harnesses → Add harness.
  2. Set Name (e.g. Ollama) and switch Type to LLM API.
  3. Choose Provider, fill in Base URL, and either API key or …or environment variable.
  4. Add model ids under Models (press Enter after each; click ★ to set the default), or click Fetch latest.
  5. Save. The harness appears under LLM APIs and in every step’s Harness menu, marked (API).

The step’s Timeout (minutes) applies to API calls too; a timed-out call fails with Timed out after N min. Cancelling a run aborts the in-flight request.