Skip to content

Local models with Ollama

You can run some or all of a pipeline on models served by Ollama on your own machine. There are two ways, and they suit different steps:

Approach Harness Can edit code? Good for
Ollama’s OpenAI-compatible API an LLM API harness No — one prompt in, text out triage, plans, summaries, reviewing a diff you pass in
A coding CLI pointed at Ollama Aider (or OpenCode / Goose configured for Ollama) Yes implementing changes

The same steps work for LM Studio (http://localhost:1234/v1) and vLLM (http://<host>:8000/v1).

Terminal window
# install from ollama.com, then:
ollama pull qwen2.5-coder:14b
ollama serve # if it isn't already running as a background service
curl http://localhost:11434/v1/models

You can reconfigure the built-in OpenAI-compatible API harness, but adding a separate one keeps OpenAI available too.

  1. Harnesses → Add harness.
  2. Name: Ollama. Type: LLM API.
  3. Provider: OpenAI-compatible (OpenAI, OpenRouter, Ollama, LM Studio…).
  4. Base URL: http://localhost:11434/v1. Leave API key and …or environment variable empty.
  5. Under Models, click Fetch latest to list your local models, and click ★ on the one to use by default.
  6. Save. The card shows API key configured: a localhost base URL counts as available without a key.

If the server rejects the request because of the output-token field, set Max output tokens to 0; Jimothy then leaves max_completion_tokens out.

In any pipeline, set an agent step’s Harness to Ollama (API) and pick a model. Good candidates: Triage, Plan (when you paste the relevant context into the prompt), and Review.

API steps can’t look at the repository, so give a local reviewer the diff. Add a shell step before it:

Step Type Setting
diff — Collect diff Shell git diff {{run.baseBranch}}...HEAD (runs after Implement)
review — Local review Agent · Ollama runs after diff; prompt below; Pass if output matches (regex) VERDICT:\s*APPROVE; On failure, loop back to implement
You are a strict code reviewer. Review this change for {{issue.key}}: {{issue.title}}
{{issue.description}}
Diff against {{run.baseBranch}}:
{{steps.diff.output | truncate:30000}}
End with exactly one line, VERDICT: APPROVE or VERDICT: CHANGES_REQUESTED,
followed by a numbered list of required changes if any.

Step output is capped at the last 20,000 characters, so very large diffs are trimmed from the start. Keep changes small, or diff only the relevant paths (git diff {{run.baseBranch}}...HEAD -- src/).

Aider can drive Ollama models and edits files in the workspace.

  1. Install Aider: python -m pip install aider-install && aider-install.
  2. Tell Aider where Ollama is: in Settings → Execution → Environment variables add OLLAMA_API_BASE=http://127.0.0.1:11434 and click Save.
  3. Harnesses → Re-detect; Aider should be green.
  4. On your implement step, choose Harness Aider and type the model in the model box: ollama_chat/qwen2.5-coder:14b. (Or add that id to Aider’s Models in the harness editor and ★ it as default.)

Aider receives the prompt through a file (--message-file), so long prompts are fine, and it commits its changes by default. Its output is plain text: you’ll see Aider’s transcript in the logs, but no token or cost figures.

Step Harness · model
Triage Ollama (API) · your model
Implement Aider · ollama_chat/<model>
Run tests Shell · {{vars.testCommand | raw}}, loop back to Implement
Collect diff Shell · git diff {{run.baseBranch}}...HEAD
Local review Ollama (API), verdict + loop back to Implement
Human approval Approval
Open pull request Shell

Nothing here leaves your machine except the final git push. Expect local models to need more loops than frontier models; raise Max loops or keep tasks small.

Harnesses are chosen per step, so a common compromise is local models for cheap, high-volume steps (triage, summaries, first-pass review) and a hosted coding agent (Claude Code, Codex) for implementation.