Local models with Ollama
You can run some or all of a pipeline on models served by Ollama on your own machine. There are two ways, and they suit different steps:
| Approach | Harness | Can edit code? | Good for |
|---|---|---|---|
| Ollama’s OpenAI-compatible API | an LLM API harness | No — one prompt in, text out | triage, plans, summaries, reviewing a diff you pass in |
| A coding CLI pointed at Ollama | Aider (or OpenCode / Goose configured for Ollama) | Yes | implementing changes |
The same steps work for LM Studio (http://localhost:1234/v1) and vLLM (http://<host>:8000/v1).
1. Run Ollama and pull a model
Section titled “1. Run Ollama and pull a model”# install from ollama.com, then:ollama pull qwen2.5-coder:14bollama serve # if it isn't already running as a background servicecurl http://localhost:11434/v1/models2. Add an Ollama API harness
Section titled “2. Add an Ollama API harness”You can reconfigure the built-in OpenAI-compatible API harness, but adding a separate one keeps OpenAI available too.
- Harnesses → Add harness.
- Name:
Ollama. Type: LLM API. - Provider: OpenAI-compatible (OpenAI, OpenRouter, Ollama, LM Studio…).
- Base URL:
http://localhost:11434/v1. Leave API key and …or environment variable empty. - Under Models, click Fetch latest to list your local models, and click ★ on the one to use by default.
- Save. The card shows API key configured: a
localhostbase URL counts as available without a key.
If the server rejects the request because of the output-token field, set Max output tokens to 0; Jimothy then leaves max_completion_tokens out.
3. Use it for text steps
Section titled “3. Use it for text steps”In any pipeline, set an agent step’s Harness to Ollama (API) and pick a model. Good candidates: Triage, Plan (when you paste the relevant context into the prompt), and Review.
API steps can’t look at the repository, so give a local reviewer the diff. Add a shell step before it:
| Step | Type | Setting |
|---|---|---|
diff — Collect diff |
Shell | git diff {{run.baseBranch}}...HEAD (runs after Implement) |
review — Local review |
Agent · Ollama | runs after diff; prompt below; Pass if output matches (regex) VERDICT:\s*APPROVE; On failure, loop back to implement |
You are a strict code reviewer. Review this change for {{issue.key}}: {{issue.title}}
{{issue.description}}
Diff against {{run.baseBranch}}:{{steps.diff.output | truncate:30000}}
End with exactly one line, VERDICT: APPROVE or VERDICT: CHANGES_REQUESTED,followed by a numbered list of required changes if any.Step output is capped at the last 20,000 characters, so very large diffs are trimmed from the start. Keep changes small, or diff only the relevant paths (git diff {{run.baseBranch}}...HEAD -- src/).
4. Use Aider for code-editing steps
Section titled “4. Use Aider for code-editing steps”Aider can drive Ollama models and edits files in the workspace.
- Install Aider:
python -m pip install aider-install && aider-install. - Tell Aider where Ollama is: in Settings → Execution → Environment variables add
OLLAMA_API_BASE=http://127.0.0.1:11434and click Save. - Harnesses → Re-detect; Aider should be green.
- On your implement step, choose Harness Aider and type the model in the model box:
ollama_chat/qwen2.5-coder:14b. (Or add that id to Aider’s Models in the harness editor and ★ it as default.)
Aider receives the prompt through a file (--message-file), so long prompts are fine, and it commits its changes by default. Its output is plain text: you’ll see Aider’s transcript in the logs, but no token or cost figures.
A fully local pipeline
Section titled “A fully local pipeline”| Step | Harness · model |
|---|---|
| Triage | Ollama (API) · your model |
| Implement | Aider · ollama_chat/<model> |
| Run tests | Shell · {{vars.testCommand | raw}}, loop back to Implement |
| Collect diff | Shell · git diff {{run.baseBranch}}...HEAD |
| Local review | Ollama (API), verdict + loop back to Implement |
| Human approval | Approval |
| Open pull request | Shell |
Nothing here leaves your machine except the final git push. Expect local models to need more loops than frontier models; raise Max loops or keep tasks small.
Mixing local and hosted
Section titled “Mixing local and hosted”Harnesses are chosen per step, so a common compromise is local models for cheap, high-volume steps (triage, summaries, first-pass review) and a hosted coding agent (Claude Code, Codex) for implementation.