
Connect OpenCode, pi, and DeepSeek Harness to the same model and compare the differences between model-agnostic coding agent harnesses
This page has been translated by machine translation. View original
Introduction
Hello, I'm Shimada from Classmethod's Manufacturing Business Technology Department.
In August 2026, DeepSeek released DeepSeek Harness.
There have long been several coding agents that allow model swapping, including Codex CLI released by a model company.
This article covers three tools: OpenCode, which is the most widely used; pi, which I use on a daily basis; and DeepSeek Harness, which is the newest.
I wanted to properly understand what the other two changed and how, and what sets these three apart from Claude Code and Codex CLI, even though I normally use pi.
First, I'll organize the origins and design differences of the three tools.
Next, I'll cover "what happens when you run Claude or OpenAI models on these," split into two background stories worth knowing and one concept about model-harness compatibility.
Finally, I'll connect all three to the same model, have them solve the same task, and observe the differences in how they behave.
The information in this article is current as of September 14, 2026.
The versions confirmed are OpenCode 1.18.30, pi 0.85.1, and DeepSeek Harness 0.1.5-rc.1.
DeepSeek Harness is in developer preview, and the README explicitly states that breaking changes may occur.
What Is a Harness
A harness is a collective term for the layer outside the model itself.
It includes tool definitions, system prompts, context management, session storage, permissions, and sandboxing.
DeepSeek Harness's design document describes this as "Agent = Model + Harness."

The harness sits outside the model, mediating back-and-forth with the model and read/write operations to the working directory
Claude Code and Codex CLI are harnesses that model companies built for their own models.
The three tools covered in this article are harnesses built without assuming a specific model.
Both share the same structure, and some like Codex CLI allow pointing to other companies' models via configuration, but the difference lies in whether the model and harness were designed simultaneously by the same company.
This difference is the premise for the discussion in the latter half.
Origins and Design of the Three Tools
| OpenCode | pi | DeepSeek Harness (dsh) |
|
|---|---|---|---|
| Author | Anomaly Innovations | Mario Zechner | DeepSeek |
| Released | 2025 | 2025 | August 13, 2026 (developer preview) |
| License | MIT | MIT | MIT |
| Design philosophy | All-inclusive | Minimal | Everything as plugins |
| Default tools | bash, edit, write, read, grep, glob, lsp, apply_patch, skill, todowrite, webfetch, websearch, question | read, write, edit, bash | bash, read, write, edit, glob, grep, str_replace_editor, skill, todo, subagent, web, and others |
| Omitted features | Nothing in particular | MCP, sub-agents, plan mode, permission popups, TODO, background bash | Fixed workflows |
| Extension method | plugins, MCP, agents, skills | TypeScript extensions, Skills, prompt templates, themes | Plugin swapping on Cordis |
| Sessions | Client/server. Shared across TUI, desktop, and IDE | JSONL tree. Branching and rewinding | Append log. resume, fork, search, replay |
| Non-interactive execution | opencode run |
pi -p |
dsh --profile headless |
OpenCode Is an All-Inclusive Tool Aimed at Being the OSS Version of Claude Code
OpenCode is a harness developed by Anomaly Innovations, the company behind SST, aiming for the same usability as Claude Code.
It comes with LSP auto-loading, MCP, multi-agent support, and plugins from the start.
Its client/server architecture, where a single server process drives the TUI, desktop app, and IDE extension, also allows running opencode serve as a background process and connecting from a separate client.

The OpenCode TUI. Adding a provider makes it appear directly in the model selector
Some articles describe OpenCode as written in Go, but this reflects a confusion about its history.
The original Go implementation moved to Charm and became a separate project called Crush, while the current OpenCode is a rewrite in TypeScript.
In terms of pricing, there are three paid plans: Zen, which charges per model usage; Black, which accesses Claude through an API-billed gateway; and Go, which is exclusively for open-weight models.
The reason for this structure is explained in the next section.
pi Is a Minimal Harness with Only Four Tools
pi is a harness developed by Mario Zechner, known as the author of libGDX.
Its default tools are just four: read, write, edit, and bash, and its system prompt is kept short as well.
pi intentionally omits MCP, sub-agents, plan mode, permission popups, TODO, and background bash.
The README explains the reason for each omission, taking the position that MCP can be replaced by CLI tools or extensions, sub-agents by tmux or extensions, and permissions by containers.
It's designed so that missing functionality can be added via TypeScript extensions, and in fact there are packages published around pi, including MCP adapters, sub-agent support, and permission control extensions.
Sessions are saved as a JSONL tree, allowing you to branch from any point and continue from there.
Compaction is also performed without deleting history, so you can return to the original branch.

pi's /tree. You can reselect a branched path and continue from there
DeepSeek Harness Is a Model-Agnostic Harness Released by a Model Company
DeepSeek Harness is a harness released under the MIT license by DeepSeek on August 13, 2026, with the command name dsh.
Built on a dependency injection framework called Cordis, it treats model adapters, tool registration, sessions, agent loops, sandboxing, and the UI all as plugins.
In the 0.1.5-rc.1 I installed locally, over 300 packages were bundled under the @deepseek-ai scope.
One thing I found surprising upon investigation is that the multi-provider layer of DeepSeek Harness is a pi package.
A plugin called dsh-llm-pi-ai depends on @earendil-works/pi-ai from pi's repository, and connections to OpenAI-compatible and Anthropic-compatible endpoints are handled there.
DeepSeek's own models have a separate dedicated adapter alongside this.
Another notable feature is that plugins are bundled to load hooks.json from both Claude Code and Codex, respectively.
This means room was provided from the start to bring over configuration assets written for existing harnesses.

dsh's Web UI. After adding a provider, you can select models just like in OpenCode or pi
What Happens When You Run Claude or OpenAI Models on These
You may come across statements like "it's not appropriate to run Claude or OpenAI models on these harnesses."
Upon investigation, I found that this single statement mixes together discussions of different natures.
One involves the background behind decisions made by Anthropic and OpenAI respectively.
The other is the concept of model-harness compatibility, which serves as a basis for choosing a harness.
I'll cover each in turn.
Anthropic Closed the Door on Subscription Misuse via Policy
On January 9, 2026, Anthropic blocked server-side the pathway by which OAuth tokens issued for Claude Pro or Max subscriptions were used from third-party tools.
Users of OpenCode, Cline, and Roo Code were affected.
In February, the Legal and compliance documentation for Claude Code was revised to explicitly state that third-party developers must use API keys issued through Claude Console, and must not route requests on behalf of users using credentials from Free, Pro, or Max plans.
On April 4, the subscription usage allowance itself became unavailable for use with third-party harnesses.
The reasons Anthropic cited were: fixed-plan compute resources were being consumed through autonomous loops without going through pay-as-you-go billing; headers were being sent that impersonated Claude Code clients to pass authentication; and the resulting technical instability.
An important distinction here is that what was prohibited was use via subscriptions, not Claude itself.
The pay-as-you-go path using an API key still works.
OpenCode Black is a product that sells Claude through an API-billed gateway.
In other words, the accurate statement is not that Claude "cannot be used" with these harnesses, but that it "cannot be used at a flat rate, making it more expensive."
OpenAI Permitted Use from Third-Party Harnesses
OpenAI moved in the opposite direction.
OpenAI executives Tibo Sottiaux and Sam Altman have publicly stated that ChatGPT Plus and higher subscriptions may be used from third-party harnesses such as pi and OpenCode.
Sottiaux also noted that pi and OpenCode each account for 5% of Codex traffic.
There is no explicit permission written in the terms of service, but from a management perspective, the direction is permissive.
Therefore, writing that Claude and OpenAI should collectively not be placed on third-party harnesses does not align with the facts.
The policy discussion is limited to Anthropic.
Model-Harness Compatibility Shows Up as Benchmark Differences
Separate from the policy background, there is a concept that serves as a basis for choosing a harness.
This is the idea that a model's performance is determined not by the model alone, but by the combination of the model and the harness.
You may come across explanations that each company's models undergo post-training on their own harnesses, with the tool vocabulary baked into the weights.
The clearest example is file editing tools.
Codex CLI edits using a patch format called apply_patch, while Claude Code edits by passing old_string and new_string for string replacement to an Edit tool.
Even the names of context files they recognize differ: CLAUDE.md versus AGENTS.md.
The explanation goes that when a model is given a tool format it's unfamiliar with, it uses more tokens for reasoning and makes more mistakes.
There is published data comparing the same model and benchmark with only the harness changed.
Terminal-Bench is a benchmark that has models solve multi-step tasks in a terminal, and through version 2.1, the operators also published scores for their reference implementation agent, Terminus 2, which they use as a comparison baseline.
| Condition | Terminal-Bench 2.1 |
|---|---|
| Claude Opus 4.6 + Claude Code | 70.1% |
| Claude Opus 4.6 + Terminus 2 | 63.8% |
Even with the same model, there is a 6-point gap between the first-party harness and a general-purpose harness.
As of September 2026, Terminal-Bench 4.0 is the latest version, but from 3.0 onward, only results from running each model with a single harness are listed, with no rows comparing the same model across different harnesses.
Since 2.1 is the last version where Claude Code and Terminus 2 appear side by side for the same model, I'm using the 2.1 numbers here.
The arXiv paper Harness-Bench also ran 5,194 executions across combinations of multiple models and harnesses, concluding that completion rates and failure patterns vary significantly by model-harness combination.
However, the causal claim that "post-training causes the difference" is an inference based on what is publicly available.
The existence of a difference has been measured, but I could not find experiments that isolate the cause.
Two things can be said from this.
When you run Claude or GPT on a harness that isn't the lab's own, performance varies by combination, so you need to re-measure.
And open-weight models that either don't have a dedicated harness or have one without lock-in (GLM, Kimi, DeepSeek, Qwen) can be combined with model-agnostic harnesses without issue.
I think OpenCode Go emerging as a plan exclusively for open-weight models is a reflection of this situation.
How to Classify Similar Tools
Sorting similar-looking tools by their relationship to a model looks like this:
- Made by a model company for their own models: Claude Code, Codex CLI, Gemini CLI, Copilot CLI. These are designed alongside the model. Codex CLI and Gemini CLI have open source code, and Codex CLI allows specifying other companies' endpoints in configuration, but the design origin is their own model.
- Model-swappable OSS for the terminal: OpenCode, pi, Crush, Aider, Goose. The main subject of this article.
- IDE extensions: Cline, Roo Code, Kilo Code, Cursor. These were affected by the January blockade in the same way as OpenCode.
- Things that run autonomously as background services: OpenHands, OpenClaw, Hermes. Used for tasks received from sources like Slack.
DeepSeek Harness sits between the first and second categories.
It was released by a model company, but the layer handling connections to other companies' models is a pi package, with the adapter for its own model placed separately alongside it.
The fact that tools and sessions can also be swapped as plugins is also a characteristic of the second category.
Including the fact that they released the harness for free in the same week they raised API prices for their own model, viewing it as a contrast between Anthropic, which locks down its harness, and DeepSeek and OpenAI, which open theirs, makes the positioning easier to understand.
Connecting All Three to the Same Model and Solving the Same Task
To see how the design differences manifest in actual behavior, I connected all three to the same model and had them solve the same task in non-interactive mode.
Connection Target
I set the model endpoint to an OpenAI-compatible endpoint on NeMo Switchyard running in a local Docker container.
I specified Switchyard's weak-only route identically from all three, and confirmed from the responseModel in session logs that the actual model is DeepSeek V4 Flash (0731) on Fireworks AI.
Since the model is the same, any differences that appear are differences on the harness side.

All three connect to the same Switchyard route, with the underlying model standardized to the same one
For all three, registering an OpenAI-compatible endpoint requires only a single configuration file.
The format differs for each.
OpenCode is written in opencode.json directly under the project.
For Chat Completions format endpoints, specify @ai-sdk/openai-compatible.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"switchyard": {
"npm": "@ai-sdk/openai-compatible",
"name": "Switchyard",
"options": {
"baseURL": "http://127.0.0.1:4100/v1",
"apiKey": "{env:SWITCHYARD_API_KEY}"
},
"models": {
"weak-only": { "name": "weak-only", "limit": { "context": 262144, "output": 16384 } }
}
}
},
"model": "switchyard/weak-only"
}
pi is written in ~/.pi/agent/models.json.
{
"providers": {
"switchyard": {
"baseUrl": "http://127.0.0.1:4100/v1",
"api": "openai-completions",
"apiKey": "$SWITCHYARD_API_KEY",
"models": [
{ "id": "weak-only", "input": ["text"], "contextWindow": 262144, "maxOutputTokens": 16384 }
]
}
}
}
DeepSeek Harness is written in ~/.dsh/settings.yaml (the location can be changed with DSH_HOME), using plugin IDs as keys.
The fact that the api: openai-completions specification shares the same name as pi is because, as mentioned earlier, the underlying implementation is a pi package.
llm-pi-ai:
providers:
switchyard:
apiKeyEnv: SWITCHYARD_API_KEY
api: openai-completions
baseURL: http://127.0.0.1:4100/v1
models:
- id: weak-only
contextWindow: 262144
maxTokens: 16384
input: [text]
agent-default-model:
provider: switchyard
model: weak-only
session-telemetry-otel:
mode: DISABLED
Task
The task was a small Python CLI fix.
I planted two problems in a script that counts word frequency in a text file: punctuation attaching to words, and uppercase and lowercase being counted separately.
wordfreq.py has 2 problems. Punctuation before and after words is being counted as part of the word, and uppercase and lowercase are being counted as separate words.
Please do the following:
1. Fix count_words to strip punctuation and count case-insensitively
2. Add a `--top N` option to display only the top N results
3. Add pytest tests to tests/test_wordfreq.py (at least 3 cases covering the 2 fixed issues and --top)
4. Run pytest and confirm all tests pass
When finished, briefly report the files changed and the confirmation results.
I prepared three identical directories and passed this text to each in non-interactive mode.
opencode run --pure --format json -m switchyard/weak-only "$(cat TASK.md)"
pi -p --mode json --no-extensions --no-skills "$(cat TASK.md)"
dsh --profile headless "$(cat TASK.md)"
OpenCode's --pure and pi's --no-extensions are specified to disable external plugins and compare with the bare harness.
DeepSeek Harness's headless profile runs with the default workspace-write mode (only allowing writes within the working directory).
Results
All three completed the task, reported the changed files and confirmation results, and finished.
| OpenCode | pi | DeepSeek Harness | |
|---|---|---|---|
| Time elapsed | 68 seconds | 45 seconds | 71 seconds |
| Number of requests to model | 19 | 6 | 13 |
| Tool calls | 19 (bash 7, edit 7, read 2, write 2, glob 1) | 7 (read 3, bash 2, write 2) | 14 (bash 5, read 4, edit 3, write 2) |
| Input tokens on first request | 8,084 | 1,744 | 7,626 |
| Total input tokens (including cache reads) | 238,856 | 19,899 | 148,638 |
| Total output tokens | 4,617 | 2,007 | 4,277 |
| Tests added | 8 | 6 | 8 |
| pytest result reported | 8 passed | 6 passed | 8 passed |
All three produced working outputs.
The reported results also matched the actual results.
However, there are differences in the numbers.
The size of the first request directly reflects the size of the system prompt and tool definitions.
Since the task text is the same for all three, the difference between OpenCode's 8,084 tokens and pi's 1,744 tokens is the difference in the preamble that the harness passes to the model.
DeepSeek Harness had a volume close to OpenCode.
This difference accumulates with each request, so in total input tokens, OpenCode is about 12 times that of pi.
If using a route with prompt caching, the cost impact will be smaller, but when running with a local model, it adds directly to the prefill time for every request.
Guidelines for Choosing Between the Three
These are guidelines based on my hands-on experience with all three.
- OpenCode: When you want to keep the Claude Code feel while just swapping the model. LSP, MCP, and IDE integration are available from the start. On the other hand, the preamble per request is large, and that weight becomes visible with local models.
- pi: When you want to keep the preamble small, or when you want to understand the harness internals yourself. Since the premise is adding missing functionality through extensions, there's some upfront effort in choosing what to add yourself. Because you start from a minimal configuration and add from there, this is the one I use most.
- DeepSeek Harness: When you want to experiment by swapping tools, sessions, and sandboxing at the plugin level. Even though it's a developer preview, I found the Web UI to be polished and practical.
As I mentioned earlier, none of these allow using Claude at a flat rate, so if you want to use Anthropic's models, it's best to use Claude Code or Claude Code on the web.
If you're using these harnesses, the options as of September 2026 are: Claude with API key pay-as-you-go, OpenAI models with a ChatGPT subscription, and everything else with open-weight models via API or locally.
Closing
OpenCode, pi, and DeepSeek Harness are all model-swappable harnesses, but their design directions differed: all-inclusive, minimal, and everything-as-plugins.
Through my research, I found that even with closed models like Opus, benchmark scores change depending on which harness is used.
When I connected all three to the same model and actually ran them, differences also appeared in elapsed time and the number of tokens required.
I also didn't know until I looked into it that DeepSeek Harness's provider layer is a pi package.
Beyond the three covered in this article, coding agent harnesses continue to appear one after another.
If any of them catch your interest, try them out in practice and find what works for your own use case.
References
- OpenCode
- OpenCode Docs: Providers
- OpenCode Docs: Tools
- pi coding agent README (GitHub)
- deepseek-ai/deepseek-harness (GitHub)
- DeepSeek Harness launches as open source rival to Claude Code (VentureBeat)
- Anthropic cracks down on unauthorized Claude usage by third-party harnesses (VentureBeat)
- Anthropic officially bans using subscription authentication for third-party Claude use (AlternativeTo)
- ChatGPT Plus: Enjoy $200 of Tokens for $20 While It Lasts (manifest.build)
- Codex for Open Source (OpenAI Developers)
- Model-Harness-Fit (Nicolas Bustamante)
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows (arXiv)
- NeMo Switchyard (GitHub)
- Part 2: Routing Requests with NeMo Switchyard and a Post-Trained Judge
- Terminal-Bench 2.1 (tbench.ai news)
