Connect OpenCode, pi, and DeepSeek Harness to the same model and compare the differences between model-agnostic coding agent harnesses

Connect OpenCode, pi, and DeepSeek Harness to the same model and compare the differences between model-agnostic coding agent harnesses

I researched the design philosophies and behavioral differences of three coding agent harnesses that allow model swapping: OpenCode, pi, and DeepSeek Harness. I connected the same model to each and had them solve the same task, then organized the basis for choosing a harness.
2026.09.14

This page has been translated by machine translation. View original

Introduction

Hello, I'm Shimada from Classmethod's Manufacturing Business Technology Department.

In August 2026, DeepSeek released DeepSeek Harness.
There have long been several coding agents that allow model swapping, including Codex CLI released by a model company.
This article covers three tools: OpenCode, which is the most widely used; pi, which I use on a daily basis; and DeepSeek Harness, which is the newest.
I wanted to properly understand what the other two changed and how, and what sets these three apart from Claude Code and Codex CLI, even though I normally use pi.

First, I'll organize the origins and design differences of the three tools.
Next, I'll cover "what happens when you run Claude or OpenAI models on these," split into two background stories worth knowing and one concept about model-harness compatibility.
Finally, I'll connect all three to the same model, have them solve the same task, and observe the differences in how they behave.

The information in this article is current as of September 14, 2026.
The versions confirmed are OpenCode 1.18.30, pi 0.85.1, and DeepSeek Harness 0.1.5-rc.1.
DeepSeek Harness is in developer preview, and the README explicitly states that breaking changes may occur.

What Is a Harness

A harness is a collective term for the layer outside the model itself.
It includes tool definitions, system prompts, context management, session storage, permissions, and sandboxing.
DeepSeek Harness's design document describes this as "Agent = Model + Harness."

A configuration where a harness, receiving a developer's request, holds system prompts, tool definitions, sessions, and permissions, and mediates requests, tool calls, and read/write operations to the working directory with the model
The harness sits outside the model, mediating back-and-forth with the model and read/write operations to the working directory

Claude Code and Codex CLI are harnesses that model companies built for their own models.
The three tools covered in this article are harnesses built without assuming a specific model.
Both share the same structure, and some like Codex CLI allow pointing to other companies' models via configuration, but the difference lies in whether the model and harness were designed simultaneously by the same company.
This difference is the premise for the discussion in the latter half.

Origins and Design of the Three Tools

OpenCode pi DeepSeek Harness (dsh)
Author Anomaly Innovations Mario Zechner DeepSeek
Released 2025 2025 August 13, 2026 (developer preview)
License MIT MIT MIT
Design philosophy All-inclusive Minimal Everything as plugins
Default tools bash, edit, write, read, grep, glob, lsp, apply_patch, skill, todowrite, webfetch, websearch, question read, write, edit, bash bash, read, write, edit, glob, grep, str_replace_editor, skill, todo, subagent, web, and others
Omitted features Nothing in particular MCP, sub-agents, plan mode, permission popups, TODO, background bash Fixed workflows
Extension method plugins, MCP, agents, skills TypeScript extensions, Skills, prompt templates, themes Plugin swapping on Cordis
Sessions Client/server. Shared across TUI, desktop, and IDE JSONL tree. Branching and rewinding Append log. resume, fork, search, replay
Non-interactive execution opencode run pi -p dsh --profile headless

OpenCode Is an All-Inclusive Tool Aimed at Being the OSS Version of Claude Code

OpenCode is a harness developed by Anomaly Innovations, the company behind SST, aiming for the same usability as Claude Code.
It comes with LSP auto-loading, MCP, multi-agent support, and plugins from the start.
Its client/server architecture, where a single server process drives the TUI, desktop app, and IDE extension, also allows running opencode serve as a background process and connecting from a separate client.

The OpenCode TUI with the model selector open, showing providers including weak-only in the Switchyard provider
The OpenCode TUI. Adding a provider makes it appear directly in the model selector

Some articles describe OpenCode as written in Go, but this reflects a confusion about its history.
The original Go implementation moved to Charm and became a separate project called Crush, while the current OpenCode is a rewrite in TypeScript.

In terms of pricing, there are three paid plans: Zen, which charges per model usage; Black, which accesses Claude through an API-billed gateway; and Go, which is exclusively for open-weight models.
The reason for this structure is explained in the next section.

pi Is a Minimal Harness with Only Four Tools

pi is a harness developed by Mario Zechner, known as the author of libGDX.
Its default tools are just four: read, write, edit, and bash, and its system prompt is kept short as well.

pi intentionally omits MCP, sub-agents, plan mode, permission popups, TODO, and background bash.
The README explains the reason for each omission, taking the position that MCP can be replaced by CLI tools or extensions, sub-agents by tmux or extensions, and permissions by containers.
It's designed so that missing functionality can be added via TypeScript extensions, and in fact there are packages published around pi, including MCP adapters, sub-agent support, and permission control extensions.

Sessions are saved as a JSONL tree, allowing you to branch from any point and continue from there.
Compaction is also performed without deleting history, so you can return to the original branch.

The pi /tree screen showing multiple branches from a single session displayed as a tree
pi's /tree. You can reselect a branched path and continue from there

DeepSeek Harness Is a Model-Agnostic Harness Released by a Model Company

DeepSeek Harness is a harness released under the MIT license by DeepSeek on August 13, 2026, with the command name dsh.
Built on a dependency injection framework called Cordis, it treats model adapters, tool registration, sessions, agent loops, sandboxing, and the UI all as plugins.
In the 0.1.5-rc.1 I installed locally, over 300 packages were bundled under the @deepseek-ai scope.

One thing I found surprising upon investigation is that the multi-provider layer of DeepSeek Harness is a pi package.
A plugin called dsh-llm-pi-ai depends on @earendil-works/pi-ai from pi's repository, and connections to OpenAI-compatible and Anthropic-compatible endpoints are handled there.
DeepSeek's own models have a separate dedicated adapter alongside this.

Another notable feature is that plugins are bundled to load hooks.json from both Claude Code and Codex, respectively.
This means room was provided from the start to bring over configuration assets written for existing harnesses.

The DeepSeek Harness Web UI showing model selection with models from an added provider listed
dsh's Web UI. After adding a provider, you can select models just like in OpenCode or pi

What Happens When You Run Claude or OpenAI Models on These

You may come across statements like "it's not appropriate to run Claude or OpenAI models on these harnesses."
Upon investigation, I found that this single statement mixes together discussions of different natures.
One involves the background behind decisions made by Anthropic and OpenAI respectively.
The other is the concept of model-harness compatibility, which serves as a basis for choosing a harness.
I'll cover each in turn.

Anthropic Closed the Door on Subscription Misuse via Policy

On January 9, 2026, Anthropic blocked server-side the pathway by which OAuth tokens issued for Claude Pro or Max subscriptions were used from third-party tools.
Users of OpenCode, Cline, and Roo Code were affected.
In February, the Legal and compliance documentation for Claude Code was revised to explicitly state that third-party developers must use API keys issued through Claude Console, and must not route requests on behalf of users using credentials from Free, Pro, or Max plans.
On April 4, the subscription usage allowance itself became unavailable for use with third-party harnesses.

The reasons Anthropic cited were: fixed-plan compute resources were being consumed through autonomous loops without going through pay-as-you-go billing; headers were being sent that impersonated Claude Code clients to pass authentication; and the resulting technical instability.

An important distinction here is that what was prohibited was use via subscriptions, not Claude itself.
The pay-as-you-go path using an API key still works.
OpenCode Black is a product that sells Claude through an API-billed gateway.
In other words, the accurate statement is not that Claude "cannot be used" with these harnesses, but that it "cannot be used at a flat rate, making it more expensive."

OpenAI Permitted Use from Third-Party Harnesses

OpenAI moved in the opposite direction.
OpenAI executives Tibo Sottiaux and Sam Altman have publicly stated that ChatGPT Plus and higher subscriptions may be used from third-party harnesses such as pi and OpenCode.
Sottiaux also noted that pi and OpenCode each account for 5% of Codex traffic.
There is no explicit permission written in the terms of service, but from a management perspective, the direction is permissive.

Therefore, writing that Claude and OpenAI should collectively not be placed on third-party harnesses does not align with the facts.
The policy discussion is limited to Anthropic.

Model-Harness Compatibility Shows Up as Benchmark Differences

Separate from the policy background, there is a concept that serves as a basis for choosing a harness.
This is the idea that a model's performance is determined not by the model alone, but by the combination of the model and the harness.

You may come across explanations that each company's models undergo post-training on their own harnesses, with the tool vocabulary baked into the weights.
The clearest example is file editing tools.
Codex CLI edits using a patch format called apply_patch, while Claude Code edits by passing old_string and new_string for string replacement to an Edit tool.
Even the names of context files they recognize differ: CLAUDE.md versus AGENTS.md.
The explanation goes that when a model is given a tool format it's unfamiliar with, it uses more tokens for reasoning and makes more mistakes.

There is published data comparing the same model and benchmark with only the harness changed.
Terminal-Bench is a benchmark that has models solve multi-step tasks in a terminal, and through version 2.1, the operators also published scores for their reference implementation agent, Terminus 2, which they use as a comparison baseline.

Condition Terminal-Bench 2.1
Claude Opus 4.6 + Claude Code 70.1%
Claude Opus 4.6 + Terminus 2 63.8%

Even with the same model, there is a 6-point gap between the first-party harness and a general-purpose harness.
As of September 2026, Terminal-Bench 4.0 is the latest version, but from 3.0 onward, only results from running each model with a single harness are listed, with no rows comparing the same model across different harnesses.
Since 2.1 is the last version where Claude Code and Terminus 2 appear side by side for the same model, I'm using the 2.1 numbers here.
The arXiv paper Harness-Bench also ran 5,194 executions across combinations of multiple models and harnesses, concluding that completion rates and failure patterns vary significantly by model-harness combination.

However, the causal claim that "post-training causes the difference" is an inference based on what is publicly available.
The existence of a difference has been measured, but I could not find experiments that isolate the cause.

Two things can be said from this.
When you run Claude or GPT on a harness that isn't the lab's own, performance varies by combination, so you need to re-measure.
And open-weight models that either don't have a dedicated harness or have one without lock-in (GLM, Kimi, DeepSeek, Qwen) can be combined with model-agnostic harnesses without issue.
I think OpenCode Go emerging as a plan exclusively for open-weight models is a reflection of this situation.

How to Classify Similar Tools

Sorting similar-looking tools by their relationship to a model looks like this:

  • Made by a model company for their own models: Claude Code, Codex CLI, Gemini CLI, Copilot CLI. These are designed alongside the model. Codex CLI and Gemini CLI have open source code, and Codex CLI allows specifying other companies' endpoints in configuration, but the design origin is their own model.
  • Model-swappable OSS for the terminal: OpenCode, pi, Crush, Aider, Goose. The main subject of this article.
  • IDE extensions: Cline, Roo Code, Kilo Code, Cursor. These were affected by the January blockade in the same way as OpenCode.
  • Things that run autonomously as background services: OpenHands, OpenClaw, Hermes. Used for tasks received from sources like Slack.

DeepSeek Harness sits between the first and second categories.
It was released by a model company, but the layer handling connections to other companies' models is a pi package, with the adapter for its own model placed separately alongside it.
The fact that tools and sessions can also be swapped as plugins is also a characteristic of the second category.
Including the fact that they released the harness for free in the same week they raised API prices for their own model, viewing it as a contrast between Anthropic, which locks down its harness, and DeepSeek and OpenAI, which open theirs, makes the positioning easier to understand.

Connecting All Three to the Same Model and Solving the Same Task

To see how the design differences manifest in actual behavior, I connected all three to the same model and had them solve the same task in non-interactive mode.

Connection Target

I set the model endpoint to an OpenAI-compatible endpoint on NeMo Switchyard running in a local Docker container.
I specified Switchyard's weak-only route identically from all three, and confirmed from the responseModel in session logs that the actual model is DeepSeek V4 Flash (0731) on Fireworks AI.
Since the model is the same, any differences that appear are differences on the harness side.

A configuration where all three of OpenCode, pi, and DeepSeek Harness connect via OpenAI-compatible API to the weak-only route of NeMo Switchyard, which relays to DeepSeek V4 Flash on Fireworks AI
All three connect to the same Switchyard route, with the underlying model standardized to the same one

For all three, registering an OpenAI-compatible endpoint requires only a single configuration file.
The format differs for each.

OpenCode is written in opencode.json directly under the project.
For Chat Completions format endpoints, specify @ai-sdk/openai-compatible.

opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "switchyard": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Switchyard",
      "options": {
        "baseURL": "http://127.0.0.1:4100/v1",
        "apiKey": "{env:SWITCHYARD_API_KEY}"
      },
      "models": {
        "weak-only": { "name": "weak-only", "limit": { "context": 262144, "output": 16384 } }
      }
    }
  },
  "model": "switchyard/weak-only"
}

pi is written in ~/.pi/agent/models.json.

~/.pi/agent/models.json
{
  "providers": {
    "switchyard": {
      "baseUrl": "http://127.0.0.1:4100/v1",
      "api": "openai-completions",
      "apiKey": "$SWITCHYARD_API_KEY",
      "models": [
        { "id": "weak-only", "input": ["text"], "contextWindow": 262144, "maxOutputTokens": 16384 }
      ]
    }
  }
}

DeepSeek Harness is written in ~/.dsh/settings.yaml (the location can be changed with DSH_HOME), using plugin IDs as keys.
The fact that the api: openai-completions specification shares the same name as pi is because, as mentioned earlier, the underlying implementation is a pi package.

~/.dsh/settings.yaml
llm-pi-ai:
  providers:
    switchyard:
      apiKeyEnv: SWITCHYARD_API_KEY
      api: openai-completions
      baseURL: http://127.0.0.1:4100/v1
      models:
        - id: weak-only
          contextWindow: 262144
          maxTokens: 16384
          input: [text]
agent-default-model:
  provider: switchyard
  model: weak-only
session-telemetry-otel:
  mode: DISABLED

Task

The task was a small Python CLI fix.
I planted two problems in a script that counts word frequency in a text file: punctuation attaching to words, and uppercase and lowercase being counted separately.

TASK.md
wordfreq.py has 2 problems. Punctuation before and after words is being counted as part of the word, and uppercase and lowercase are being counted as separate words.
Please do the following:
1. Fix count_words to strip punctuation and count case-insensitively
2. Add a `--top N` option to display only the top N results
3. Add pytest tests to tests/test_wordfreq.py (at least 3 cases covering the 2 fixed issues and --top)
4. Run pytest and confirm all tests pass
When finished, briefly report the files changed and the confirmation results.

I prepared three identical directories and passed this text to each in non-interactive mode.

opencode run --pure --format json -m switchyard/weak-only "$(cat TASK.md)"
pi -p --mode json --no-extensions --no-skills "$(cat TASK.md)"
dsh --profile headless "$(cat TASK.md)"

OpenCode's --pure and pi's --no-extensions are specified to disable external plugins and compare with the bare harness.
DeepSeek Harness's headless profile runs with the default workspace-write mode (only allowing writes within the working directory).

Results

All three completed the task, reported the changed files and confirmation results, and finished.

OpenCode pi DeepSeek Harness
Time elapsed 68 seconds 45 seconds 71 seconds
Number of requests to model 19 6 13
Tool calls 19 (bash 7, edit 7, read 2, write 2, glob 1) 7 (read 3, bash 2, write 2) 14 (bash 5, read 4, edit 3, write 2)
Input tokens on first request 8,084 1,744 7,626
Total input tokens (including cache reads) 238,856 19,899 148,638
Total output tokens 4,617 2,007 4,277
Tests added 8 6 8
pytest result reported 8 passed 6 passed 8 passed

All three produced working outputs.
The reported results also matched the actual results.
However, there are differences in the numbers.

The size of the first request directly reflects the size of the system prompt and tool definitions.
Since the task text is the same for all three, the difference between OpenCode's 8,084 tokens and pi's 1,744 tokens is the difference in the preamble that the harness passes to the model.
DeepSeek Harness had a volume close to OpenCode.
This difference accumulates with each request, so in total input tokens, OpenCode is about 12 times that of pi.
If using a route with prompt caching, the cost impact will be smaller, but when running with a local model, it adds directly to the prefill time for every request.

Guidelines for Choosing Between the Three

These are guidelines based on my hands-on experience with all three.

  • OpenCode: When you want to keep the Claude Code feel while just swapping the model. LSP, MCP, and IDE integration are available from the start. On the other hand, the preamble per request is large, and that weight becomes visible with local models.
  • pi: When you want to keep the preamble small, or when you want to understand the harness internals yourself. Since the premise is adding missing functionality through extensions, there's some upfront effort in choosing what to add yourself. Because you start from a minimal configuration and add from there, this is the one I use most.
  • DeepSeek Harness: When you want to experiment by swapping tools, sessions, and sandboxing at the plugin level. Even though it's a developer preview, I found the Web UI to be polished and practical.

As I mentioned earlier, none of these allow using Claude at a flat rate, so if you want to use Anthropic's models, it's best to use Claude Code or Claude Code on the web.
If you're using these harnesses, the options as of September 2026 are: Claude with API key pay-as-you-go, OpenAI models with a ChatGPT subscription, and everything else with open-weight models via API or locally.

Closing

OpenCode, pi, and DeepSeek Harness are all model-swappable harnesses, but their design directions differed: all-inclusive, minimal, and everything-as-plugins.
Through my research, I found that even with closed models like Opus, benchmark scores change depending on which harness is used.
When I connected all three to the same model and actually ran them, differences also appeared in elapsed time and the number of tokens required.
I also didn't know until I looked into it that DeepSeek Harness's provider layer is a pi package.

Beyond the three covered in this article, coding agent harnesses continue to appear one after another.
If any of them catch your interest, try them out in practice and find what works for your own use case.

References


Claudeならクラスメソッドにお任せください

クラスメソッドは、Anthropic社とリセラー契約を締結しています。各種製品ガイドから、業種別の活用法、フェーズごとのお悩み解決などサービス支援ページにまとめております。まずはご覧いただき、お気軽にご相談ください。

サービス詳細を見る

Share this article

AI白書