# How to benchmark inference providers across coding-agent harnesses (https://tenki.cloud/blog/benchmark-inference-providers-coding-agent-harnesses)

> For the complete documentation index, see [llms.txt](https://tenki.cloud/llms.txt)

- Author: Guzman Pintos (Product)
- Published: Oct 5, 2026
- Category: [Sandbox](https://tenki.cloud/blog/category/sandbox)
- Reading time: 6 min

Test any model or inference endpoint across coding-agent harnesses like Claude Code and OpenCode on Terminal-Bench, with every trial in a Tenki sandbox.

A model can pass a benchmark and still fail in the coding agent your customers use, because each agent calls a different part of the inference API. This guide shows how to test one model across six coding agents on Terminal-Bench with [Harbor](https://github.com/harbor-framework/harbor), every trial in its own [Tenki Sandbox](https://tenki.cloud/products/sandbox), and sort each failure by cause.

## Why the same model behaves differently in each harness

A harness is the program around the model: it holds the conversation, offers tools, runs the commands the model asks for, and decides when the task is done. Harnesses use different parts of an inference API, and each part can fail on its own:

| API                                | Used by                          | What it exercises                                                            |
| ---------------------------------- | -------------------------------- | ---------------------------------------------------------------------------- |
| Chat Completions, tool calls       | Goose, mini-swe-agent, Qwen Code | Tool schemas, tool calls in responses                                        |
| Chat Completions, commands in text | Terminus-2                       | Long conversations, following an output format without tools                 |
| OpenAI Responses                   | OpenCode, Codex                  | Conversation state across turns, tool calls and their outputs as input items |
| Anthropic Messages                 | Claude Code                      | Streaming events, tool use blocks, thinking blocks                           |

A quick check that sends a prompt to the endpoint and reads the answer exercises almost none of this. A real harness sends what it actually sends, and that is what the benchmark needs to reproduce.

## 1. Choose tasks the environment can run

Use a benchmark that grades with tests, so a pass means the task was solved. Terminal-Bench 2.0 fits: each task is a container image with its own tests, and the tasks are the kind of terminal work coding agents do.

Before you blame the model for a failure, make sure the task can pass at all. Some tasks break over time, for example when they install package versions that are no longer published. Run Harbor's reference solutions first, and keep the tasks that pass. With Harbor and the Tenki environment installed (see the [Harbor guide](https://tenki.cloud/docs/sandbox/harbor.md)):

```bash
harbor run -d terminal-bench@2.0 -e tenki_harbor:TenkiEnvironment \
  --agent oracle --n-concurrent 16
```

## 2. Run the matrix

[`tenki-harness-evals`](https://github.com/LuxorLabs/tenki-harness-evals) builds the Harbor job, runs it and writes the report. Point it at your endpoint:

```bash
git clone https://github.com/LuxorLabs/tenki-harness-evals.git
cd tenki-harness-evals && uv sync
cp .env.example .env   # Tenki API key, endpoint URL and key
uv run evals models    # the model ids your endpoint serves
uv run evals run --model your-org/your-model \
  --task break-filter-js-from-html --task largest-eigenval   # repeat --task for each task you kept
```

Each trial gets a fresh Tenki VM sized to the task, with the task's container running under Docker inside it, and the VM is deleted when the trial ends. Six harnesses on ten tasks is 60 trials. At 16 at a time, that took about an hour per model in our runs. The [Harbor guide](https://tenki.cloud/docs/sandbox/harbor.md) covers the environment itself, including baking Docker into a [snapshot](https://tenki.cloud/docs/sandbox/snapshots.md) so each trial starts faster.

## 3. Read the report

Every trial lands in one bucket:

* **Passed:** the task's tests passed.
* **Failed:** the harness finished and the tests failed. This is the signal about the model.
* **API error:** the endpoint rejected or broke one of the harness's requests. This is a bug in the endpoint, not the model.
* **Timeout:** the agent ran out of the task's time budget.
* **Infra error:** the sandbox or the task's verifier failed. It counts against neither the model nor the harness.

Here is the report from one of our runs, a model served by an OpenAI-compatible endpoint:

| Harness        | Pass rate | Passed | Failed | API errors | Timeouts |
| -------------- | --------- | ------ | ------ | ---------- | -------- |
| Goose          | 80%       | 8      | 1      | 0          | 1        |
| Terminus-2     | 80%       | 8      | 1      | 0          | 1        |
| mini-swe-agent | 70%       | 7      | 1      | 0          | 2        |
| Claude Code    | 20%       | 2      | 0      | 7          | 1        |
| OpenCode       | 0%        | 0      | 0      | 10         | 0        |
| Qwen Code      | 0%        | 0      | 0      | 10         | 0        |

Read the API errors column first. The top three harnesses show what the model can do: 70 to 80% of tasks solved. The bottom three are held back by API errors. A whole row of API errors means that harness never had a working conversation with the endpoint, and fixing the endpoint is what moves it, not changing the model.

The report then groups API errors by their message, with an example trial for each, so 27 failed trials read as the three bugs they are. Open the example in `harbor view jobs` to see every step the harness took before the error.

## 4. Check the three most common endpoint bugs

These three caused every API error in our runs. Each one passes a simple prompt test and breaks a real harness.

**Tool definitions without `parameters` are rejected.** Qwen Code failed on its first request:

```text
API Error: 400 ... Invalid JSON data: Failed to deserialize the JSON body into the
target type: tools[2].function: missing field `parameters`
```

The OpenAI spec makes `function.parameters` optional, and some harnesses leave it out. Accept a tool without it and treat it as an empty object schema.

**Responses API input items don't parse.** OpenCode's first request went through and a later one failed:

```text
Invalid JSON data: Failed to deserialize the JSON body into the target type:
input: data did not match any variant of untagged enum ResponseInput
```

Later requests in a Responses session carry the conversation so far as `input` items, and the likely cause is the item types for the model's earlier tool calls and their outputs. The endpoint's schema has to accept every item type a client can send back, not only messages.

**Anthropic streams break mid-response.** Claude Code reported:

```text
API Error: The response stream was malformed. The response above may be incomplete.
```

Claude Code streams every response and expects the full Messages event sequence. Check that streams with tool use and thinking blocks arrive complete and in order. In our runs this bug appeared in 8 of 10 tasks with one model and 4 of 10 with another on the same endpoint, so test it with each model you serve.

## 5. Repeat it for every model

Some endpoint bugs depend on the model behind them, like the streaming one above, so run the matrix again whenever you add or update a model, and before you announce it. To compare models with each other, rather than to find bugs, give it more data: agent runs are noisy, so use `--attempts 3` and more tasks before reading much into a 10-point difference.

The [Harbor guide](https://tenki.cloud/docs/sandbox/harbor.md) covers installing Harbor with the Tenki environment, and [`tenki-harness-evals`](https://github.com/LuxorLabs/tenki-harness-evals) has every option for the matrix runs.