# Harbor (https://tenki.cloud/docs/sandbox/harbor)

> For the complete documentation index, see [llms.txt](https://tenki.cloud/llms.txt)

Run Harbor benchmarks like Terminal-Bench with every trial in its own Tenki Sandbox VM, to evaluate coding agents and models in parallel.

[Harbor](https://github.com/harbor-framework/harbor) is an open-source framework for evaluating coding agents and models. It runs an agent (Claude Code, Codex, OpenCode, Goose, mini-swe-agent and others) against the tasks of a benchmark such as Terminal-Bench, grades every attempt with the task's own tests, and records the full trajectory. [`tenki-harbor`](https://github.com/LuxorLabs/tenki-harbor) adds Tenki as a place to run those trials: each one gets its own sandbox session, and you can run dozens in parallel.

## The short version

```bash
uv tool install harbor \
  --with "tenki-harbor @ git+https://github.com/LuxorLabs/tenki-harbor" \
  --with-executables-from tenki-harbor
export TENKI_API_KEY=tk_your_api_key

harbor run -t hello-world/hello-world \
  -e tenki_harbor:TenkiEnvironment --agent oracle
```

The last command runs Harbor's hello-world task with its reference solution, so no model is involved. It creates a sandbox, runs the task, grades it, prints a reward of `1.0`, and terminates the sandbox, in about 35 seconds. The rest of this page covers running real agents, faster startup, and what is supported.

## How it fits together

Harbor owns the evaluation: which tasks, which agent, how many attempts, the grading, and the results on disk. Tenki owns the machine. For each trial, `tenki-harbor`:

1. Creates a sandbox session sized to the task's CPU, memory and storage request.
2. Installs and starts Docker in the VM.
3. Pulls the task's prebuilt image, or builds its Dockerfile in the VM.
4. Routes every command and file transfer from Harbor into that container.
5. Terminates the session when the trial ends.

Harbor tasks are defined as container images, and a Tenki session is a full Linux VM, so the task container runs under Docker inside the VM. Agents and tests see exactly the environment the task's author built.

## Before you start

You need [uv](https://docs.astral.sh/uv/) (Harbor requires Python 3.12 or later), a Tenki API key exported as `TENKI_API_KEY` (see [Manage API keys](https://tenki.cloud/docs/account/api-keys.md)), and an API key for the model your agent will use.

## 1. Install

```bash
uv tool install harbor \
  --with "tenki-harbor @ git+https://github.com/LuxorLabs/tenki-harbor" \
  --with-executables-from tenki-harbor
```

This installs the `harbor` CLI with the Tenki environment in the same tool environment, plus the `tenki-harbor` command used for setup and cleanup below. To use it from a Python project instead, add both packages: `uv add harbor "tenki-harbor @ git+https://github.com/LuxorLabs/tenki-harbor"`.

## 2. Check the setup

```bash
export TENKI_API_KEY=tk_your_api_key
harbor run -t hello-world/hello-world \
  -e tenki_harbor:TenkiEnvironment --agent oracle
```

`-e tenki_harbor:TenkiEnvironment` is how Harbor loads an environment from an installed package. A reward of `1.0` means the sandbox, Docker, file transfer and grading all work.

## 3. Run a benchmark with an agent

Point any Harbor agent at a dataset and add `-e tenki_harbor:TenkiEnvironment`:

```bash
export ANTHROPIC_API_KEY=sk-ant-...
harbor run -d terminal-bench@2.0 -e tenki_harbor:TenkiEnvironment \
  --agent claude-code --model anthropic/claude-opus-4-1 \
  --n-concurrent 16 -l 10
```

That runs Claude Code on the first 10 Terminal-Bench 2.0 tasks, one sandbox per task. `--n-concurrent` caps how many run at once, and `-l` caps how many tasks run; drop `-l` to run all 89. Pass any other variables the agent needs with `--ae KEY=value`. Results land in `jobs/`; browse every step of every trial with `harbor view jobs`.

The same settings in a job config, for runs you repeat:

```yaml title="job.yaml"
n_concurrent_trials: 16
environment:
  import_path: tenki_harbor:TenkiEnvironment
  delete: true
agents:
  - name: claude-code
    model_name: anthropic/claude-opus-4-1
datasets:
  - name: terminal-bench
    version: "2.0"
    n_tasks: 10 # remove to run the full dataset
```

```bash
harbor run -c job.yaml
```

## 4. Start trials faster from a snapshot

Every trial installs Docker in a fresh VM, which takes about 12 seconds. Do it once instead:

```bash
tenki-harbor prepare
export TENKI_HARBOR_SNAPSHOT_ID=<printed snapshot id>
```

`prepare` creates a small sandbox, installs Docker, saves it as a [snapshot](https://tenki.cloud/docs/sandbox/snapshots.md), and terminates the sandbox. Trials then start from the snapshot with Docker already running. On Terminal-Bench, median environment setup went from 28 to 19 seconds. A snapshot belongs to the workspace it was taken in, so run `prepare` once per workspace.

## Options

Pass options with `--ek key=value` on the command line, or under `environment.kwargs` in a job config.

| Option             | Default                                            | Description                                                                                                  |
| ------------------ | -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `max_duration_sec` | `7200`                                             | Hard lifetime of each sandbox, so a crashed run cannot leave one running. Must cover build, agent and tests. |
| `disk_size_gb`     | Task storage plus 30 GB, at least 50               | Root disk of the VM, which also holds the image layers.                                                      |
| `snapshot_id`      | `TENKI_HARBOR_SNAPSHOT_ID`                         | Start sandboxes from this snapshot, usually the one from `tenki-harbor prepare`.                             |
| `image`            | The base image                                     | Start sandboxes from a Tenki image instead. Pass `image` or `snapshot_id`, not both.                         |
| `base_url`         | `TENKI_API_ENDPOINT`, or `https://api.tenki.cloud` | Tenki API endpoint.                                                                                          |

## What is supported

| Feature                                                                   | Status                                    |
| ------------------------------------------------------------------------- | ----------------------------------------- |
| Tasks with a prebuilt `docker_image`, including all of Terminal-Bench 2.0 | Supported                                 |
| Tasks with a Dockerfile, built in the VM                                  | Supported                                 |
| `network_mode = "no-network"`, including switching between phases         | Supported                                 |
| Tasks with a `docker-compose.yaml`                                        | Not yet, rejected with an error           |
| Network allowlists                                                        | Not yet, rejected with an error           |
| GPUs, Windows containers                                                  | Not supported                             |
| Up to 16 vCPU, 64 GB of memory and a 100 GB disk per trial                | Larger requests are rejected, not reduced |

We ran every Terminal-Bench 2.0 task with Harbor's reference solutions on Tenki, 16 at a time: 79 of 89 passed. Nine of the other ten fail because of the tasks themselves, such as package versions that are no longer published, and one hangs intermittently in its own test suite. The [`tenki-harbor` README](https://github.com/LuxorLabs/tenki-harbor) lists each one.

## Compare harnesses on one model

[`tenki-harness-evals`](https://github.com/LuxorLabs/tenki-harness-evals) builds on this to answer a different question: does a model work with every coding agent your users run? It runs one model across six harnesses on Terminal-Bench and reports, per harness, which trials passed, which failed the task, and which hit an API error that comes from the model endpoint rather than the model.

```bash
git clone https://github.com/LuxorLabs/tenki-harness-evals.git
cd tenki-harness-evals && uv sync
uv run evals run --model your-org/your-model
```

## Cost and cleanup

* Each trial uses one sandbox for as long as the trial runs, sized to the task, and terminates it at the end, whether the trial passed or failed.
* If Harbor is killed mid-run, its sandboxes still end on their own after `max_duration_sec`.
* Every sandbox is tagged `harbor` and carries the Harbor trial name in its metadata. To see and end any that are still running:

```bash
tenki-harbor sessions
tenki-harbor cleanup
```

## Troubleshooting

**`Tenki requires TENKI_API_KEY`.** Export `TENKI_API_KEY` in the shell that runs `harbor`. The CLI's saved login is not used.

**`docker pull ... no space left on device`.** The task's image is larger than the disk. Pass `--ek disk_size_gb=80`.

**`Tenki does not support docker-compose tasks yet`.** The task defines several services with Docker Compose, which is not supported yet. Run it on another Harbor environment for now.

**A trial was cut off before its tests ran.** Its build, agent and test timeouts add up to more than `max_duration_sec`. Raise it, for example `--ek max_duration_sec=14400`. If your workspace limits session lifetime, Tenki warns at creation and uses the limit instead.

## What's next

* [Sessions](https://tenki.cloud/docs/sandbox/sessions.md) for everything a sandbox can do, and [Snapshots](https://tenki.cloud/docs/sandbox/snapshots.md) for how `prepare` works.
* [Harbor documentation](https://harborframework.com/docs) for agents, datasets, job configs and the results viewer.
* [`tenki-harbor`](https://github.com/LuxorLabs/tenki-harbor) and [`tenki-harness-evals`](https://github.com/LuxorLabs/tenki-harness-evals) on GitHub.