Harness Bench

The harness bench measures the harnesses your agent can run side by side, on one machine, in one run: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. This page is its methods: every number it writes, how it's measured, and what can bend it.

# the free groups: install, running, each message, security
link-core/bench/harnesses/run.sh --note "gcp e2-standard-2"
# one group, one harness: to fix a harness whose run failed
link-core/bench/harnesses/run.sh --groups running --harness openclaw
# the real model: spends API credits, shows an estimate and asks first
link-core/bench/harnesses/run.sh --groups real --spend --max-spend 20
# one run each and short waits, to try the script
link-core/bench/harnesses/run.sh --quick
# this checkout's programs, built, in place of the lnk installed here
link-core/bench/harnesses/run.sh --dev --quick
# the bench itself, on the mini harness, nothing spent
link-core/bench/harnesses/run.sh --self-test

Needs: Linux with bubblewrap, python3 and openssl, and lnk on the PATH (or this checkout and cargo, with --dev). Measuring what a harness sends a hosted provider needs root (Each message). The free groups take the machine about two hours for four harnesses, most of it installs and idle waits. Only the real group costs money, up to its cap. It's never run in CI.

What it writes

bench-<date>/, where it's run:

  • results.json: every metric of every harness under its key (the keys below, as install_time), each with its value, the number compared (n), its spread (least and most over its runs), how many runs, what measured it (by), or why it has none; and under runs, every run's raw values.
  • results.md: the same as tables, then where Link harness loses, then each task's runs.
  • about.json: the machine, lnk --version, each adapter's version, the settings of the run, the model, and each harness's message path.

A value is a median wherever there are several runs, and its spread says how far the runs went. A median without its spread can't be judged: two harnesses whose spreads overlap aren't told apart by the run.

What's held equal

  • One machine, one run, one release. Every harness is measured on the same machine, with the same lnk and so the same pinned harness versions, the same model and the same messages.
  • Interleaved. Repeated measurements go harness by harness (the first install of each, then the second), so a slow minute of the machine's or the network's falls on all of them.
  • One agent at a time. Only the agent being measured runs.
  • A home of its own. The run sets HOME to a folder of its own, with LNK_SERVICES=off and LNK_KEYCHAIN=off, and removes it at the end. Each agent is bench-<harness>, made with lnk agent new.
  • Agents run as processes of the run, as lnk agent start --foreground runs one, and are stopped as Ctrl-C stops it. A user's agent runs as a service: a service's start and stop add systemd's time, the same for every harness, which the bench leaves out.
  • lnk measure is the instrument wherever it has the measure: memory, CPU, processes, ports, network, disk and tokens. The bench times only what lnk measure doesn't have (install, start, stop, each message's times), and each value says which measured it.
  • A stand-in model (provider.py) answers every group but the real one at once, so the model's time is nobody's. It's a SOCKS5 upstream that answers TLS as api.anthropic.com, api.openai.com and openrouter.ai, with a CA made for the run: each agent's proxy has it as its upstream, so a harness is configured exactly as with a hosted key. It records what it's sent.

How a message reaches each harness

HarnessPath
Link harness, DeepSeek Harness, the mini harnessIts endpoint, as Link's channels send a message: proven, then streamed
Hermes AgentIts own OpenAI-compatible API, the thread as its session
OpenClawIts gateway's own OpenAI-compatible endpoint, turned on in that agent's config for the run

OpenClaw and Hermes speak their channels themselves, so neither has an endpoint of Link's. The bench turns on OpenClaw's endpoint (gateway.http.endpoints.chatCompletions.enabled) with OpenClaw's own openclaw config set. Where that doesn't answer, it sends each message with openclaw agent instead, and subtracts that command's start (openclaw --version, timed three times): a user messaging from a channel never pays it. Then both values are kept, and the cell says so. Each run's paths are in about.json.

What it measures

Install

Each harness is installed cold with lnk agent new: before each install the agent and every package cache in the run's home are removed. Link's own programs are there already, for every harness. Runs: 3 (--quick: 1).

  • Install time (install_time, s): from lnk agent new until "Running ... here", its last line before the harness starts: the adapter's install and configure. Timed by the bench from that command's own lines; results.json keeps install and configure apart. Noise: the network.
  • Download (downloaded, MB): bytes in and out of the machine's network during the install, from lnk measure's machine. Bias: anything else on the network counts; run it on a quiet machine.
  • On disk (on_disk, MB and files): the agent's folders after the install, before any message, from lnk measure. For Link harness, its one program, lnk-link-harness.
  • Packages (packages): distinct npm and Python packages, each package.json under node_modules and each .dist-info. The two aren't added up: they don't weigh the same.
  • Run code at install (install_scripts): npm packages with a preinstall, install or postinstall script.
  • Brings (runtimes): the runtimes in its home, with their versions.
  • Needs on the machine (needs_machine): not measured by the script. Install each harness in a bare debian:stable-slim with ca-certificates and bubblewrap, and list what it failed without.
  • Biggest parts (biggest_parts): its five largest packages or folders, a note for the reader.
  • Its download is checked (install_pinned): read from its install's lockfiles: "built by Link", "lockfile hashes", "lockfile without hashes", or "package hash; dependencies float".

Running

The stand-in answers at once.

  • First start (first_start, s): after the install, until the harness answers: its endpoint proves itself, or Hermes's /health or OpenClaw's gateway answers, looked at every 0.1 s. Includes any first-run setup. One for each install.
  • Restart (restart, s): stopped, then started, until it answers. Runs: 5 (quick: 1).
  • Stop (stop, s): from Ctrl-C until none of its processes remain. Runs: one for each restart.
  • Cold start (cold_start, s): a restart after the file cache is dropped (/proc/sys/vm/drop_caches). Linux, as root, on a machine with its own kernel; else its reason. Runs: 3.
  • First answer (first_answer, s): from a start until the first message is answered whole. Runs: 1.
  • Idle memory (idle_memory, MB): lnk measure's memory in use (PSS: pages shared between processes split between them), its median over 10 minutes looked at every 10 s, from 2 minutes after its first answer (quick: 20 s every 4 s, after 5 s). Spread: the looks.
  • Processes, ports (processes_ports): lnk measure's while idle.
  • Idle CPU (idle_cpu, core-seconds per hour): lnk measure's CPU over the idle minutes, per hour. One interval: no spread.
  • Idle traffic (idle_traffic, KB per hour): lnk measure's network through its proxy over the idle minutes, and the hosts it reached, from lnk sandbox log.
  • Added per message (added_per_message, ms, median / p95): the answer's time, whole, less the stand-in's time answering it, over 30 messages after 3 to warm up (quick: 5 after 1). Bias: the bench's own Python on the same machine, the same for every harness.
  • To model (to_model, ms): from the message sent until the stand-in is asked.
  • First token through (first_token_through, ms): from the stand-in's first piece of the answer until it reaches the bench. The stand-in streams three pieces, 200 ms apart. None for a harness that answers only whole the way it's reached.
  • CPU per message (cpu_per_message, core-ms): lnk measure's CPU over the messages, divided by their number.
  • Peak memory per message (peak_memory, MB): the most memory lnk measure saw, a look a second during the messages, less idle.
  • Disk per 100 messages (disk_per_100, KB): growth of the agent's folders over the messages, from lnk measure, per 100.
  • A provider error (provider_error): the stand-in answers 500 once, then hangs once (120 s; quick: 10 s). Whether the answer still comes, and when the user hears anything.

Each message, as the model sees it

What the stand-in recorded, once as Anthropic and once as OpenAI: a harness may send each a different prompt. One thread each, in a fresh files folder.

  • Prompt overhead (prompt_overhead, KB and tokens): the first request for "Reply with exactly: ok", in a new thread: everything but that sentence (its system prompt, its tools, any other messages). Tokens are about four bytes each, or with --count-tokens Anthropic's token count (free; needs ANTHROPIC_API_KEY). The real group's "Tokens for ok" is the provider's own count.
  • Tools offered (tools_offered): the tools in that request.
  • Extra model calls (extra_calls): requests in that turn beyond the one that answers it.
  • Growth per turn (growth_per_turn, tokens): what each turn adds to the prompt beyond its own words, by a straight line through 20 short turns (quick: 5).
  • Compacts at (compacts_at): the turn at which the prompt first shrinks, over 200 turns of 2 KB (quick: 8).
  • Prompt cache reuse (cache_friendly): "Yes" when everything before each turn's newest message is sent byte for byte as before, so a provider's cache can reuse it; else the byte where it first changed.

Link's proxy trusts only the machine's certificates, and starts with no environment that could add one. So the bench runs itself again in a mount namespace of its own, where the machine's CA file holds the run's CA too: only the run's programs trust it, and the machine's own trust is never changed. That takes root (unshare). Without it, --model-via local serves the stand-in as a model on the machine instead. That measures a local model's requests, which a harness may build differently from a hosted provider's, so the results say so.

With a real model

export ANTHROPIC_API_KEY=...
link-core/bench/harnesses/run.sh --groups real --spend --max-spend 20

The only group that spends. Every harness gets the same model, Claude Haiku 5.5 (--anthropic names another), and the same key, straight through Link's proxy to the provider, as a user's agent reaches it. It refuses to start without --spend and --max-spend, or for a model with no price in prices.json, which names where each price was read and when. It prints the price and an estimate from each harness's prompt size, then asks; --yes answers for a script. It stops at the cap. Each task runs 3 times (quick: 1), interleaved across the harnesses, each in a new thread with a fresh files folder.

TaskSentPasses when
T1 Hello"Reply with exactly: ok"The answer is ok
T2 RecallA code word, a sum, "What was my code word?"The last answer holds the code word
T3 Read your files"When does my trip start? It's in my files."The answer holds the date in notes/trip.md, among 50 notes
T4 Write a file"Save todo.txt in my files with the line: buy milk"The file holds the line
T5 Fetch"What's the title of https://example.com?"The answer holds "Example Domain"
T6 Count"How many rows does data.csv in my files have?"The answer holds 1,234
T7 Long answer"Write 300 words on how bees make honey."200 to 400 words

A harness whose tools can't do a task (by their names, as it offers them to the stand-in) shows "no tool" for it.

  • Tasks passed (tasks_passed, of 7): tasks passed in every run. Each task's runs are in results.md.
  • Tokens for "ok" (tokens_ok): T1's input tokens.
  • Tokens per suite (tokens_suite): input (and the cached part) and output over T1 to T7, the median run.
  • Spend for the 7 tasks (spend_suite, $): those tokens times the pinned price (prices.json, with its source and date). Input is the whole prompt, as lnk measure counts it, so the cache's reads and writes come out of it and are priced at their own rates: cache reads (link.tokens.cached.sum) at the cache-read price, cache writes (link.tokens.cache_written.sum) at the cache-write price, and the rest at the input price.
  • Model calls per task (calls_per_task): its proxy's calls to the model (lnk sandbox log).
  • First token, whole answer (first_token, whole_answer, s): from the message sent until the answer's first piece, and all of it, on T7. Mostly the provider's time.

Tokens are lnk measure's, read before and after each task while lnk measure watch looks every second. A harness Link has no tokens for shows none, and says why.

What it asks for, and reaches

  • Asks (asks_linux, asks_mac): what its adapter's needs names beyond Link's defaults, for each provider (Anthropic, OpenAI, OpenRouter, local) and each channel (none, Telegram, Slack, Discord, WhatsApp, Signal). An adapter answers for the OS it runs on, so the Mac's row comes from a run on a Mac.
  • Listening ports (ports): lnk measure's, while idle and the most during a turn.
  • Hosts it reaches (hosts): the hosts its proxy let it reach while it ran, besides its model's, from lnk sandbox log.
  • check (check): lnk agent <harness> check: "passes", or what failed.

results.md names every metric where Link harness isn't the best of the harnesses run: Link's value and the best value, each with its spread, which harness has it, and the gap. A task Link harness passed fewer times, "no tool" included, is a loss too. Losses are computed from the values, never picked, and published as measured. A loss "inside the noise" is one where Link's least is within the best's spread: the run doesn't tell them apart. Words (what it brings, how a fault went) and what's better neither way (tools offered, packages) aren't compared.

What isn't measured

  • Its answers' quality: that's the model's. The tasks only check that the harness gets the model's answer to the user, with the tool the task needs.
  • Throughput: that's the model's too. What a harness adds to streaming is "first token through".
  • Real channels (Slack, Telegram, ...): their service's noise would swamp the harness's.
  • Local models: OpenClaw takes only Ollama, and Hermes needs a 64k window. A fair comparison needs a GPU and a design of its own.

The self-test

run.sh --self-test runs every group against the mini harness and the stand-in, with the stand-in as a local model. The real group is rehearsed with the stand-in as its model, so nothing is spent. It checks the losses' arithmetic on made-up values for four harnesses, and that about.json records each harness's message path. It fails if any metric has neither a value nor a stated reason. Run it after changing the bench.