Harness Bench
The harness bench measures the harnesses your agent can run side by side, on one machine, in one run: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. This page is its methods: every number it writes, how it's measured, and what can bend it.
# the free groups: install, running, each message, security
link-core/bench/harnesses/run.sh --note "gcp e2-standard-2"
# one group, one harness: to fix a harness whose run failed
link-core/bench/harnesses/run.sh --groups running --harness openclaw
# the real model: spends API credits, shows an estimate and asks first
link-core/bench/harnesses/run.sh --groups real --spend --max-spend 20
# one run each and short waits, to try the script
link-core/bench/harnesses/run.sh --quick
# this checkout's programs, built, in place of the lnk installed here
link-core/bench/harnesses/run.sh --dev --quick
# the bench itself, on the mini harness, nothing spent
link-core/bench/harnesses/run.sh --self-testNeeds: Linux with bubblewrap, python3 and openssl, and lnk on the
PATH (or this checkout and cargo, with --dev). Measuring what a
harness sends a hosted provider needs root (Each
message). The free groups take
the machine about two hours for four harnesses, most of it installs and
idle waits. Only the real group costs money, up to its cap. It's never
run in CI.
What it writes
bench-<date>/, where it's run:
results.json: every metric of every harness under its key (the keys below, asinstall_time), each with itsvalue, the number compared (n), itsspread(least and most over its runs), how manyruns, what measured it (by), orwhyit has none; and underruns, every run's raw values.results.md: the same as tables, then where Link harness loses, then each task's runs.about.json: the machine,lnk --version, each adapter's version, the settings of the run, the model, and each harness's message path.
A value is a median wherever there are several runs, and its spread says how far the runs went. A median without its spread can't be judged: two harnesses whose spreads overlap aren't told apart by the run.
What's held equal
- One machine, one run, one release. Every harness is measured on
the same machine, with the same
lnkand so the same pinned harness versions, the same model and the same messages. - Interleaved. Repeated measurements go harness by harness (the first install of each, then the second), so a slow minute of the machine's or the network's falls on all of them.
- One agent at a time. Only the agent being measured runs.
- A home of its own. The run sets
HOMEto a folder of its own, withLNK_SERVICES=offandLNK_KEYCHAIN=off, and removes it at the end. Each agent isbench-<harness>, made withlnk agent new. - Agents run as processes of the run, as
lnk agent start --foregroundruns one, and are stopped as Ctrl-C stops it. A user's agent runs as a service: a service's start and stop add systemd's time, the same for every harness, which the bench leaves out. lnk measureis the instrument wherever it has the measure: memory, CPU, processes, ports, network, disk and tokens. The bench times only whatlnk measuredoesn't have (install, start, stop, each message's times), and each value says which measured it.- A stand-in model (
provider.py) answers every group but the real one at once, so the model's time is nobody's. It's a SOCKS5 upstream that answers TLS asapi.anthropic.com,api.openai.comandopenrouter.ai, with a CA made for the run: each agent's proxy has it as its upstream, so a harness is configured exactly as with a hosted key. It records what it's sent.
How a message reaches each harness
| Harness | Path |
|---|---|
| Link harness, DeepSeek Harness, the mini harness | Its endpoint, as Link's channels send a message: proven, then streamed |
| Hermes Agent | Its own OpenAI-compatible API, the thread as its session |
| OpenClaw | Its gateway's own OpenAI-compatible endpoint, turned on in that agent's config for the run |
OpenClaw and Hermes speak their channels themselves, so neither has an
endpoint of Link's. The bench turns on OpenClaw's endpoint
(gateway.http.endpoints.chatCompletions.enabled) with OpenClaw's own
openclaw config set. Where that doesn't answer, it sends each
message with openclaw agent instead, and subtracts that command's
start (openclaw --version, timed three times): a user messaging from
a channel never pays it. Then both values are kept, and the cell says
so. Each run's paths are in about.json.
What it measures
Install
Each harness is installed cold with lnk agent new: before each
install the agent and every package cache in the run's home are
removed. Link's own programs are there already, for every harness.
Runs: 3 (--quick: 1).
- Install time (
install_time, s): fromlnk agent newuntil "Running ... here", its last line before the harness starts: the adapter's install and configure. Timed by the bench from that command's own lines;results.jsonkeeps install and configure apart. Noise: the network. - Download (
downloaded, MB): bytes in and out of the machine's network during the install, fromlnk measure's machine. Bias: anything else on the network counts; run it on a quiet machine. - On disk (
on_disk, MB and files): the agent's folders after the install, before any message, fromlnk measure. For Link harness, its one program,lnk-link-harness. - Packages (
packages): distinct npm and Python packages, eachpackage.jsonundernode_modulesand each.dist-info. The two aren't added up: they don't weigh the same. - Run code at install (
install_scripts): npm packages with apreinstall,installorpostinstallscript. - Brings (
runtimes): the runtimes in its home, with their versions. - Needs on the machine (
needs_machine): not measured by the script. Install each harness in a baredebian:stable-slimwithca-certificatesandbubblewrap, and list what it failed without. - Biggest parts (
biggest_parts): its five largest packages or folders, a note for the reader. - Its download is checked (
install_pinned): read from its install's lockfiles: "built by Link", "lockfile hashes", "lockfile without hashes", or "package hash; dependencies float".
Running
The stand-in answers at once.
- First start (
first_start, s): after the install, until the harness answers: its endpoint proves itself, or Hermes's/healthor OpenClaw's gateway answers, looked at every 0.1 s. Includes any first-run setup. One for each install. - Restart (
restart, s): stopped, then started, until it answers. Runs: 5 (quick: 1). - Stop (
stop, s): from Ctrl-C until none of its processes remain. Runs: one for each restart. - Cold start (
cold_start, s): a restart after the file cache is dropped (/proc/sys/vm/drop_caches). Linux, as root, on a machine with its own kernel; else its reason. Runs: 3. - First answer (
first_answer, s): from a start until the first message is answered whole. Runs: 1. - Idle memory (
idle_memory, MB):lnk measure's memory in use (PSS: pages shared between processes split between them), its median over 10 minutes looked at every 10 s, from 2 minutes after its first answer (quick: 20 s every 4 s, after 5 s). Spread: the looks. - Processes, ports (
processes_ports):lnk measure's while idle. - Idle CPU (
idle_cpu, core-seconds per hour):lnk measure's CPU over the idle minutes, per hour. One interval: no spread. - Idle traffic (
idle_traffic, KB per hour):lnk measure's network through its proxy over the idle minutes, and the hosts it reached, fromlnk sandbox log. - Added per message (
added_per_message, ms, median / p95): the answer's time, whole, less the stand-in's time answering it, over 30 messages after 3 to warm up (quick: 5 after 1). Bias: the bench's own Python on the same machine, the same for every harness. - To model (
to_model, ms): from the message sent until the stand-in is asked. - First token through (
first_token_through, ms): from the stand-in's first piece of the answer until it reaches the bench. The stand-in streams three pieces, 200 ms apart. None for a harness that answers only whole the way it's reached. - CPU per message (
cpu_per_message, core-ms):lnk measure's CPU over the messages, divided by their number. - Peak memory per message (
peak_memory, MB): the most memorylnk measuresaw, a look a second during the messages, less idle. - Disk per 100 messages (
disk_per_100, KB): growth of the agent's folders over the messages, fromlnk measure, per 100. - A provider error (
provider_error): the stand-in answers 500 once, then hangs once (120 s; quick: 10 s). Whether the answer still comes, and when the user hears anything.
Each message, as the model sees it
What the stand-in recorded, once as Anthropic and once as OpenAI: a harness may send each a different prompt. One thread each, in a fresh files folder.
- Prompt overhead (
prompt_overhead, KB and tokens): the first request for "Reply with exactly: ok", in a new thread: everything but that sentence (its system prompt, its tools, any other messages). Tokens are about four bytes each, or with--count-tokensAnthropic's token count (free; needsANTHROPIC_API_KEY). The real group's "Tokens for ok" is the provider's own count. - Tools offered (
tools_offered): the tools in that request. - Extra model calls (
extra_calls): requests in that turn beyond the one that answers it. - Growth per turn (
growth_per_turn, tokens): what each turn adds to the prompt beyond its own words, by a straight line through 20 short turns (quick: 5). - Compacts at (
compacts_at): the turn at which the prompt first shrinks, over 200 turns of 2 KB (quick: 8). - Prompt cache reuse (
cache_friendly): "Yes" when everything before each turn's newest message is sent byte for byte as before, so a provider's cache can reuse it; else the byte where it first changed.
Link's proxy trusts only the machine's certificates, and starts with
no environment that could add one. So the bench runs itself again in
a mount namespace of its own, where the machine's CA file holds the
run's CA too: only the run's programs trust it, and the machine's own
trust is never changed. That takes root (unshare). Without it,
--model-via local serves the stand-in as a model on the machine
instead. That measures a local model's requests, which a harness may
build differently from a hosted provider's, so the results say so.
With a real model
export ANTHROPIC_API_KEY=...
link-core/bench/harnesses/run.sh --groups real --spend --max-spend 20The only group that spends. Every harness gets the same model, Claude
Haiku 5.5 (--anthropic names another), and the same key, straight
through Link's proxy to the provider, as a user's agent reaches it. It
refuses to start without --spend and --max-spend, or for a model
with no price in prices.json, which names where each price was read
and when. It prints the price and an estimate from each harness's
prompt size, then asks; --yes answers for a script. It stops at the
cap. Each task runs 3 times (quick: 1), interleaved across the
harnesses, each in a new thread with a fresh files folder.
| Task | Sent | Passes when |
|---|---|---|
| T1 Hello | "Reply with exactly: ok" | The answer is ok |
| T2 Recall | A code word, a sum, "What was my code word?" | The last answer holds the code word |
| T3 Read your files | "When does my trip start? It's in my files." | The answer holds the date in notes/trip.md, among 50 notes |
| T4 Write a file | "Save todo.txt in my files with the line: buy milk" | The file holds the line |
| T5 Fetch | "What's the title of https://example.com?" | The answer holds "Example Domain" |
| T6 Count | "How many rows does data.csv in my files have?" | The answer holds 1,234 |
| T7 Long answer | "Write 300 words on how bees make honey." | 200 to 400 words |
A harness whose tools can't do a task (by their names, as it offers them to the stand-in) shows "no tool" for it.
- Tasks passed (
tasks_passed, of 7): tasks passed in every run. Each task's runs are inresults.md. - Tokens for "ok" (
tokens_ok): T1's input tokens. - Tokens per suite (
tokens_suite): input (and the cached part) and output over T1 to T7, the median run. - Spend for the 7 tasks (
spend_suite, $): those tokens times the pinned price (prices.json, with its source and date). Input is the whole prompt, aslnk measurecounts it, so the cache's reads and writes come out of it and are priced at their own rates: cache reads (link.tokens.cached.sum) at the cache-read price, cache writes (link.tokens.cache_written.sum) at the cache-write price, and the rest at the input price. - Model calls per task (
calls_per_task): its proxy's calls to the model (lnk sandbox log). - First token, whole answer (
first_token,whole_answer, s): from the message sent until the answer's first piece, and all of it, on T7. Mostly the provider's time.
Tokens are lnk measure's, read before and after each task while
lnk measure watch looks every second. A harness Link has no tokens
for shows none, and says why.
What it asks for, and reaches
- Asks (
asks_linux,asks_mac): what its adapter'sneedsnames beyond Link's defaults, for each provider (Anthropic, OpenAI, OpenRouter, local) and each channel (none, Telegram, Slack, Discord, WhatsApp, Signal). An adapter answers for the OS it runs on, so the Mac's row comes from a run on a Mac. - Listening ports (
ports):lnk measure's, while idle and the most during a turn. - Hosts it reaches (
hosts): the hosts its proxy let it reach while it ran, besides its model's, fromlnk sandbox log. check(check):lnk agent <harness> check: "passes", or what failed.
Where Link harness loses
results.md names every metric where Link harness isn't the best of
the harnesses run: Link's value and the best value, each with its
spread, which harness has it, and the gap. A task Link harness passed
fewer times, "no tool" included, is a loss too. Losses are computed
from the values, never picked, and published as measured. A loss
"inside the noise" is one where Link's least is within the best's
spread: the run doesn't tell them apart. Words (what it brings, how a
fault went) and what's better neither way (tools offered, packages)
aren't compared.
What isn't measured
- Its answers' quality: that's the model's. The tasks only check that the harness gets the model's answer to the user, with the tool the task needs.
- Throughput: that's the model's too. What a harness adds to streaming is "first token through".
- Real channels (Slack, Telegram, ...): their service's noise would swamp the harness's.
- Local models: OpenClaw takes only Ollama, and Hermes needs a 64k window. A fair comparison needs a GPU and a design of its own.
The self-test
run.sh --self-test runs every group against the mini harness and
the stand-in, with the stand-in as a local model. The real group is
rehearsed with the stand-in as its model, so nothing is spent. It
checks the losses' arithmetic on made-up values for four harnesses,
and that about.json records each harness's message path. It fails if
any metric has neither a value nor a stated reason. Run it after
changing the bench.
