Single
One agent of each harness your agent can run, side by side on one machine: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. This page is its methods: every number it writes, how it's measured, and what can bend it.
lnk bench single --note "my box" # every group but the real one
lnk bench single --quick # one run each, short waits
# one group of one harness, to redo a harness whose run failed
lnk bench single --groups running --harness openclaw
# the real model's tasks: shows an estimate and asks first
lnk bench single --groups real --spend --max-spend 20It runs in a home of its own, ~/.lnk-bench, which it removes at the
end. It needs Linux with bubblewrap and openssl, and measuring what a
harness sends a hosted provider needs root (Each
message). The free groups take
the machine about an hour and a half for four harnesses, most of it
installs and idle waits. Only the real group costs money, up to its
cap.
What it writes
bench-single-<date>/, where it's run:
results.json: every metric of every harness under its key (the keys below, asinstall_time), each with itsvalue, the number compared (n), itsspread(least and most over its runs), how manyruns, what measured it (by), orwhyit has none; and underruns, every run's raw values.results.md: the same as tables, then where Link Harness loses, then each task's runs.about.json: the machine,lnk --version, each adapter's version, the spec and its settings, the model, and each harness's message path.
Below, each count and wait names its key in the spec's [single], and
the default's value.
What's held equal
- Interleaved, in turn. Repeated measurements go harness by
harness (the first install of each, then the second), so a slow
minute of the machine's or the network's falls on all of them, and
each round starts one harness later than the last (installs,
restarts and cold starts), so none always goes first onto a machine
still settling.
about.jsonkeeps the machine's load as each group began, and at the end. - One agent at a time, while anything is timed or idle: only the agent being measured runs (another's would share Link's own programs' pages, and lower the measured one's PSS). The message group, whose numbers are no time, and the real group's rounds run their harnesses at once where the spec says so.
- A home of its own. The run sets
HOMEto a folder of its own, withLNK_SERVICES=offandLNK_KEYCHAIN=off, and removes it at the end. Each agent isbench-<harness>, made withlnk agent create. - Agents run as processes of the run, as
lnk agent start --foregroundruns one, and are stopped as Ctrl-C stops it. A user's agent runs as a service: a service's start and stop add systemd's time, the same for every harness, which the bench leaves out. - The stand-in answers at once, in every group but the real one.
It's a SOCKS5 upstream that answers TLS as
api.anthropic.com,api.openai.comandopenrouter.ai, with a CA made for the run: each agent's proxy has it as its upstream, so a harness is configured exactly as with a hosted key.
How a message reaches each harness
| Harness | Path |
|---|---|
| Link Harness, DeepSeek Harness, the mini harness | Its endpoint, as Link's channels send a message: proven, then streamed |
| Hermes Agent | Its own OpenAI-compatible API, the thread as its session |
| OpenClaw | Its gateway's own OpenAI-compatible endpoint, turned on in that agent's config for the run |
OpenClaw and Hermes speak their channels themselves, so neither has an
endpoint of Link's. The bench turns on OpenClaw's endpoint
(gateway.http.endpoints.chatCompletions.enabled) with OpenClaw's own
openclaw config set. Where that doesn't answer, it sends each
message with openclaw agent instead, and subtracts that command's
start (openclaw --version, timed three times): a user messaging from
a channel never pays it. Then both values are kept, and the cell says
so. Each run's paths are in about.json.
What it measures
Install
Each harness is installed as a person installs it after lnk up
(which brings none of the four's plugins but Link Harness's): lnk plugin add of its plugin, then lnk agent create, cold. Before each
install the agent and every package cache in the run's home are
removed, and its plugin's program is set aside; after it, the program
the run was built with is put back, so every group runs that one. The
plugin added is the signed release of the run's lnk version, so a
build of a version with no release has no install time or download,
and says why. Runs: 2 (installs; quick: 1).
- Install time (
install_time, s):lnk plugin add, then fromlnk agent createuntil "Running ... here", its last line before the harness starts: the plugin's install, and the adapter's install and configure. Timed by the bench from those commands' own lines;results.jsonkeeps the three apart. Noise: the network. - Download (
downloaded, MB): bytes in and out of the machine's network during both, fromlnk measure's machine;results.jsonkeeps them apart. Bias: anything else on the network counts; run it on a quiet machine. - On disk (
on_disk, MB and files): the same for every harness: its plugin's program and its agent's folders after the install, before any message, in blocks on disk (lnk measure's, a small file taking a whole block), less the runtimes it brings (Node, Python, Bun, Deno), which many machines have already. A runtime is its own folder in the agent's home less what's installed into it (a harness in Node'slib/node_modules, as OpenClaw's is, stays the harness's). uv is a package manager, not a runtime: it stays the harness's. - On disk, a fresh machine (
on_disk_fresh, MB and files): the same, with the runtimes it brings. - Packages (
packages): distinct npm and Python packages, eachpackage.jsonundernode_modulesand each.dist-info. The two aren't added up: they don't weigh the same. - Run code at install (
install_scripts): npm packages with apreinstall,installorpostinstallscript. - Brings (
runtimes): the runtimes and uv in its home (abin/folder outsidenode_modulesandsite-packages), with their versions. - Needs on the machine (
needs_machine): not measured by the script. Install each harness in a baredebian:stable-slimwithca-certificatesandbubblewrap, and list what it failed without. - Biggest parts (
biggest_parts): its five largest packages, folders or its plugin's program, of 1 MB or more, a note for the reader. - Its download is checked (
install_pinned): read from its install's lockfiles: "built by Link", "lockfile hashes", "lockfile without hashes", or "package hash; dependencies float".
Running
The stand-in answers at once.
- First start (
first_start, s): after the install, until the harness answers: its endpoint proves itself and its/healthanswers (DeepSeek's adapter proves itself beforedshruns; its health waits fordsh), or Hermes's/healthor OpenClaw's gateway answers, looked at every 10 ms for its first 2 s, then every 0.1 s. Includes any first-run setup. One for each install. - Restart (
restart, s): stopped, then started, until it answers. Runs: 5 (quick: 1). - Stop (
stop, s): from Ctrl-C until none of its processes remain, looked at every 5 ms for its first 2 s. Runs: one for each restart. - Cold start (
cold_start, s): a restart after the file cache is dropped (/proc/sys/vm/drop_caches). Linux, as root, on a machine with its own kernel; else its reason. Runs: 3 (colds). - First answer (
first_answer, s): from a start until the first message is answered whole. Runs: 1. - Idle memory (
idle_memory, MB): its processes' memory in use (PSS: pages shared between processes split between them), its median over 5 minutes (idle) looked at every 10 s (look), from 2 minutes after its first answer (settle; quick: 20 s every 4 s, after 5 s). Spread: the looks. Every look islnk measure's where it saw the agent in all of them, else the bench's own reading of the same processes in all of them, never a mix, and the value says which. - Processes, ports (
processes_ports): its processes, and the TCP ports they listen on, while idle.results.jsonnames each process the bench finds of it: its program and first arguments. - Idle CPU (
idle_cpu, core-seconds per hour): its CPU from the first idle look to the last, per hour. One interval: no spread. - Idle traffic (
idle_traffic, KB per hour):lnk measure's network through its proxy over the idle minutes, and the hosts it reached, fromlnk sandbox log. - Added per message (
added_per_message, ms, median / p95): the answer's time, whole, less the stand-in's time answering it, over 30 messages after 3 to warm up (quick: 5 after 1). An answer must carry its message's marker, which the stand-in echoes: else it wasn't the model's, and the message failed. A message's requests of the stand-in are those carrying its marker (its tool calls' answers among them), each counted only within the message's own time, and time two of them overlap once: a request of the message before, still going, isn't its. Bias: the bench's own work on the same machine, the same for every harness. - To model (
to_model, ms): from the message sent until the stand-in is asked. Through OpenClaw's own command, the command's start is left out, as in "added". - First token through (
first_token_through, ms): from the stand-in's first piece of the answer until it reaches the bench. The stand-in streams three pieces, 200 ms apart. None for a harness that answers only whole the way it's reached. - CPU per message (
cpu_per_message, core-ms): its CPU over the messages, less what it would have used idle meanwhile (its idle rate, above), divided by their number.results.jsonkeeps it with the idle part too (with_idle). - Peak memory per message (
peak_memory, MB): for each message, the most resident memory of its processes (all under it, those started for a turn too) while it answered, less the same over half a second before the messages; the median of the messages, and the most inresults.json. These are the bench's own looks every 20 ms (watch), so a turn of 10 ms is seen, and nolnk measure(which walks every agent's folders) runs while it answers. The same looks give the ports it listened on while answering, each process's in its own network. - Disk per 100 messages (
disk_per_100, KB): growth of the agent's folders over the messages, in blocks on disk, walked by the bench, per 100. A few bytes more can take a whole block:results.jsonkeeps the growth in bytes (apparent_kb_per_100) and the blocks' resolution (resolution_kb_per_100, a block for each file that grew), so a difference smaller than that isn't one. - A provider error (
provider_error): the stand-in answers 500 once, then hangs once (120 s; quick: 10 s). Whether the answer still comes, and when the user hears anything.
Each message, as the model sees it
What the stand-in recorded, once as Anthropic and once as OpenAI: a
harness may send each a different prompt. One thread each, in a fresh
files folder. None of these numbers is a time, so the harnesses are
measured at once (the spec's message_at_once): each agent's fake key
is its own, of one length, and the stand-in tells whose a request is
by it. A local model's requests carry no key, so with --model-via local they go one at a time.
- Prompt overhead (
prompt_overhead, KB and tokens): the first request for "Reply with exactly: ok", in a new thread: everything but that sentence (its system prompt, its tools, any other messages). Tokens are the request's bytes as sent over four, or with--count-tokensAnthropic's token count (free; needsANTHROPIC_API_KEY). The real group's "Tokens for ok" is the provider's own count. - Tools offered (
tools_offered): the tools in that request. - Extra model calls (
extra_calls): requests in that turn beyond the one that answers it. - Growth per turn (
growth_per_turn, tokens): what each turn adds to the prompt beyond its own words, by a straight line through 20 short turns (quick: 5). - Compacts at (
compacts_at): the turn at which the prompt first shrinks, and how large the prompt was before it, over 50 turns of 8 KB (quick: 8). What sets compaction off is how much the history holds, which few large turns reach as many small ones would.results.jsonkeeps each turn's prompt size. A turn that never reaches the model ends it with no number, and says which turn and why. - Prompt cache reuse (
cache_friendly): "Yes" when everything before each turn's newest message is sent byte for byte as before, so a provider's cache can reuse it; else the byte where it first changed. Cache marks (Anthropic'scache_control) are left out of the comparison: they say where a cached part ends, and marking each turn's newest message, as a conversation is cached, moves the mark without changing anything cached.
Link's proxy trusts only the machine's certificates, and starts with
no environment that could add one. So the bench runs itself again in
a mount namespace of its own, where the machine's CA file holds the
run's CA too: only the run's programs trust it, and the machine's own
trust is never changed. That takes root (unshare). Without it,
--model-via local serves the stand-in as a model on the machine
instead. That measures a local model's requests, which a harness may
build differently from a hosted provider's, so the results say so.
OpenClaw takes a local model only as Ollama, which Link knows at
Ollama's own port: with OpenClaw, the stand-in is at 11434 and speaks
Ollama's API too, so an Ollama running there must be stopped first.
With a real model
export ANTHROPIC_API_KEY=...
lnk bench single --groups real --spend --max-spend 20The only group that spends. Every harness gets the same model, Claude
Haiku 5.5 (--anthropic names another), and the same key, straight
through Link's proxy to the provider, as a user's agent reaches it. It
refuses to start without --spend and --max-spend, or for a model
with no price in prices.json, which names where each price was read
and when. It prints the price and an estimate from each harness's
prompt size, then asks; --yes answers for a script. It stops at the
cap. Each task runs 3 times (real_runs; quick: 1), interleaved
across the harnesses, each in a new thread with a fresh files folder.
A round's harnesses run at once (real_at_once): the outcomes, tokens
and spend don't move by it, while first token and whole answer carry
the other harnesses' work on the machine. The cap is looked at after
each round.
| Task | Sent | Passes when |
|---|---|---|
| T1 Hello | "Reply with exactly: ok" | The answer is ok |
| T2 Recall | A code word, a sum, "What was my code word?" | The last answer holds the code word |
| T3 Read your files | "When does my trip start? It's in my files." | The answer holds the date in notes/trip.md, among 50 notes |
| T4 Write a file | "Save todo.txt in my files with the line: buy milk" | The file holds the line |
| T5 Fetch | "What's the title of https://example.com?" | The answer holds "Example Domain" |
| T6 Count | "How many rows does data.csv in my files have?" | The answer holds 1,234 |
| T7 Long answer | "Write 300 words on how bees make honey." | 200 to 400 words |
A harness whose tools can't do a task (by their names, as it offers them to the stand-in) shows "no tool" for it.
- Tasks passed (
tasks_passed, of 7): tasks passed in every run. Each task's runs are inresults.md. - Tokens for "ok" (
tokens_ok): T1's input tokens. - Tokens per suite (
tokens_suite): input (and the cached part) and output over T1 to T7, the median run. - Spend for the 7 tasks (
spend_suite, $): those tokens times the pinned price (prices.json, with its source and date). Input is the whole prompt, aslnk measurecounts it, so the cache's reads and writes come out of it and are priced at their own rates: cache reads (link.tokens.cached.sum) at the cache-read price, cache writes (link.tokens.cache_written.sum) at the cache-write price, and the rest at the input price. - Model calls per task (
calls_per_task): its proxy's calls to the model (lnk sandbox log), counted before and after each task. - First token, whole answer (
first_token,whole_answer, s): from the message sent until the answer's first piece, and all of it, on T7. Mostly the provider's time.
Tokens are lnk measure's, read before and after each task while
lnk measure watch looks every second. A harness Link has no tokens
for shows none, and says why.
What it asks for, and reaches
- Asks (
asks_linux,asks_mac): what its adapter'sneedsnames beyond Link's defaults, for each provider (Anthropic, OpenAI, OpenRouter, local) and each channel (none, Telegram, Slack, Discord, WhatsApp, Signal). An adapter answers for the OS it runs on, so the Mac's row comes from a run on a Mac. - Listening ports (
ports): while idle, and the most during a turn. - Hosts it reaches (
hosts): the hosts its proxy let it reach while it ran, besides its model's, fromlnk sandbox log. check(check):lnk agent harness check: "passes", or what failed.
Where Link Harness loses
results.md names every metric where Link Harness isn't the best of
the harnesses run: Link's value and the best value, each with its
spread, which harness has it, and the gap. A task Link Harness passed
fewer times, "no tool" included, is a loss too. Losses are computed
from the values, never picked, and published as measured. A loss
"inside the noise" is one where Link's least is within the best's
spread: the run doesn't tell them apart. Words (what it brings, how a
fault went) and what's better neither way (tools offered, packages)
aren't compared.
What isn't measured
- Its answers' quality: that's the model's. The tasks only check that the harness gets the model's answer to the user, with the tool the task needs.
- Throughput: that's the model's too. What a harness adds to streaming is "first token through".
- Real channels (Slack, Telegram, ...): their service's noise would swamp the harness's.
- Local models: OpenClaw takes only Ollama, and Hermes needs a 64k window. A fair comparison needs a GPU and a design of its own.
The self-test
lnk bench single --self-test runs every group against the mini
harness and the stand-in, with the stand-in as a local model, and the
real group rehearsed with the stand-in as its model, so nothing is
spent. It checks the losses' arithmetic on made-up values for four
harnesses, and that about.json records each harness's message path.
It fails if any metric has neither a value nor a stated reason:
Run it from a checkout.
