Single

One agent of each harness your agent can run, side by side on one machine: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. This page is its methods: every number it writes, how it's measured, and what can bend it.

lnk bench single --note "my box"   # every group but the real one
lnk bench single --quick           # one run each, short waits
# one group of one harness, to redo a harness whose run failed
lnk bench single --groups running --harness openclaw
# the real model's tasks: shows an estimate and asks first
lnk bench single --groups real --spend --max-spend 20

It runs in a home of its own, ~/.lnk-bench, which it removes at the end. It needs Linux with bubblewrap and openssl, and measuring what a harness sends a hosted provider needs root (Each message). The free groups take the machine about an hour and a half for four harnesses, most of it installs and idle waits. Only the real group costs money, up to its cap.

What it writes

bench-single-<date>/, where it's run:

  • results.json: every metric of every harness under its key (the keys below, as install_time), each with its value, the number compared (n), its spread (least and most over its runs), how many runs, what measured it (by), or why it has none; and under runs, every run's raw values.
  • results.md: the same as tables, then where Link Harness loses, then each task's runs.
  • about.json: the machine, lnk --version, each adapter's version, the spec and its settings, the model, and each harness's message path.

Below, each count and wait names its key in the spec's [single], and the default's value.

What's held equal

  • Interleaved, in turn. Repeated measurements go harness by harness (the first install of each, then the second), so a slow minute of the machine's or the network's falls on all of them, and each round starts one harness later than the last (installs, restarts and cold starts), so none always goes first onto a machine still settling. about.json keeps the machine's load as each group began, and at the end.
  • One agent at a time, while anything is timed or idle: only the agent being measured runs (another's would share Link's own programs' pages, and lower the measured one's PSS). The message group, whose numbers are no time, and the real group's rounds run their harnesses at once where the spec says so.
  • A home of its own. The run sets HOME to a folder of its own, with LNK_SERVICES=off and LNK_KEYCHAIN=off, and removes it at the end. Each agent is bench-<harness>, made with lnk agent create.
  • Agents run as processes of the run, as lnk agent start --foreground runs one, and are stopped as Ctrl-C stops it. A user's agent runs as a service: a service's start and stop add systemd's time, the same for every harness, which the bench leaves out.
  • The stand-in answers at once, in every group but the real one. It's a SOCKS5 upstream that answers TLS as api.anthropic.com, api.openai.com and openrouter.ai, with a CA made for the run: each agent's proxy has it as its upstream, so a harness is configured exactly as with a hosted key.

How a message reaches each harness

HarnessPath
Link Harness, DeepSeek Harness, the mini harnessIts endpoint, as Link's channels send a message: proven, then streamed
Hermes AgentIts own OpenAI-compatible API, the thread as its session
OpenClawIts gateway's own OpenAI-compatible endpoint, turned on in that agent's config for the run

OpenClaw and Hermes speak their channels themselves, so neither has an endpoint of Link's. The bench turns on OpenClaw's endpoint (gateway.http.endpoints.chatCompletions.enabled) with OpenClaw's own openclaw config set. Where that doesn't answer, it sends each message with openclaw agent instead, and subtracts that command's start (openclaw --version, timed three times): a user messaging from a channel never pays it. Then both values are kept, and the cell says so. Each run's paths are in about.json.

What it measures

Install

Each harness is installed as a person installs it after lnk up (which brings none of the four's plugins but Link Harness's): lnk plugin add of its plugin, then lnk agent create, cold. Before each install the agent and every package cache in the run's home are removed, and its plugin's program is set aside; after it, the program the run was built with is put back, so every group runs that one. The plugin added is the signed release of the run's lnk version, so a build of a version with no release has no install time or download, and says why. Runs: 2 (installs; quick: 1).

  • Install time (install_time, s): lnk plugin add, then from lnk agent create until "Running ... here", its last line before the harness starts: the plugin's install, and the adapter's install and configure. Timed by the bench from those commands' own lines; results.json keeps the three apart. Noise: the network.
  • Download (downloaded, MB): bytes in and out of the machine's network during both, from lnk measure's machine; results.json keeps them apart. Bias: anything else on the network counts; run it on a quiet machine.
  • On disk (on_disk, MB and files): the same for every harness: its plugin's program and its agent's folders after the install, before any message, in blocks on disk (lnk measure's, a small file taking a whole block), less the runtimes it brings (Node, Python, Bun, Deno), which many machines have already. A runtime is its own folder in the agent's home less what's installed into it (a harness in Node's lib/node_modules, as OpenClaw's is, stays the harness's). uv is a package manager, not a runtime: it stays the harness's.
  • On disk, a fresh machine (on_disk_fresh, MB and files): the same, with the runtimes it brings.
  • Packages (packages): distinct npm and Python packages, each package.json under node_modules and each .dist-info. The two aren't added up: they don't weigh the same.
  • Run code at install (install_scripts): npm packages with a preinstall, install or postinstall script.
  • Brings (runtimes): the runtimes and uv in its home (a bin/ folder outside node_modules and site-packages), with their versions.
  • Needs on the machine (needs_machine): not measured by the script. Install each harness in a bare debian:stable-slim with ca-certificates and bubblewrap, and list what it failed without.
  • Biggest parts (biggest_parts): its five largest packages, folders or its plugin's program, of 1 MB or more, a note for the reader.
  • Its download is checked (install_pinned): read from its install's lockfiles: "built by Link", "lockfile hashes", "lockfile without hashes", or "package hash; dependencies float".

Running

The stand-in answers at once.

  • First start (first_start, s): after the install, until the harness answers: its endpoint proves itself and its /health answers (DeepSeek's adapter proves itself before dsh runs; its health waits for dsh), or Hermes's /health or OpenClaw's gateway answers, looked at every 10 ms for its first 2 s, then every 0.1 s. Includes any first-run setup. One for each install.
  • Restart (restart, s): stopped, then started, until it answers. Runs: 5 (quick: 1).
  • Stop (stop, s): from Ctrl-C until none of its processes remain, looked at every 5 ms for its first 2 s. Runs: one for each restart.
  • Cold start (cold_start, s): a restart after the file cache is dropped (/proc/sys/vm/drop_caches). Linux, as root, on a machine with its own kernel; else its reason. Runs: 3 (colds).
  • First answer (first_answer, s): from a start until the first message is answered whole. Runs: 1.
  • Idle memory (idle_memory, MB): its processes' memory in use (PSS: pages shared between processes split between them), its median over 5 minutes (idle) looked at every 10 s (look), from 2 minutes after its first answer (settle; quick: 20 s every 4 s, after 5 s). Spread: the looks. Every look is lnk measure's where it saw the agent in all of them, else the bench's own reading of the same processes in all of them, never a mix, and the value says which.
  • Processes, ports (processes_ports): its processes, and the TCP ports they listen on, while idle. results.json names each process the bench finds of it: its program and first arguments.
  • Idle CPU (idle_cpu, core-seconds per hour): its CPU from the first idle look to the last, per hour. One interval: no spread.
  • Idle traffic (idle_traffic, KB per hour): lnk measure's network through its proxy over the idle minutes, and the hosts it reached, from lnk sandbox log.
  • Added per message (added_per_message, ms, median / p95): the answer's time, whole, less the stand-in's time answering it, over 30 messages after 3 to warm up (quick: 5 after 1). An answer must carry its message's marker, which the stand-in echoes: else it wasn't the model's, and the message failed. A message's requests of the stand-in are those carrying its marker (its tool calls' answers among them), each counted only within the message's own time, and time two of them overlap once: a request of the message before, still going, isn't its. Bias: the bench's own work on the same machine, the same for every harness.
  • To model (to_model, ms): from the message sent until the stand-in is asked. Through OpenClaw's own command, the command's start is left out, as in "added".
  • First token through (first_token_through, ms): from the stand-in's first piece of the answer until it reaches the bench. The stand-in streams three pieces, 200 ms apart. None for a harness that answers only whole the way it's reached.
  • CPU per message (cpu_per_message, core-ms): its CPU over the messages, less what it would have used idle meanwhile (its idle rate, above), divided by their number. results.json keeps it with the idle part too (with_idle).
  • Peak memory per message (peak_memory, MB): for each message, the most resident memory of its processes (all under it, those started for a turn too) while it answered, less the same over half a second before the messages; the median of the messages, and the most in results.json. These are the bench's own looks every 20 ms (watch), so a turn of 10 ms is seen, and no lnk measure (which walks every agent's folders) runs while it answers. The same looks give the ports it listened on while answering, each process's in its own network.
  • Disk per 100 messages (disk_per_100, KB): growth of the agent's folders over the messages, in blocks on disk, walked by the bench, per 100. A few bytes more can take a whole block: results.json keeps the growth in bytes (apparent_kb_per_100) and the blocks' resolution (resolution_kb_per_100, a block for each file that grew), so a difference smaller than that isn't one.
  • A provider error (provider_error): the stand-in answers 500 once, then hangs once (120 s; quick: 10 s). Whether the answer still comes, and when the user hears anything.

Each message, as the model sees it

What the stand-in recorded, once as Anthropic and once as OpenAI: a harness may send each a different prompt. One thread each, in a fresh files folder. None of these numbers is a time, so the harnesses are measured at once (the spec's message_at_once): each agent's fake key is its own, of one length, and the stand-in tells whose a request is by it. A local model's requests carry no key, so with --model-via local they go one at a time.

  • Prompt overhead (prompt_overhead, KB and tokens): the first request for "Reply with exactly: ok", in a new thread: everything but that sentence (its system prompt, its tools, any other messages). Tokens are the request's bytes as sent over four, or with --count-tokens Anthropic's token count (free; needs ANTHROPIC_API_KEY). The real group's "Tokens for ok" is the provider's own count.
  • Tools offered (tools_offered): the tools in that request.
  • Extra model calls (extra_calls): requests in that turn beyond the one that answers it.
  • Growth per turn (growth_per_turn, tokens): what each turn adds to the prompt beyond its own words, by a straight line through 20 short turns (quick: 5).
  • Compacts at (compacts_at): the turn at which the prompt first shrinks, and how large the prompt was before it, over 50 turns of 8 KB (quick: 8). What sets compaction off is how much the history holds, which few large turns reach as many small ones would. results.json keeps each turn's prompt size. A turn that never reaches the model ends it with no number, and says which turn and why.
  • Prompt cache reuse (cache_friendly): "Yes" when everything before each turn's newest message is sent byte for byte as before, so a provider's cache can reuse it; else the byte where it first changed. Cache marks (Anthropic's cache_control) are left out of the comparison: they say where a cached part ends, and marking each turn's newest message, as a conversation is cached, moves the mark without changing anything cached.

Link's proxy trusts only the machine's certificates, and starts with no environment that could add one. So the bench runs itself again in a mount namespace of its own, where the machine's CA file holds the run's CA too: only the run's programs trust it, and the machine's own trust is never changed. That takes root (unshare). Without it, --model-via local serves the stand-in as a model on the machine instead. That measures a local model's requests, which a harness may build differently from a hosted provider's, so the results say so. OpenClaw takes a local model only as Ollama, which Link knows at Ollama's own port: with OpenClaw, the stand-in is at 11434 and speaks Ollama's API too, so an Ollama running there must be stopped first.

With a real model

export ANTHROPIC_API_KEY=...
lnk bench single --groups real --spend --max-spend 20

The only group that spends. Every harness gets the same model, Claude Haiku 5.5 (--anthropic names another), and the same key, straight through Link's proxy to the provider, as a user's agent reaches it. It refuses to start without --spend and --max-spend, or for a model with no price in prices.json, which names where each price was read and when. It prints the price and an estimate from each harness's prompt size, then asks; --yes answers for a script. It stops at the cap. Each task runs 3 times (real_runs; quick: 1), interleaved across the harnesses, each in a new thread with a fresh files folder. A round's harnesses run at once (real_at_once): the outcomes, tokens and spend don't move by it, while first token and whole answer carry the other harnesses' work on the machine. The cap is looked at after each round.

TaskSentPasses when
T1 Hello"Reply with exactly: ok"The answer is ok
T2 RecallA code word, a sum, "What was my code word?"The last answer holds the code word
T3 Read your files"When does my trip start? It's in my files."The answer holds the date in notes/trip.md, among 50 notes
T4 Write a file"Save todo.txt in my files with the line: buy milk"The file holds the line
T5 Fetch"What's the title of https://example.com?"The answer holds "Example Domain"
T6 Count"How many rows does data.csv in my files have?"The answer holds 1,234
T7 Long answer"Write 300 words on how bees make honey."200 to 400 words

A harness whose tools can't do a task (by their names, as it offers them to the stand-in) shows "no tool" for it.

  • Tasks passed (tasks_passed, of 7): tasks passed in every run. Each task's runs are in results.md.
  • Tokens for "ok" (tokens_ok): T1's input tokens.
  • Tokens per suite (tokens_suite): input (and the cached part) and output over T1 to T7, the median run.
  • Spend for the 7 tasks (spend_suite, $): those tokens times the pinned price (prices.json, with its source and date). Input is the whole prompt, as lnk measure counts it, so the cache's reads and writes come out of it and are priced at their own rates: cache reads (link.tokens.cached.sum) at the cache-read price, cache writes (link.tokens.cache_written.sum) at the cache-write price, and the rest at the input price.
  • Model calls per task (calls_per_task): its proxy's calls to the model (lnk sandbox log), counted before and after each task.
  • First token, whole answer (first_token, whole_answer, s): from the message sent until the answer's first piece, and all of it, on T7. Mostly the provider's time.

Tokens are lnk measure's, read before and after each task while lnk measure watch looks every second. A harness Link has no tokens for shows none, and says why.

What it asks for, and reaches

  • Asks (asks_linux, asks_mac): what its adapter's needs names beyond Link's defaults, for each provider (Anthropic, OpenAI, OpenRouter, local) and each channel (none, Telegram, Slack, Discord, WhatsApp, Signal). An adapter answers for the OS it runs on, so the Mac's row comes from a run on a Mac.
  • Listening ports (ports): while idle, and the most during a turn.
  • Hosts it reaches (hosts): the hosts its proxy let it reach while it ran, besides its model's, from lnk sandbox log.
  • check (check): lnk agent harness check: "passes", or what failed.

results.md names every metric where Link Harness isn't the best of the harnesses run: Link's value and the best value, each with its spread, which harness has it, and the gap. A task Link Harness passed fewer times, "no tool" included, is a loss too. Losses are computed from the values, never picked, and published as measured. A loss "inside the noise" is one where Link's least is within the best's spread: the run doesn't tell them apart. Words (what it brings, how a fault went) and what's better neither way (tools offered, packages) aren't compared.

What isn't measured

  • Its answers' quality: that's the model's. The tasks only check that the harness gets the model's answer to the user, with the tool the task needs.
  • Throughput: that's the model's too. What a harness adds to streaming is "first token through".
  • Real channels (Slack, Telegram, ...): their service's noise would swamp the harness's.
  • Local models: OpenClaw takes only Ollama, and Hermes needs a 64k window. A fair comparison needs a GPU and a design of its own.

The self-test

lnk bench single --self-test runs every group against the mini harness and the stand-in, with the stand-in as a local model, and the real group rehearsed with the stand-in as its model, so nothing is spent. It checks the losses' arithmetic on made-up values for four harnesses, and that about.json records each harness's message path. It fails if any metric has neither a value nor a stated reason: Run it from a checkout.