Get Started

Measure Link's harnesses on your own machine: one agent of each side by side, or as many agents of one as the machine runs. Every number says how it was measured and how far its runs spread, so anyone can run it again and check.

You need Linux with bubblewrap, and lnk with its agents (lnk up agents). It's free unless you run the real model's tasks, which spend API credits up to a cap you set.

Quickstart

lnk plugin add bench              # Link Bench, at lnk's version
lnk bench single --quick          # one of each harness, quickly
lnk bench single --note "my box"  # the full run, about 90 minutes
lnk bench many                    # Link Harness, in growing steps

Each run writes its results where it's run: results.md to read, results.json with every value, and about.json with the machine, the versions and the settings. Results says what each holds, and Security what a run does to the machine.

Two ways to measure

  • Single runs one agent of each harness, side by side: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. It's for picking a harness, and for finding where Link Harness loses.
  • Many runs agents of one harness in steps of 1, 5, 20, 50, 100 and on, an hour each, until a message isn't answered in time. It's for sizing a machine for a team of agents.

They're one bench. An agent is made, started, stopped and sent a message the same way in both, and read the same way, so a measure of many agents is the same measure as one, under the same name. Idle memory at 100 agents is the same number as idle memory in Single, read across 100 of them.

How it measures

  • One machine, one run, one release. Every harness runs on the same machine with the same lnk, so the same pinned harness versions, the same model and the same messages.
  • A stand-in model answers every message but the real tasks', in each provider's format, at once or after a set delay. So the model's time is nobody's, and nothing is spent. It records what each harness sends it.
  • lnk measure reads the agents wherever it sees them: memory, processes, ports, CPU and network, from outside. Where it doesn't, the bench reads the same processes itself, and each value says which read it. While an agent answers, the bench looks at its processes every few milliseconds itself, since a turn of 10 ms is gone before lnk measure looks.
  • A value is a median wherever there are several runs, and its spread is the least and the most of them. Two harnesses whose spreads overlap aren't told apart by the run.
  • The method doesn't bend. A measure, a task or a threshold never changes after its results are seen just to move Link Harness up. A change of method runs every harness again.

How it's built

Link Bench is a plugin of Link's, lnk-bench, in its own folder of Link's repository (link-bench/), and it reaches Link only as a user does, through lnk. One set of parts does all of its work, and each command only says what to measure and how often:

PartDoes
agent/Makes an agent with lnk agent create, starts and stops it, knows when it's ready, finds its processes, and sends it a message the way its harness takes one
standin.rsThe stand-in model: each provider's format, behind Link's proxy or as a model on the machine, and a record of each request
reading.rsReads an agent idle, from lnk measure or /proc, watches it while it answers, and walks its folders
send.rsSends one message and measures it, and turns a set of messages and an idle look into the measures both commands write
results.rsEvery measure's value, spread, runs and source, or why it has none, as results.json and results.md
single/, many.rsThe two commands: Single's groups, a file each, and Many's steps

Its tests run with every folder's in CI, and each command has a self-test on the mini harness (Run it from a checkout).

Change what it measures

How often and for how long it measures is its spec. The default is the one published numbers come from, and --quick is a short one to try the bench. Your own spec is a file holding only what it changes, and a flag on the command line wins over both:

lnk bench single --spec my-spec.toml
lnk bench many --steps 1,5,20 --hold 600

The default spec, link-bench/specs/default.toml in Link's repository, says what each setting is. Each run's about.json keeps the spec it ran.

Run it from a checkout

# lnk and its plugins, then lnk-bench, on your PATH
cargo build --release --manifest-path link-core/Cargo.toml
cargo install --path link-bench
lnk bench single --self-test --programs link-core/target/release
lnk bench many --self-test --programs link-core/target/release

--programs runs a checkout's built programs in place of the lnk installed here, in a home of the run's own. Name it once for each folder, link-harness/target/release for Link Harness among them. Each self-test runs its command on the mini harness against the stand-in in a few minutes, and fails if any measure has neither a value nor a reason. Run them after changing the bench.

Why the bench is built this way: Decisions.