Get Started
Measure Link's harnesses on your own machine: one agent of each side by side, or as many agents of one as the machine runs. Every number says how it was measured and how far its runs spread, so anyone can run it again and check.
You need Linux with bubblewrap, and lnk with its agents (lnk up agents). It's free unless you run the real model's tasks, which spend
API credits up to a cap you set.
Quickstart
lnk plugin add bench # Link Bench, at lnk's version
lnk bench single --quick # one of each harness, quickly
lnk bench single --note "my box" # the full run, about 90 minutes
lnk bench many # Link Harness, in growing stepsEach run writes its results where it's run: results.md to read,
results.json with every value, and about.json with the machine,
the versions and the settings. Results says what each
holds, and Security what a run does to the machine.
Two ways to measure
- Single runs one agent of each harness, side by side: what each costs to install, to keep running and to send a message, what it sends its model, and what it asks for. It's for picking a harness, and for finding where Link Harness loses.
- Many runs agents of one harness in steps of 1, 5, 20, 50, 100 and on, an hour each, until a message isn't answered in time. It's for sizing a machine for a team of agents.
They're one bench. An agent is made, started, stopped and sent a message the same way in both, and read the same way, so a measure of many agents is the same measure as one, under the same name. Idle memory at 100 agents is the same number as idle memory in Single, read across 100 of them.
How it measures
- One machine, one run, one release. Every harness runs on the same
machine with the same
lnk, so the same pinned harness versions, the same model and the same messages. - A stand-in model answers every message but the real tasks', in each provider's format, at once or after a set delay. So the model's time is nobody's, and nothing is spent. It records what each harness sends it.
lnk measurereads the agents wherever it sees them: memory, processes, ports, CPU and network, from outside. Where it doesn't, the bench reads the same processes itself, and each value says which read it. While an agent answers, the bench looks at its processes every few milliseconds itself, since a turn of 10 ms is gone beforelnk measurelooks.- A value is a median wherever there are several runs, and its spread is the least and the most of them. Two harnesses whose spreads overlap aren't told apart by the run.
- The method doesn't bend. A measure, a task or a threshold never changes after its results are seen just to move Link Harness up. A change of method runs every harness again.
How it's built
Link Bench is a plugin of Link's, lnk-bench, in its own folder of
Link's repository (link-bench/), and it reaches Link only as a user
does, through lnk. One set of parts does all of its work, and each
command only says what to measure and how often:
| Part | Does |
|---|---|
agent/ | Makes an agent with lnk agent create, starts and stops it, knows when it's ready, finds its processes, and sends it a message the way its harness takes one |
standin.rs | The stand-in model: each provider's format, behind Link's proxy or as a model on the machine, and a record of each request |
reading.rs | Reads an agent idle, from lnk measure or /proc, watches it while it answers, and walks its folders |
send.rs | Sends one message and measures it, and turns a set of messages and an idle look into the measures both commands write |
results.rs | Every measure's value, spread, runs and source, or why it has none, as results.json and results.md |
single/, many.rs | The two commands: Single's groups, a file each, and Many's steps |
Its tests run with every folder's in CI, and each command has a self-test on the mini harness (Run it from a checkout).
Change what it measures
How often and for how long it measures is its spec. The default is the
one published numbers come from, and --quick is a short one to try
the bench. Your own spec is a file holding only what it changes, and a
flag on the command line wins over both:
lnk bench single --spec my-spec.toml
lnk bench many --steps 1,5,20 --hold 600The default spec, link-bench/specs/default.toml in Link's repository,
says what each setting is. Each run's about.json keeps the spec it
ran.
Run it from a checkout
# lnk and its plugins, then lnk-bench, on your PATH
cargo build --release --manifest-path link-core/Cargo.toml
cargo install --path link-bench
lnk bench single --self-test --programs link-core/target/release
lnk bench many --self-test --programs link-core/target/release--programs runs a checkout's built programs in place of the lnk
installed here, in a home of the run's own. Name it once for each
folder, link-harness/target/release for Link Harness among them. Each
self-test runs its command on the mini harness against the stand-in
in a few minutes, and fails if any measure has neither a value nor a
reason. Run them after changing the bench.
Why the bench is built this way: Decisions.
