Decisions
Why Link Bench works the way it does.
Why is it one bench with two ways to measure?
One agent of a harness and many of it are the same agent, so they're measured the same way. Single and Many make, start, stop, read and message an agent with the same code, and write the same results, so a measure taken across 100 agents is the measure taken on one, under the same name. Single answers which harness to pick, and Many answers how many agents a machine runs. Neither is a second program to drift from the first.
Why is it written in Rust?
It started as two Python scripts, and became one Rust program: the rest of Link is Rust, so it reuses Link's own proof of an agent's endpoint and its help, instead of copies. Its tests run in CI with every folder's. The Rust measured what the Python did, the same way, before the Python was deleted.
Why is it a plugin?
Its results are meant to be checked: anyone should be able to run it
again at a new release and get the same answer within the noise the
tables show. So it ships with every release, at the core's version, and
lnk plugin add bench installs it like any other part of Link. It
measures with what a user has: lnk agent create, lnk measure, lnk sandbox log.
Why does Many measure Link Harness unless told otherwise?
Each agent installs its own harness. Link Harness takes a few megabytes
and a second for that, while the other three take hundreds of megabytes
and minutes, so a step of 100 of their agents takes hours and tens of
gigabytes. Every harness can still be measured, with --harness,
under the same rules, so a small new harness can be compared the day
it comes out.
Why a stand-in model?
A real model's time and failures would swamp what the harness adds, and cost money on every run. The stand-in answers at once, or after a set delay, in each provider's format, behind Link's own proxy as a hosted provider is reached. So a harness runs exactly as with a key, and the model's time is nobody's. Only Single's real group uses a real model, for what a stand-in can't show: whether a task gets done, and what it costs.
Why does it read with lnk measure first?
lnk measure is how a user sees what an agent costs, so the bench
uses it and so checks it. Where it doesn't see an agent, the bench
reads the same processes itself, and each value says which read it.
A set of looks is all of one or all of the other, so no value mixes
the two. While an agent answers, the bench looks itself, every few
milliseconds, as lnk measure walks every agent's folders, which
would load the machine and miss a turn of 10 ms.
Why a spec?
How often and for how long it measures decides both how long a run
takes and how far its numbers can be trusted. The spec says each in
one file, with the default the one published numbers come from, so a
run's choices are on record in its about.json and a quicker run is a
choice anyone can see.
