Decisions

Why Link Bench works the way it does.

Why is it one bench with two ways to measure?

One agent of a harness and many of it are the same agent, so they're measured the same way. Single and Many make, start, stop, read and message an agent with the same code, and write the same results, so a measure taken across 100 agents is the measure taken on one, under the same name. Single answers which harness to pick, and Many answers how many agents a machine runs. Neither is a second program to drift from the first.

Why is it written in Rust?

It started as two Python scripts, and became one Rust program: the rest of Link is Rust, so it reuses Link's own proof of an agent's endpoint and its help, instead of copies. Its tests run in CI with every folder's. The Rust measured what the Python did, the same way, before the Python was deleted.

Why is it a plugin?

Its results are meant to be checked: anyone should be able to run it again at a new release and get the same answer within the noise the tables show. So it ships with every release, at the core's version, and lnk plugin add bench installs it like any other part of Link. It measures with what a user has: lnk agent create, lnk measure, lnk sandbox log.

Each agent installs its own harness. Link Harness takes a few megabytes and a second for that, while the other three take hundreds of megabytes and minutes, so a step of 100 of their agents takes hours and tens of gigabytes. Every harness can still be measured, with --harness, under the same rules, so a small new harness can be compared the day it comes out.

Why a stand-in model?

A real model's time and failures would swamp what the harness adds, and cost money on every run. The stand-in answers at once, or after a set delay, in each provider's format, behind Link's own proxy as a hosted provider is reached. So a harness runs exactly as with a key, and the model's time is nobody's. Only Single's real group uses a real model, for what a stand-in can't show: whether a task gets done, and what it costs.

Why does it read with lnk measure first?

lnk measure is how a user sees what an agent costs, so the bench uses it and so checks it. Where it doesn't see an agent, the bench reads the same processes itself, and each value says which read it. A set of looks is all of one or all of the other, so no value mixes the two. While an agent answers, the bench looks itself, every few milliseconds, as lnk measure walks every agent's folders, which would load the machine and miss a turn of 10 ms.

Why a spec?

How often and for how long it measures decides both how long a run takes and how far its numbers can be trusted. The spec says each in one file, with the default the one published numbers come from, so a run's choices are on record in its about.json and a quicker run is a choice anyone can see.