Security
What Link Bench does to the machine it runs on, what it trusts, and what it leaves behind. Run it on a machine of its own where you can: it makes and deletes agents, and Single runs as root to see what a harness sends a hosted provider.
Your agents
- Single never touches your agents. It runs in a home of its own,
~/.lnk-bench(LNK_BENCH_ROOTnames another), with Link's services and Keychain off, and removes it at the end, all but the results. A second run refuses to start while one is going there. - Many runs in your home on a machine of its own, as your
agents would, its agents as your services. They're named
bench-many-<harness>-<n>(--prefix), it refuses to start if any agent here has that name, and it deletes them at the end unless--keep. With--programsit runs in a home of its own, as Single does. - A run killed outright leaves its agents running. The next run in
a home of its own stops every process whose program is in that home,
or whose
HOMEis, before it starts. In yours, Many names the agents it finds and stops.
The test certificate
Single sees what a harness sends Anthropic or OpenAI by playing them,
behind Link's proxy, with a certificate authority made for the run
(openssl, valid two days). Link's proxy trusts only the machine's
certificates, so the bench runs itself again in a mount namespace of
its own (unshare, as root) where the machine's certificate file has
the run's added. Only the run's own programs see it, and the
machine's trust is never changed. The keys are in the results folder's
pki/, removed at the end. A run killed outright leaves them until
they expire.
--model-via local needs none of this, and no root: the stand-in is a
model on the machine instead.
What it sends and spends
- The stand-in is local. It listens on 127.0.0.1 only, and every group's messages but the real one's go to it, never to a provider. Each agent's key for it is made up.
- The real group spends your money. It sends seven tasks to
Anthropic with your
ANTHROPIC_API_KEY, through Link's proxy as a user's agent reaches it. It never starts without--spendand--max-spend, shows its estimate and asks first, and stops at the cap. Its tasks' files are made up. --count-tokenssends each message group prompt it counts to Anthropic's token counting, with your key. Nothing else leaves the machine but each harness's own install and what it reaches, which the results list (hosts).
The machine
- Cold starts drop the machine's file cache (Linux, as root), so everything else on it reads from disk again for a while.
- Installs set a harness's plugin aside. Single moves
lnk-<plugin>on your PATH aside while it installs the released one, and puts it back after. A run killed between leaveslnk-<plugin>.bench-aside, which the next Single run puts back. - Many takes Ollama's port (11434) when it measures OpenClaw, which takes a local model only from there: an Ollama running there must be stopped first.
What the results hold
The results name the machine, its CPU and kernel, lnk's and each
adapter's versions, and your --note. Single's stand-in log keeps every
prompt each harness sent, the bench's own messages and each harness's
own instructions, never a person's. Read them before you publish them.
