Results
Every file a run writes, and every field of them, for reading a run's numbers in a script or checking where one came from. The measures themselves are explained on Single and Many.
The folder
A run writes into bench-single-<date>-<time>/ or
bench-many-<date>-<time>/ where it's run (--out names another), and
rewrites its files whole as each group or step ends, so a run stopped
halfway leaves what it had measured.
| File | Holds |
|---|---|
results.md | The tables, to read |
results.json | Every value, with how it was measured |
about.json | The machine, the versions and the settings |
logs/ | Each agent's lnk output, with the time of each line |
stand-in/requests.jsonl | Each request the stand-in model was sent |
messages.jsonl | Every message Many sent, a line each |
Every value: results.json
Each harness's measures are under harnesses, each by its key, as the
Single and Many pages name them (idle_memory, added_per_message):
{"harnesses": {"link": {"idle_memory": {"value": 21.6, "n": 21.6, "spread": [21.5, 21.7], "runs": 6, "by": "lnk measure", "raw": []}}}}| Field | What |
|---|---|
value | The value, as the tables show it: a number, or words with their number in n |
n | The number compared across harnesses |
spread | The least and the most of its runs, where it has more than one |
runs | How many runs its value is the median of |
by | What measured it: lnk measure, the bench's own reading of /proc, a proxy's log |
why | In place of all of these, why it has no value |
raw | What it was worked out from: each look, each run's numbers |
A measure taken once in each provider's format (Single's message group)
has an entry for each, under anthropic and openai, or local. One
of Many's has an entry for each step, under its number of agents
("20"), and steps lists each harness's steps in order.
runs holds what each measure came from, by harness: each install
(setup), each message (messages), each real task's run (real), and
what its adapter's needs said for each provider and channel (needs).
Many's are each step's messages, under its number of agents. In
Single, a harness whose group failed lists why in its errors.
Single's real group adds prices, the pinned prices it charged by, and
real_spent, the dollars it spent.
The machine and the settings: about.json
| Field | What |
|---|---|
date | When it started |
lnk, adapters | lnk --version, and each harness's adapter's --version |
machine | Its name, OS and kernel, CPUs and their model, memory |
note | --note's line |
harnesses | The harnesses measured, in the tables' order |
spec, quick, settings | The spec it ran, whether it's a quick one, and every count and wait it used, flags included |
load | The machine's load averages (1, 5 and 15 minutes) as each group or harness began, and at the end |
agents run as | Processes of the run, or the machine's services |
Single adds groups, model_via, the real group's model, the order
it went in, and each harness's message paths: how the bench sent it a
message. Many adds its steps, and messages, the load in words.
Each message: messages.jsonl
Many writes a line for each message as each step ends, and Single keeps
its messages in results.json's runs. Each holds the agent, its marker,
when it was sent (sent), when its first piece came back
(first_piece) and its end (end), all in Unix seconds, the answer's
text or its error, whether it was the agent's first since it
started, the model's requests, when the first came (model_at) and
their time (model_time), and its processes' memory beyond idle at
most (peak), the most ports they listened on, and the memory servers
started under them (tool_servers).
What the model was sent: stand-in/requests.jsonl
A line for each request, with when it came (at), when its answer was
decided and ended (decided, end), the provider it played and the
path, the message's marker, and the agent whose key it carried. Single
keeps each request's body whole (body, and raw as sent), which its
message group reads. Many keeps no bodies, as thousands of turns would
fill the machine's memory the agents are measured by.
