Results

Every file a run writes, and every field of them, for reading a run's numbers in a script or checking where one came from. The measures themselves are explained on Single and Many.

The folder

A run writes into bench-single-<date>-<time>/ or bench-many-<date>-<time>/ where it's run (--out names another), and rewrites its files whole as each group or step ends, so a run stopped halfway leaves what it had measured.

FileHolds
results.mdThe tables, to read
results.jsonEvery value, with how it was measured
about.jsonThe machine, the versions and the settings
logs/Each agent's lnk output, with the time of each line
stand-in/requests.jsonlEach request the stand-in model was sent
messages.jsonlEvery message Many sent, a line each

Every value: results.json

Each harness's measures are under harnesses, each by its key, as the Single and Many pages name them (idle_memory, added_per_message):

{"harnesses": {"link": {"idle_memory": {"value": 21.6, "n": 21.6, "spread": [21.5, 21.7], "runs": 6, "by": "lnk measure", "raw": []}}}}
FieldWhat
valueThe value, as the tables show it: a number, or words with their number in n
nThe number compared across harnesses
spreadThe least and the most of its runs, where it has more than one
runsHow many runs its value is the median of
byWhat measured it: lnk measure, the bench's own reading of /proc, a proxy's log
whyIn place of all of these, why it has no value
rawWhat it was worked out from: each look, each run's numbers

A measure taken once in each provider's format (Single's message group) has an entry for each, under anthropic and openai, or local. One of Many's has an entry for each step, under its number of agents ("20"), and steps lists each harness's steps in order.

runs holds what each measure came from, by harness: each install (setup), each message (messages), each real task's run (real), and what its adapter's needs said for each provider and channel (needs). Many's are each step's messages, under its number of agents. In Single, a harness whose group failed lists why in its errors.

Single's real group adds prices, the pinned prices it charged by, and real_spent, the dollars it spent.

The machine and the settings: about.json

FieldWhat
dateWhen it started
lnk, adapterslnk --version, and each harness's adapter's --version
machineIts name, OS and kernel, CPUs and their model, memory
note--note's line
harnessesThe harnesses measured, in the tables' order
spec, quick, settingsThe spec it ran, whether it's a quick one, and every count and wait it used, flags included
loadThe machine's load averages (1, 5 and 15 minutes) as each group or harness began, and at the end
agents run asProcesses of the run, or the machine's services

Single adds groups, model_via, the real group's model, the order it went in, and each harness's message paths: how the bench sent it a message. Many adds its steps, and messages, the load in words.

Each message: messages.jsonl

Many writes a line for each message as each step ends, and Single keeps its messages in results.json's runs. Each holds the agent, its marker, when it was sent (sent), when its first piece came back (first_piece) and its end (end), all in Unix seconds, the answer's text or its error, whether it was the agent's first since it started, the model's requests, when the first came (model_at) and their time (model_time), and its processes' memory beyond idle at most (peak), the most ports they listened on, and the memory servers started under them (tool_servers).

What the model was sent: stand-in/requests.jsonl

A line for each request, with when it came (at), when its answer was decided and ended (decided, end), the provider it played and the path, the message's marker, and the agent whose key it carried. Single keeps each request's body whole (body, and raw as sent), which its message group reads. Many keeps no bodies, as thousands of turns would fill the machine's memory the agents are measured by.