Many

As many agents of one harness as this machine runs: each with Link Memory, each sent a message every ten minutes, in steps of 1, 5, 20, 50, 100 and on, an hour each, until one isn't answered in time. What each agent costs idle and per message is measured at every step, the same measures as Single's, under the same names.

lnk bench many --note "my box"         # Link Harness, as services
lnk bench many --harness link,hermes   # each harness named, in turn
lnk bench many --quick                 # two short steps, to try it

It runs Link Harness unless --harness names others, each measured in turn, its agents deleted before the next's. Any harness can be measured, though a harness that installs hundreds of megabytes takes that, and minutes, for each agent. On a machine of its own it runs as the machine's user, with its agents as that user's services. With --programs it runs a checkout's programs in a home of its own, and where there's no user's systemd, as in a container, its agents are processes of the run. It needs Linux with bubblewrap, and Link Memory (lnk plugin add link-memory).

How it measures

Each step's numbers are in the spec's [many]: steps, hold (an hour), every (ten minutes), and the rest below.

  • Each turn calls memory once, as a user's agent with memory would: the stand-in answers a message with a call to memory__search (Link Memory's, before any search of the harness's own), then the tool's answer with text, each after delay (1 s). So a turn is two model requests.
  • In time is the answer whole within slack (2 s) beyond what the model took over the turn. A message's answer must carry its marker, whatever the harness: else it wasn't the model's.
  • Messages are spread evenly over the ten minutes, not all at once, and each goes in its own thread of the run, as a person's would. A message reaches each harness as Single sends it, in one thread of conversation a step, so every step's agents answer with the same history behind them.
  • The stand-in is a local model at Ollama's own port, 11434, when OpenClaw is measured, else at any free port. OpenClaw takes a local model only as Ollama, so the stand-in speaks Ollama's API too, and an Ollama running there must be stopped first.
  • Ready is the harness answering, as in Single: its endpoint proves itself and its /health answers, or Hermes's API or OpenClaw's gateway does.
  • A step ends at its first failure, and the harness's steps with it: a message late or unanswered, an agent not made or not starting, or one not running after the agents settle. That step's measures are kept.

What it measures

At each step, per agent where it says so:

  • Most agents in time (held): the largest step whose every message was in time, and what failed first (failed_first).
  • Made in (make, s): from lnk agent create until the agent answers with its memory, the median of the step's new agents.
  • Idle memory, processes, ports and CPU (idle_memory, processes_ports, idle_cpu): each agent looked at every 5 s (look) for 15 s (idle), after the step's agents settle for 15 s (settle), read as in Single. Each is the median of the agents, with their spread.
  • Machine memory (machine_memory, MB per agent): the machine's memory in use beyond what it used before the harness's first agent, over the step's agents. A short run's can read below zero, as the machine's cache is freed.
  • Services (units): each agent's systemd units, which are none in the foreground.
  • Added per message, to model, model requests (added_per_message, to_model, model_requests): as in Single, over the step's messages. To model is apart for an agent's first message since it started (first_to_model), which starts what's idle.
  • CPU per message (cpu_per_message, core-ms): the agents' CPU over the step, less what they'd have used idle meanwhile, per message.
  • Machine busy (machine_busy, %): the machine's CPU over the step, over every CPU /proc/stat counts, without steal (time a cloud's hypervisor gave another machine).
  • Memory per message (peak_memory, MB): as Single's peak memory, each answering agent's beyond its idle, looked at every 0.1 s (watch).
  • Processes and tool servers started (tasks_started, tool_servers, per message): the machine's processes and threads started (/proc/stat, this run's own threads among them, so an upper bound), and the memory servers seen under answering agents.
  • Disk (disk_per_100, disk_per_1000_events, KB): the agents' folders' growth in blocks, their harness's install aside, per 100 messages, and for Link Harness per 1,000 events of its conversations, its memory's index among them.

What it writes

bench-many-<date>/, where it's run: results.md (each harness's most agents in time and what failed first, then a row of its measures a step), results.json (every value, its spread, what read it, or why it has none), about.json, messages.jsonl (every message), stand-in/requests.jsonl (when each model request came and its marker, never its body) and each agent's log in logs/.

At the end its agents are stopped and deleted (--keep leaves them running). A run finding an earlier run's agents (named --prefix, bench-many-) stops before it starts, naming them. The first Ctrl-C ends the step and stops the agents, and a second stops at once.

The self-test

lnk bench many --self-test runs two short steps of the mini harness against the stand-in, without memory, in about a minute. It fails if the last step didn't hold, or if any measure has neither a value nor a stated reason: Run it from a checkout.