Decisions

Why Link Harness works the way it does. Link's own decisions about harnesses are on its Agents Decisions page.

Design

Link Harness, the plain word, no brand. You type lnk agent use link; its program is lnk-link-harness and its plugin link-harness. KNOWN (harness settings) maps the name you type to the plugin, and an adapter from outside Link stays lnk-<name>, found by its own name.

For the densest deployment, run at the safest: the agent is always an argument of a turn, and each request names it; settings are read at each turn; every path comes from the agent through one function; and one turn runs at a time per conversation, a lock in the process. Each has one implementation now, one agent per process.

Nothing that matters when its process dies. It's a small router: each turn reads what it needs through its store, and writes there. The events are the truth, each number written once. Everything else (where the window starts, the copy for people, its settings as last read, each tool's list of calls) is derived, and made again by a process without it. So any process on any machine can take the next turn.

Everything it reads and writes: its settings (config.json, written by configure through the store too), its instructions, its conversations and where each one's window starts. One seam per resource, keyed by agent: the store has one implementation, the agent's folders on disk, and the code is the same on every machine. The workspace, the folder an agent's tools work in, is the agent's files folder, named by the same one function. One reader goes around the store: Link Memory reads the conversations' folders on disk itself, so a store kept anywhere else would need memory to read it there too.

From sections, each with an owner and a most, composed in the order of how often each changes, least first: Link's base, the model's, the tools', the environment (its files folder and channels, as Link resolved them at its start), its skills, your standing instructions (instructions/instruction.md, read at each turn, so you edit it like any other file), its notes, and now. A provider caches a prompt up to its first change, so what changes least goes first, and the system prompt stays the same from turn to turn: the date, the time, the channel and who's speaking change every message, so they head each message instead, kept with it so it reads the same at every turn after. Where each comes from, and how large each may be, are opinions, the spec's [instructions]; every section can be replaced, the base included. Everything is read through the store, so a file of yours is read as the conversations are, never through a link.

Conversations

As events: a folder per conversation, named by its thread, and a file per event, named by a zero-padded number so a listing is the conversation in order, never rewritten. An event holds the text, its time, the model, provider and response id, the tokens in, out and cached, the latency, and the cost only when the provider reports it, never estimated. A model's thinking is kept with its answer and sent back with its calls, never shown. A markdown copy beside them, added to as the agent answers, is for people and memory (what it holds). The folder is in the agent's files, so memory, a backup, a move and the next harness read it. Storage is one seam, keyed by agent and conversation, with this one implementation.

What does a conversation's copy hold?

What was said, up to the agent's last answer: the owner's messages and the agent's answers, never a tool's call or its answer, and never the owner's message while it's being answered. Memory searches the copy, so a search would otherwise find the question it's searching for, and the next search what the last one returned, an echo that grows with each call. The events keep everything, calls and answers included. A test of four turns shows it (Link Harness's tests/memory.rs), with a gate code from the chat's own past, beyond the window, found by memory.

Why is the copy added to rather than made again?

The copy holds what was said up to the agent's last answer, so it only grows, and only when the agent answers. Each answer adds what was said since the last one to its end, from the messages the turn already holds, rather than reading every message and writing the whole copy again. When the copy doesn't end with the last answer (it was cut short, or written elsewhere), it's made again from every message.

An agent runs in one place at a time, and a move carries its conversations with its files folder, so on your own machines nothing needs a store two machines share. Keeping them in the bucket instead, for your Mac and a box to answer in turn, would take each cloud's own API (rclone can't refuse a write whose name is taken, so it can't keep one turn at a time), a bucket tool outside the sandbox, and a copy here for memory. Backups already put them in your buckets, and the storage seam stays, with this one implementation.

Each group is a conversation of its own, conversations/<channel>-<group id>. Link's bridge takes your mention of the agent out of the message before it goes on, and a mention with nothing else is ignored. Each message is sent to Link Harness alone, and Link Harness adds the conversation before it.

The window

How much of a conversation does the model see?

A window: the conversation sent the same way each turn, so the provider's cache holds, as much of it as the model takes. Past that, the oldest turns drop half the window at a time, at points fixed by the conversation itself, so the prompt's start stays the same for many turns between drops.

Where does a model's window come from?

From what's known of the model, which is data, not code: Link's file of models' facts, refreshed at every release, under yours (~/.config/lnk/model-facts.toml), which wins field by field, so a model newer than your Link works without a release (lnk model facts). Link Harness takes the window from your file first, then from what a runtime on the machine says it runs the model with (it knows better than any file), then Link's, and gives a model nobody described the spec's small window ([window] unknown, 32,768 tokens) rather than a guess that overflows. Ollama runs a model with a smaller window than it was trained for unless told, so a trained window Ollama reports counts only up to that small one, unless you say otherwise.

How much of it to send is the spec's. Link's [window] max is held to the model's window; one you or a team set is taken as set, whatever Link knows of the model: the window you set always wins, and the only limit is the provider refusing a prompt as too long, which the harness reports as that, saying what it sent. The conversation's share is the window less the answer's ([model] answer, else [window] answer's share, at most the model's longest), the instructions and tools, and the margin ([window] margin).

Tokens are counted roughly before a turn, by characters (four ASCII characters to a token, one for any other, so Chinese or an emoji isn't undercounted), since each provider's tokenizer is its own; the margin covers the difference. After a turn, the provider's own counts are what's kept, measured and priced. The window's drop points stay on the rough count, which never changes for a message once written, so the provider's cache holds.

How does a turn find the window without reading every message?

Where the window starts depends on every message before it, so window.json beside the messages keeps how far the count got: a message the count goes on from, its time, and the tokens before it. A turn reads the messages from there, or from the agent's last answer if that's earlier, and counts on. The count only moves forward, since a drop point that didn't fit never fits again, so going on from any earlier mark finds the same start. The file is derived from the messages, never the only copy of anything: when it's missing, belongs to other messages or was made for another budget, the turn reads every message and writes it again. So any process on any machine can still take the next turn.

Tools

As an MCP client of the servers Link's tool bridge already runs outside the sandbox ($LNK_TOOL_<NAME>, and memory at $LNK_MEMORY), never as code of its own. A tool written once works for every harness, the calls you allow and the ones that ask are the same for each, and Link Harness still runs no code. Memory is the first. Each tool is offered to the model as <tool>__<call>, so two tools can have calls of the same name, and memory's are memory__search, memory__read, memory__list and memory__links. The tools are listed the same way each turn, so the provider's cache holds.

A subagent's task can come with a shorter list of tools (Link-Tools), which Link Harness keeps in tools.json beside the conversation's events. A later list narrows it again and never widens it, and a call to a tool the subagent wasn't offered is answered as one there isn't.

How are a tool's calls kept?

As events of the conversation, like any message. The model's answer that makes calls holds them, and each tool's answer is an event of its own after it, written as it happens. A turn cut short leaves the conversation as far as it got. A call left without an answer is answered in the next prompt as unanswered, so the prompt stays valid. With the tools gone, the calls and answers go to the model as text.

Does a turn hold anything of its tools?

Only memory's session, kept between turns. A turn opens a session with a tool only when the model calls it, so a message that calls no tool starts none. Memory's is kept for the turns after, and ended once none has used it for 10 minutes, or an hour after it started, so what it started without (its model, not up yet) is tried again: ending it with each turn started its process again with every message that called it (its sandbox, its model server and its index), and its embedding ran only inside turns. Every other tool's session ends with its turn. A kept session is one process for all of an agent's conversations at once, and a tool can keep what a conversation did in it, such as a browser's pages or a shell's folder: memory is the agent's own in every conversation, but another tool's state from one person's turn mustn't reach another's. Each tool's list of calls is kept in the process between turns, and asked again whenever a session starts. Both are derived from the tools: a process that hasn't listed them yet lists all of them at once when its first turn starts, and starts its own sessions. The conversation's events hold everything else the next turn needs, so any process on any machine can take it.

A tool that fails to start after it was listed is still offered, and a call to it is answered as unavailable. So the tools, and the prompt they start, stay the same, and the provider's cache holds. A tool that was never listed and doesn't answer is left out of that turn only.

How many calls does a turn make?

At most eight rounds. The model is then asked once more and told to call nothing, so a model stuck in a loop still answers. Each answer is kept up to 32 KiB, so one answer can't fill the window. The tools one round calls run at once. A message that steers a running turn is answered even past the last round of calls.