Decisions
Why the models plugin works the way it does. What it does is the feature page; what it protects is Security.
What does Link manage for models?
Ollama. Link wraps it rather than running models itself: it installs the
official Ollama if it's missing, starts it, and pulls and deletes models
through Ollama's API. Link has no llama.cpp builds, model catalog or GPU
matrix of its own. install and remove work with Ollama only; other
runtimes the user installs, and Link finds and serves them.
What is the contract with a runtime?
The OpenAI-compatible /v1 API. list asks each runtime's default port
for GET /v1/models and takes whatever answers with an OpenAI-style
model list. serve forwards only /v1/... and streams the response
back, so a runtime's management API, such as Ollama's /api/pull and
/api/delete, stays private.
Where does Ollama come from?
Ollama's own downloads, not a copy Link pins and runs as its own service. The app on macOS starts itself at login and updates itself, and the Linux install script's system service survives reboots. The cost: Link trusts ollama.com's HTTPS and, on macOS, Apple's notarization, rather than checksums of its own.
Which model does Link recommend?
Qwen3, sized by memory, on Apple Silicon only. It calls tools, which
agents need, and runs well in 8 to 32 GB of unified memory. On Intel Macs
and servers without a GPU a local model is too slow for an agent, so an
API key is the better default there. lnk model install <model> still
works.
Where do a model's figures come from?
From a file of facts Link ships, not code: a model's window, longest
answer, how it thinks and its price change with every model released,
and a table in a harness would need a release to change. The file is
refreshed at every release from OpenRouter's public models API, which
needs no key and lists the models of every provider Link reaches, with
their windows, longest answers and prices per token; how a model thinks
is kept as Link wrote it, since OpenRouter says only whether it
reasons. A release whose refreshed file doesn't read back whole, or no
longer describes Link's own default models with figures a model could
have, stops. A change to a model the committed file has is bound by it:
a window more than halved, or a price moved more than tenfold, keeps
the committed figure and is said in the release's log, so a wrong
answer from OpenRouter can't shrink every agent's window or zero its
costs. The refreshed file is built in, not committed back, so the
release needs no write access to the repository. Yours
(model-facts.toml) sits over it, field by field, so a model released
after your Link works for years without a new one. Each harness gets
its agent's model's facts with its settings (model_facts), so any
harness, not only Link's, reads the same figures. Link Harness takes a
model's window from yours first, then from what a runtime on the machine
reports, then from Link's, and prices a turn its provider didn't price.
How does a model get a public URL?
Through the tunnel plugin. lnk tunnel open <model> starts lnk model serve <model> --json on a local port, tunnels to it, and stops it with
the tunnel. The models plugin never talks to a relay.
How does a model at its own address reach an agent?
lnk model add <url> remembers it, and discovery asks it as it asks the
default ports. So lnk model list --json, and with it lnk agent start,
the tunnel and every harness, see it without a flag of their own, and a
model someone is building needs no port trick. Only an address on this
machine is added: discovery never goes through a proxy, and the sandbox
lets a harness reach a local port, not another host. lnk model forget
drops it, and list shows an added address while nothing answers there,
so a server being restarted isn't forgotten.
