On Your Computer

Free and private: nothing you say leaves this computer. lnk installs a model with Ollama, or uses the runtime you already run, and serves it as an OpenAI-compatible API for anything else you run.

Install a model

lnk model recommend        # what install would pick here, and why
lnk model install          # install that one, and Ollama if it's missing
lnk model install qwen3:8b # any model in Ollama's library

install asks before installing Ollama; --yes doesn't ask. On Apple Silicon it picks the largest Qwen3 that leaves room for everything else:

MemoryModel
8 GBqwen3:4b, about 2.5 GB
16 GBqwen3:8b, 5.2 GB
32 GB or moreqwen3:14b, 9.3 GB

Anywhere but an Apple Silicon Mac with 8 GB or more, it recommends none, but install <model> still works.

See what's running

lnk model list           # the runtimes answering here, and their models
lnk model list --start   # start Ollama first if it's installed but stopped

list finds Ollama (port 11434), LM Studio (1234), llama.cpp (8080), vLLM (8000) and Jan (1337) at their default ports, and the addresses you add (For Developers).

Serve one on this machine

lnk model serve qwen3 --port 9100 # just this model, at http://127.0.0.1:9100/v1
lnk model serve llm               # every model of the runtime running here

It serves until you stop it, and starts Ollama if it's installed but stopped. A named model is used for every request, whatever model the request asks for.

  • --port <n> picks the local port. Any free one otherwise.
  • --upstream <url> points at any other OpenAI-compatible server, or a runtime on a port that isn't its default.
  • --model <name> with llm is the same as naming the model.

Remove a model

lnk model remove qwen3   # qwen3 finds qwen3:8b; --yes doesn't ask

This deletes the model from Ollama and frees its disk space.

Use it from your agent

lnk agent start offers the models running here, or the recommended one to install, so your agent needs no API key. See Switch Models.

lnk model install qwen3-embedding   # your agent's memory searches by meaning with it

An embedding model on this machine lets memory find a note by what it means, not only its words.