On Your Computer
Free and private: nothing you say leaves this computer. lnk
installs a model with Ollama, or uses the runtime you already run, and
serves it as an OpenAI-compatible API for anything else you run.
Install a model
lnk model recommend # what install would pick here, and why
lnk model install # install that one, and Ollama if it's missing
lnk model install qwen3:8b # any model in Ollama's libraryinstall asks before installing Ollama; --yes doesn't ask. On Apple
Silicon it picks the largest Qwen3 that leaves room for everything else:
| Memory | Model |
|---|---|
| 8 GB | qwen3:4b, about 2.5 GB |
| 16 GB | qwen3:8b, 5.2 GB |
| 32 GB or more | qwen3:14b, 9.3 GB |
Anywhere but an Apple Silicon Mac with 8 GB or more, it recommends
none, but install <model> still works.
See what's running
lnk model list # the runtimes answering here, and their models
lnk model list --start # start Ollama first if it's installed but stoppedlist finds Ollama (port 11434), LM Studio (1234), llama.cpp (8080),
vLLM (8000) and Jan (1337) at their default ports, and the addresses you
add (For Developers).
Serve one on this machine
lnk model serve qwen3 --port 9100 # just this model, at http://127.0.0.1:9100/v1
lnk model serve llm # every model of the runtime running hereIt serves until you stop it, and starts Ollama if it's installed but stopped. A named model is used for every request, whatever model the request asks for.
--port <n>picks the local port. Any free one otherwise.--upstream <url>points at any other OpenAI-compatible server, or a runtime on a port that isn't its default.--model <name>withllmis the same as naming the model.
Remove a model
lnk model remove qwen3 # qwen3 finds qwen3:8b; --yes doesn't askThis deletes the model from Ollama and frees its disk space.
Use it from your agent
lnk agent start offers the models running here, or the recommended one
to install, so your agent needs no API key. See
Switch Models.
lnk model install qwen3-embedding # your agent's memory searches by meaning with itAn embedding model on this machine lets memory find a note by what it means, not only its words.
