Models
The models on your computer, as an API. lnk installs one for you
with Ollama, or finds the one you already run, and serves it as an
OpenAI-compatible API, here or, through a tunnel, from
anywhere. Your agent can use it too, with no API key.
Needs: macOS or Linux. The model lnk picks for you is for Apple Silicon
Macs; elsewhere you name one. Sharing a model also needs
tunnels. It's free: models run on your own hardware.
Quickstart
lnk up models tunnels # add models, and tunnels to share them
lnk model install # Ollama and the model that suits this computer
lnk model list # the models running here
lnk tunnel open qwen3 --name qwen --auth friend:secret # https://qwen.<you>.local.link/v1Install a model
lnk model recommend # which model suits this computer, and why
lnk model install # install it, and Ollama if it's missing
lnk model install qwen3:8b # or pick one from Ollama's libraryIt asks before installing Ollama. On an Apple Silicon Mac it picks the largest Qwen3 your memory runs well.
Remove a model
lnk model remove qwen3Deletes the model and frees its disk space.
See what's running
lnk model listlnk looks for each runtime at its usual port: Ollama at 11434, LM
Studio at 1234, llama.cpp at 8080, vLLM at 8000 and Jan at 1337. It
also asks the addresses you add.
Use a model at its own address
lnk model add http://localhost:5000 --name mymodel # remember your server; list shows its models
lnk model forget mymodel # the address or the name; the server keeps runningYour model is then one like Ollama's: lnk agent start --model <model>
uses it, and lnk tunnel open <model> shares it. Your server needs to
answer OpenAI's /v1/models and /v1/chat/completions. Only an address
on this computer can be added.
Serve one on this computer
lnk model serve qwen3 --port 9100 # http://127.0.0.1:9100/v1, just this model
lnk model serve llm # every model of the runtime running hereTokens stream as they're made. --upstream <url> points at a runtime on
another port.
Share one from anywhere
lnk tunnel open qwen3 --name qwen --auth friend:secret
curl -N -u friend:secret https://qwen.<you>.local.link/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"qwen3","stream":true,"messages":[{"role":"user","content":"hi"}]}'Keep the password: an open model is costly to abuse. OpenAI's SDKs send
their own Authorization header, so give them the password as a Basic
Authorization header in default_headers, with any api_key.
Demo a model you're building
lnk tunnel open llm --upstream http://localhost:5000 --name mymodel --auth demo:secret # its OpenAI-style /v1 API
lnk tunnel open 5000 --name mymodel --auth demo:secret # any other API, every route as isA model you're training or serving from your own code works like any
other. The first command needs your server to answer OpenAI's
/v1/models and /v1/chat/completions, as vLLM and llama.cpp do. The
second forwards whatever your server answers, such as a Gradio app.
Restart your server as often as you like: the URL stays the same.
For your agent to use it too, add it with lnk model add.
Use it from your agent
lnk agent start offers the models it finds here before it asks for an
API key, and on an Apple Silicon Mac with none, installs one. See
Agents.
lnk model install qwen3-embeddingWith an embedding model here, your agent's memory finds a note by what it means, not only its words (Agents).
What works
- Ollama, LM Studio, llama.cpp, vLLM and Jan, at their usual ports.
- Only the OpenAI-compatible
/v1/API is served: the runtime's own management endpoints, such as Ollama's pulls and deletes, stay private. lnk model installandremovework with Ollama; other runtimes you manage yourself.
Every flag: lnk model <command> --help.
