The Protocol
The tunnel's Service interface, the agent↔relay wire, changing it
without breaking old peers, and login.
The tunnel's Service interface
pub trait Service: Send + Sync + 'static {
fn call(&self, req: http::Request<Body>) -> ResponseFuture<'_>; // -> Result<http::Response<Body>, BoxError>
fn websocket(&self, req: http::Request<()>) -> WsFuture<'_> { /* 501 by default */ }
}link-agentturns each relay stream into anhttp::Requestand streams the returned body back chunk by chunk. The request body streams in as the visitor uploads it (link_agent::read_bodycollects it); reading slowly is fine, flow control makes the upload wait.- A service never sees stream ids, the relay, or its WebSocket.
- For visitors' WebSockets, return
WsUpgrade::Acceptedwith aWsSocket(a pair of message channels) fromwebsocket, or a normal response.link_agent::websocket::connectbridges to a localws://server. Errbecomes a 502 for the visitor.exposes_an_in_process_serviceinsrc/tests/e2eis the smallest example.
Wire protocol
The agent (lnk tunnel open) connects to wss://<relay>/_link/connect
(ws:// when the relay URL is http://). Each binary WebSocket frame
carries one MessagePack message, encoded with named fields
(to_vec_named): AgentMsg and RelayMsg in
src/shared/protocol/src/lib.rs.
agent → relay Hello { version, token, name?, client_version?, basic_auth?, websocket, flow_control, stream_window, stream_window_max, github }
relay → agent Welcome { name, public_url, latest_client?, notice?, auth_enforced, flow_control, stream_window, stream_window_max, github_enforced }
| Rejected { reason } (also sent later: kicked, blocked, or its name taken by a newer connection)
relay → agent Request { stream, method, uri, headers, body, websocket, body_follows, body_length? }
relay → agent RequestBody { stream, chunk } … RequestEnd { stream } (streamed upload, flow control)
agent → relay ResponseHead { stream, status, headers }
agent → relay ResponseBody { stream, chunk } (0..n, at most 256 KiB each)
agent → relay ResponseEnd { stream } | ResponseAbort { stream, reason }
both ways WsFrame { stream, frame: Text | Binary | Close } (after a 101 head)
both ways Credit { stream, bytes } (flow control: room to send more)
relay → agent Cancel { stream } (the visitor left)- Handshake.
Hellocarries the protocolversion, the token, the app name asked for, the tunnel's password (basic_auth) and the GitHub accounts it lets in (github). The relay answersWelcome(the name and URL it assigned, the latestlnkrelease, an operator notice, andauth_enforcedandgithub_enforcedconfirming it checks the password and the GitHub sign-in) orRejected, after which the agent stops instead of retrying. The upgrade request carries the token too, asAuthorization: Bearer, so the relay can let the connection skip the handshake slots strangers share beforeHelloarrives; a relay that doesn't read it usesHello's alone. - Streams. Each public request is a stream: a relay-allocated
u64id, unique per connection; an agent closes a connection whose relay sends aRequestfor a stream still running. Any number are in flight at once, and their chunks interleave on the one connection. Every stream ends withResponseEndorResponseAbort, or withCancelfrom the relay. Keep that invariant:dispatchremoves a stream on End/Abort, andStreamGuard::dropremoves it and sendsCancelotherwise. - Flow control. Each stream has a window in each direction,
counting
MESSAGE_COST(64 bytes) per message on top of the payload; the receiver grants more withCreditas data is actually read. The agent asks for a window inHello.stream_windowand the relay answers the one both use inWelcome.stream_window: at mostMAX_STREAM_WINDOW(2 MiB), andSTREAM_WINDOW(640 KiB) with a peer that omits it (link_protocol::stream_window). One stream moves at most a window per round trip between the agent and the relay, so a window can grow: a receiver may grant more credit than was used, up to what both agree on instream_window_max(at mostMAX_GROWN_WINDOW, 16 MiB; a peer that omits it doesn't grow,link_protocol::grown_window). The relay grows a download's window by what the visitor took each time the visitor has taken everything waiting, while its response budget is less than half full, and shrinks it back as the visitor reads once the budget is past half; credit never takes the budget's last quarter, kept for new streams' first windows.lnkgrows an upload's window by what the local app took each time it took everything waiting. Either side gives up on a stream afterCREDIT_TIMEOUT(5 minutes) without credit. With flow control, request bodies followRequestasRequestBodychunks (body_follows); without it they arrive whole inRequest.body, up toMAX_REQUEST_BODY(8 MiB). - Writing. Each side reads and writes its connection independently,
and keeps two queues to write from: one for what keeps streams moving
(
Request,ResponseHead,Credit,Cancel,ResponseAbort), always written first, and one for body data andWsFrames, which takes the streams in turn, a message each, so a download holds up another stream's body by one piece. Bodies go in pieces of 64 KiB at most. A stream's head still goes out before its body. Messages already waiting go out in one flush, and both sides setTCP_NODELAY, andTCP_NOTSENT_LOWATat 128 KiB where the system has it (lnkon Linux and macOS, the relay on Linux), so a small message waits behind at most that much in the kernel, not a whole send buffer of a download on a slow link. The relay sets it on agents' connections only: a visitor's carries one response at a time, and a visitor who stops reading should fill its send buffer, not the response budget that other visitors' first windows need. - WebSockets. A
Requestwithwebsocketset answered by a 101 head becomes a WebSocket: the relay and the agent each terminate their own, one to the visitor and one tows://localhost:<port>, and pass messages asWsFrames. Pings, fragmentation and compression stay per hop; the subprotocol the local server picks is passed back. - Liveness. Both sides ping every 15 s (
PING_INTERVAL) and drop a connection when nothing, not even a pong, has arrived for 45 s (IDLE_TIMEOUT), however much they are sending. On a drop the relay fails that tunnel's in-flight requests with 502 and frees its name; the agent reconnects with backoff and asks for the same name. Its backoff starts again from a second only after a connection that lasted 30 s. A newer connection with the same name takes it over, and the older one getsRejectedand stops. - Headers. Hop-by-hop headers, plus
hostandcontent-length, are stripped inis_hop_by_hopin the protocol crate; each hop sets its own. The agent sendsHost: localhost:<port>, so dev servers that check Host still work. - Size limits both sides enforce (
MAX_AGENT_MESSAGE,MAX_RELAY_MESSAGE,MAX_CHUNK,MAX_WS_MESSAGE) are constants in the protocol crate; what they mean for users is the limits table in Security.
Changing the protocol
Deployed relays and installed CLIs upgrade at different times, so every change keeps old peers working:
- Prefer additive fields with
#[serde(default)]: old peers ignore them, and they default when missing. No version bump.additive_fields_are_compatiblein the protocol crate shows the pattern. - Breaking changes (renamed or removed fields, changed meaning,
anything old peers can't decode) bump
link_protocol::VERSION. The relay serves agents fromMIN_AGENT_VERSIONup (insrc/server/relay/src/session.rs): a bump leaves it at the old version, and it is raised only after the release that dropped it has been out a while. Agents below it are told to upgrade. - Negotiate behavior the same way:
Hello.flow_controlandWelcome.flow_controlswitch on per-stream windows (Credit) and streamed request bodies only when both sides support them, andstream_windowenlarges the window the same way, so each keeps its old behavior with old peers.src/tests/e2ecovers both pairings (old_agent,new_agents_work_with_old_relays).
Login
Code: the relay's side is src/server/relay/src/login.rs (device flow,
the invite list, open signups) and login/approve.rs (the web flow and
its approval page); tokens and who they belong to are users.rs
(Origin::Login, Origin::Configured). The CLI's side,
src/local/agent/src/login.rs, only shows the link and code and polls;
lnk auth saves the token in ~/.config/lnk/config.toml.
The endpoints, JSON over HTTPS on the relay's bare domain (types in
link_protocol::login):
POST /_link/login/start StartLogin -> LoginStarted
POST /_link/login/poll PollLogin -> 202 LoginPending | 200 LoginResponse | 4xx text
GET /_link/login/confirm Authorization: Bearer <token> -> 200 [WaitingLogin] | 401
POST /_link/login/confirm Authorization: Bearer <token>, ConfirmLogin -> 204 | 401 | 404
POST /_link/login/ticket Authorization: Bearer <token>, no body -> 200 LoginTicket | 401 | 429
POST /_link/logout Authorization: Bearer <token>, no body -> 204 | 401 | 403
GET /_link/whoami Authorization: Bearer <token> -> 200 Whoami | 401Once the account has a machine signed in (a login token), a login
approved in the browser waits for one of them to confirm it: the poll
answers 202 with LoginPending.confirm ({"id", "username", "machines", "expires_in"}) until then. GET /_link/login/confirm lists
the logins waiting for the token's account (WaitingLogin: id,
device, from, same_network, started, code, expires_in), and
POST answers one ({"id", "approve"}), 404 for one that isn't
waiting for that account. Both take a login token or one from the
operator's users file, never a short token. ticket gives a
confirmation ahead ({"ticket", "expires_in"}), which a new login
carries in StartLogin.ticket; at most 5 per account at once (429).
Every field is additive: a lnk says it can wait with
StartLogin.confirm, and where a confirmation is needed a start without
it is refused at the poll (403, saying to upgrade) rather than left
waiting. An older relay ignores both fields and answers the confirm and
ticket paths 404, which lnk reads as "too old to confirm logins".
logout revokes the login token it carries, saved before the 204. An
unknown or missing token is 401, and one from the operator's users file
403. whoami says whose a login token or a short token (lnk auth token --json) is: {"username", "user_id"}, user_id the GitHub
account's id (left out for a user from the operator's users file), and
401 for a token the relay doesn't take. A service can sign a relay's
user in with a short token that way, never holding their login token.
The rules
(who can log in, what a token is bound to, which names a connection can
claim) are in Security.
src/tests/e2e/tests/login.rs runs the device flow end to end against
a fake GitHub (an axum app serving /login/device/code,
/login/oauth/access_token and /user): the relay is started in
process with its GithubLogin pointed at the fake, and the test calls
link_agent::login::login directly, then exposes a tunnel with the
token it gets, and the confirmation from a machine signed in, tickets
and an older lnk's refusal. approve.rs does the same for the web
flow, playing the browser on the relay's approval page, phished logins
included, and signups.rs covers open signups. Extend them rather than testing against real GitHub. By hand,
point a local relay at a fake with its hidden --github-api and
--github-url flags, and run LNK_NO_BROWSER=1 lnk auth login --relay http://localhost:7080 so it prints the link instead of opening a
browser; the CLI has no GitHub setting of its own.
Why it works this way: decisions.md.
