The Protocol

The tunnel's Service interface, the agent↔relay wire, changing it without breaking old peers, and login.

The tunnel's Service interface

pub trait Service: Send + Sync + 'static {
    fn call(&self, req: http::Request<Body>) -> ResponseFuture<'_>; // -> Result<http::Response<Body>, BoxError>
    fn websocket(&self, req: http::Request<()>) -> WsFuture<'_> { /* 501 by default */ }
}
  • link-agent turns each relay stream into an http::Request and streams the returned body back chunk by chunk. The request body streams in as the visitor uploads it (link_agent::read_body collects it); reading slowly is fine, flow control makes the upload wait.
  • A service never sees stream ids, the relay, or its WebSocket.
  • For visitors' WebSockets, return WsUpgrade::Accepted with a WsSocket (a pair of message channels) from websocket, or a normal response. link_agent::websocket::connect bridges to a local ws:// server.
  • Err becomes a 502 for the visitor.
  • exposes_an_in_process_service in src/tests/e2e is the smallest example.

Wire protocol

The agent (lnk tunnel open) connects to wss://<relay>/_link/connect (ws:// when the relay URL is http://). Each binary WebSocket frame carries one MessagePack message, encoded with named fields (to_vec_named): AgentMsg and RelayMsg in src/shared/protocol/src/lib.rs.

agent → relay   Hello { version, token, name?, client_version?, basic_auth?, websocket, flow_control, stream_window, stream_window_max, github }
relay → agent   Welcome { name, public_url, latest_client?, notice?, auth_enforced, flow_control, stream_window, stream_window_max, github_enforced }
              | Rejected { reason }                   (also sent later: kicked, blocked, or its name taken by a newer connection)

relay → agent   Request { stream, method, uri, headers, body, websocket, body_follows, body_length? }
relay → agent   RequestBody { stream, chunk } … RequestEnd { stream }   (streamed upload, flow control)
agent → relay   ResponseHead { stream, status, headers }
agent → relay   ResponseBody { stream, chunk }         (0..n, at most 256 KiB each)
agent → relay   ResponseEnd { stream }  |  ResponseAbort { stream, reason }
both ways       WsFrame { stream, frame: Text | Binary | Close }   (after a 101 head)
both ways       Credit { stream, bytes }               (flow control: room to send more)
relay → agent   Cancel { stream }                      (the visitor left)
  • Handshake. Hello carries the protocol version, the token, the app name asked for, the tunnel's password (basic_auth) and the GitHub accounts it lets in (github). The relay answers Welcome (the name and URL it assigned, the latest lnk release, an operator notice, and auth_enforced and github_enforced confirming it checks the password and the GitHub sign-in) or Rejected, after which the agent stops instead of retrying. The upgrade request carries the token too, as Authorization: Bearer, so the relay can let the connection skip the handshake slots strangers share before Hello arrives; a relay that doesn't read it uses Hello's alone.
  • Streams. Each public request is a stream: a relay-allocated u64 id, unique per connection; an agent closes a connection whose relay sends a Request for a stream still running. Any number are in flight at once, and their chunks interleave on the one connection. Every stream ends with ResponseEnd or ResponseAbort, or with Cancel from the relay. Keep that invariant: dispatch removes a stream on End/Abort, and StreamGuard::drop removes it and sends Cancel otherwise.
  • Flow control. Each stream has a window in each direction, counting MESSAGE_COST (64 bytes) per message on top of the payload; the receiver grants more with Credit as data is actually read. The agent asks for a window in Hello.stream_window and the relay answers the one both use in Welcome.stream_window: at most MAX_STREAM_WINDOW (2 MiB), and STREAM_WINDOW (640 KiB) with a peer that omits it (link_protocol::stream_window). One stream moves at most a window per round trip between the agent and the relay, so a window can grow: a receiver may grant more credit than was used, up to what both agree on in stream_window_max (at most MAX_GROWN_WINDOW, 16 MiB; a peer that omits it doesn't grow, link_protocol::grown_window). The relay grows a download's window by what the visitor took each time the visitor has taken everything waiting, while its response budget is less than half full, and shrinks it back as the visitor reads once the budget is past half; credit never takes the budget's last quarter, kept for new streams' first windows. lnk grows an upload's window by what the local app took each time it took everything waiting. Either side gives up on a stream after CREDIT_TIMEOUT (5 minutes) without credit. With flow control, request bodies follow Request as RequestBody chunks (body_follows); without it they arrive whole in Request.body, up to MAX_REQUEST_BODY (8 MiB).
  • Writing. Each side reads and writes its connection independently, and keeps two queues to write from: one for what keeps streams moving (Request, ResponseHead, Credit, Cancel, ResponseAbort), always written first, and one for body data and WsFrames, which takes the streams in turn, a message each, so a download holds up another stream's body by one piece. Bodies go in pieces of 64 KiB at most. A stream's head still goes out before its body. Messages already waiting go out in one flush, and both sides set TCP_NODELAY, and TCP_NOTSENT_LOWAT at 128 KiB where the system has it (lnk on Linux and macOS, the relay on Linux), so a small message waits behind at most that much in the kernel, not a whole send buffer of a download on a slow link. The relay sets it on agents' connections only: a visitor's carries one response at a time, and a visitor who stops reading should fill its send buffer, not the response budget that other visitors' first windows need.
  • WebSockets. A Request with websocket set answered by a 101 head becomes a WebSocket: the relay and the agent each terminate their own, one to the visitor and one to ws://localhost:<port>, and pass messages as WsFrames. Pings, fragmentation and compression stay per hop; the subprotocol the local server picks is passed back.
  • Liveness. Both sides ping every 15 s (PING_INTERVAL) and drop a connection when nothing, not even a pong, has arrived for 45 s (IDLE_TIMEOUT), however much they are sending. On a drop the relay fails that tunnel's in-flight requests with 502 and frees its name; the agent reconnects with backoff and asks for the same name. Its backoff starts again from a second only after a connection that lasted 30 s. A newer connection with the same name takes it over, and the older one gets Rejected and stops.
  • Headers. Hop-by-hop headers, plus host and content-length, are stripped in is_hop_by_hop in the protocol crate; each hop sets its own. The agent sends Host: localhost:<port>, so dev servers that check Host still work.
  • Size limits both sides enforce (MAX_AGENT_MESSAGE, MAX_RELAY_MESSAGE, MAX_CHUNK, MAX_WS_MESSAGE) are constants in the protocol crate; what they mean for users is the limits table in Security.

Changing the protocol

Deployed relays and installed CLIs upgrade at different times, so every change keeps old peers working:

  • Prefer additive fields with #[serde(default)]: old peers ignore them, and they default when missing. No version bump. additive_fields_are_compatible in the protocol crate shows the pattern.
  • Breaking changes (renamed or removed fields, changed meaning, anything old peers can't decode) bump link_protocol::VERSION. The relay serves agents from MIN_AGENT_VERSION up (in src/server/relay/src/session.rs): a bump leaves it at the old version, and it is raised only after the release that dropped it has been out a while. Agents below it are told to upgrade.
  • Negotiate behavior the same way: Hello.flow_control and Welcome.flow_control switch on per-stream windows (Credit) and streamed request bodies only when both sides support them, and stream_window enlarges the window the same way, so each keeps its old behavior with old peers. src/tests/e2e covers both pairings (old_agent, new_agents_work_with_old_relays).

Login

Code: the relay's side is src/server/relay/src/login.rs (device flow, the invite list, open signups) and login/approve.rs (the web flow and its approval page); tokens and who they belong to are users.rs (Origin::Login, Origin::Configured). The CLI's side, src/local/agent/src/login.rs, only shows the link and code and polls; lnk auth saves the token in ~/.config/lnk/config.toml.

The endpoints, JSON over HTTPS on the relay's bare domain (types in link_protocol::login):

POST /_link/login/start  StartLogin  -> LoginStarted
POST /_link/login/poll   PollLogin   -> 202 LoginPending | 200 LoginResponse | 4xx text
GET  /_link/login/confirm  Authorization: Bearer <token>              -> 200 [WaitingLogin] | 401
POST /_link/login/confirm  Authorization: Bearer <token>, ConfirmLogin -> 204 | 401 | 404
POST /_link/login/ticket   Authorization: Bearer <token>, no body      -> 200 LoginTicket | 401 | 429
POST /_link/logout       Authorization: Bearer <token>, no body -> 204 | 401 | 403
GET  /_link/whoami       Authorization: Bearer <token>          -> 200 Whoami | 401

Once the account has a machine signed in (a login token), a login approved in the browser waits for one of them to confirm it: the poll answers 202 with LoginPending.confirm ({"id", "username", "machines", "expires_in"}) until then. GET /_link/login/confirm lists the logins waiting for the token's account (WaitingLogin: id, device, from, same_network, started, code, expires_in), and POST answers one ({"id", "approve"}), 404 for one that isn't waiting for that account. Both take a login token or one from the operator's users file, never a short token. ticket gives a confirmation ahead ({"ticket", "expires_in"}), which a new login carries in StartLogin.ticket; at most 5 per account at once (429).

Every field is additive: a lnk says it can wait with StartLogin.confirm, and where a confirmation is needed a start without it is refused at the poll (403, saying to upgrade) rather than left waiting. An older relay ignores both fields and answers the confirm and ticket paths 404, which lnk reads as "too old to confirm logins".

logout revokes the login token it carries, saved before the 204. An unknown or missing token is 401, and one from the operator's users file 403. whoami says whose a login token or a short token (lnk auth token --json) is: {"username", "user_id"}, user_id the GitHub account's id (left out for a user from the operator's users file), and 401 for a token the relay doesn't take. A service can sign a relay's user in with a short token that way, never holding their login token. The rules (who can log in, what a token is bound to, which names a connection can claim) are in Security.

src/tests/e2e/tests/login.rs runs the device flow end to end against a fake GitHub (an axum app serving /login/device/code, /login/oauth/access_token and /user): the relay is started in process with its GithubLogin pointed at the fake, and the test calls link_agent::login::login directly, then exposes a tunnel with the token it gets, and the confirmation from a machine signed in, tickets and an older lnk's refusal. approve.rs does the same for the web flow, playing the browser on the relay's approval page, phished logins included, and signups.rs covers open signups. Extend them rather than testing against real GitHub. By hand, point a local relay at a fake with its hidden --github-api and --github-url flags, and run LNK_NO_BROWSER=1 lnk auth login --relay http://localhost:7080 so it prints the link instead of opening a browser; the CLI has no GitHub setting of its own.

Why it works this way: decisions.md.