Decisions

Why clouds and boxes work the way they do.

Why its spec is what it is: Specs Decisions.

Cloud accounts

Whose cloud account do boxes use?

Yours, connected once with lnk cloud connect for both buckets and boxes. Local Link holds no provider accounts and bills nothing.

Where do cloud connections live?

A cloud plugin keeps the accounts, and one adapter plugin per cloud (lnk-aws, lnk-gcp) signs in to them and makes, starts, stops and removes machines, the way harnesses are adapters. A cloud is an engine to swap, so the cloud and box plugins never know which one they're on, and a new cloud is a new adapter. Other plugins read an account, keys included, with lnk cloud env <name> --json, and call its adapter with it themselves, as lnk box and lnk bucket both do: one way to reach an adapter, and the cloud plugin keeps no list of which commands another plugin may run. rclone is the storage engine, the cloud's CLI the machine engine, and both act as the account's one sign-in.

How do buckets and boxes sign in to the same cloud?

The same way, as the one lnk cloud account. The adapter alone knows how to sign in, and answers storage with rclone's settings for the account, so lnk bucket holds no sign-in code for AWS or Google Cloud and a login that works for boxes works for buckets. Two ways of signing in had meant two sets of bugs, and a service account that worked for boxes but not buckets.

How does a Google Cloud account sign in?

As one credential file, kept by the adapter: a service account's key, or your login's application default credentials. gcloud can act as a service account's key, and rclone and Google's libraries as either, but nothing but gcloud can use gcloud's own login. So connecting with your login signs you in once more (gcloud auth application-default login), which refreshes itself, and a service account is connected by its key (--key-file), which gcloud never hands out. A token per command would need no second sign-in, but fails any transfer over an hour.

Making a box

What is a box?

Just another machine running lnk, as your computer is. Tunnels, models and buckets work on a box by working on a headless Linux server. Only making a box, moving an agent to it and its exit are new.

How is a box made?

With the cloud's own CLI, aws or gcloud, as buckets use rclone. Every way of signing in to the cloud works, including profiles, SSO and gcloud's accounts, and nothing of the cloud's API is rebuilt.

In Link's settings folder, for you only, from the cloud's own installer, never with administrator rights or edits to the shell's settings. What lnk installs, lnk can find, and nothing else on the machine changes. The installer isn't pinned or verified, as the harness installers aren't.

How does a box sign in to the relay?

With lnk auth login on the box, approved by you in the browser like any other machine. It gets a token of its own, bound to its device. A new login on an account with a machine signed in also waits for one of them to confirm it; lnk box start hands the box's login a ticket from this machine, a confirmation given ahead, so you approve once (accounts). The box installs lnk from, and signs in to, the relay this computer uses, else local.link.

How are SSH host keys trusted?

They're read from the box's console at its first boot and pinned in Link's own known_hosts before the first connection. The computer that pinned them keeps them in the cloud, as metadata lnk-host-keys on Google Cloud or a lnk-host-key tag on AWS, for your other computers, since the console's first-boot output doesn't last. The cloud account is already what's trusted.

Why does a box start with lnk alone?

A box is a machine running lnk; what runs on it belongs to the plugin that runs it locally. So a new box gets the core and your login only, and lnk agent move (or lnk agent new --on) adds what an agent needs there first: the harness plugin, which brings the sandbox, the bucket plugin for its backups, and lnk sandbox ready --yes, the sandbox's own setup, which installs bubblewrap and its AppArmor profile through the box's sudo. The box plugin then knows no plugin but accounts and cloud, and a box never carries plugins its work doesn't use.

Does --yes make a box from a spec not seen before?

Yes, as --yes runs a sandbox's spec. lnk box start prints what it will make before anything else, even with --yes: the machine, its disk, the lnk it gets and the price when the cloud gives one. Without --yes it asks every time, whatever spec it's from, so there's nothing to remember of the specs it made. --yes is how a script, or lnk agent move, says yes ahead. A spec can cost money, but it has no paths, programs or account, so it can't reach anything on this computer: what it decides is the bill, which the printed price and the cloud's quotas bound.

Finding your boxes

How are boxes named?

With a name you type, like aws-1, and a random machine ID, which names its machine in the cloud (lnk-<id>). Two computers never make the same one, and the name you type isn't the cloud's. Boxes from before names and IDs are brought up to date: one named after its account, like aws, becomes aws-1, and the next time Link reaches it running, it gets its ID from the box's machine.toml and its tags.

Which boxes in an account are yours?

The ones tagged with your GitHub account's id (lnk-owner=github-<id>), so a project shared with another Local Link user lists only your boxes. The id, not the username, since usernames can be renamed. It isn't hashed: it's public anyway, and a hash of it could be reversed by trying ids.

How does a new computer get into your boxes?

Through the cloud account, which already controls them: the new computer's key is added by the cloud's own means. On AWS, EC2 Instance Connect lets a key in for a minute and lnk box appends it over that session, since an adapter never reads another plugin's key.

Where does the list of boxes live?

In the clouds, by tags; boxes.toml is this computer's copy. Boxes are found by their tags, their host keys are kept on them, and a new computer's key is added through the account. No machine, and not the relay, keeps the only record: lose your computer, keep your boxes.

What happens when several lnk box commands run at once?

Each change is read, made and saved under a lock (an empty *.lock file in ~/.config/lnk/box), held only while it's saved, not for the whole command: a start takes minutes, and lnk box list in another terminal shouldn't wait for it. Each command reads boxes.toml again just before it changes it, and changes only its own box. A start takes its name before making the box, so two at once get aws-1 and aws-2; one stopped before its cloud made the box leaves the name, with no id, until lnk box remove forgets it.

Images

Which images do boxes get?

Each adapter maps the OS to its image from a fixed table, by its publisher: Canonical's SSM parameters on AWS and its ubuntu-os-cloud project on Google Cloud, and for a box with GPUs AWS's Deep Learning Base AMI and Canonical's accelerator images. Never by searching image names, which anyone can publish under. Only Ubuntu LTS: a box assumes the user ubuntu, cloud-init and apt.

Is an image a field of the spec, or an argument?

An argument of lnk box start, --image. An image is in one account of one cloud, and a spec makes the same box in any account and either cloud, so a spec naming an image would make a box in one place only. The spec says the machine; the image says where its disk starts.

Why is a box stopped to be saved?

So its disk is saved whole. AWS can save a running machine without a reboot, but files being written are then saved half-written; Google Cloud refuses a disk in use without --force, for the same reason. lnk box image save stops the box, saves it, and starts it again, which costs a few minutes and a new address.

Why is a box's login removed before it's saved?

So no two boxes share one. A box's lnk login is a token bound to it; copied into an image, every box from it would act as the first, and removing one would cut them all off. lnk box image save signs the box's lnk out with lnk auth logout, the accounts plugin's own command, never by touching its files, and a box from the image signs in as any new box does. The saved box is signed in again after, in a terminal. For the same reason a box made from an image forgets the saved box's machine ID before its setup writes its own, and a box with agents on it isn't saved.

Ubuntu's cloud images delete their SSH host keys and make new ones at the first boot of each new machine, and print them on the console, so a box from an image has keys of its own, read as any new box's are. Link pins them under the box's own name, so the saved box's are never trusted for it; deleting the keys on the box before saving would add a step that cloud-init already does.

Which images can a box start from?

Only one Link made, found by its tags (lnk-image, lnk-owner), and only the account's own: on AWS --owners self, so a public AMI can't pass for one, and on Google Cloud the project's own. An adapter from before images answers none of the image commands, so lnk box says so and makes no box from one there, while it makes every other box as before.

Why does saving refuse a box holding secrets, not remove them?

A box saved as an image hands its disk to every box made from it, and the box saved is still yours, in use: deleting its git credentials or a cloud CLI's login would break it without a word. So save looks where secrets are usually left, refuses, and lists them; removing them is yours to do. A setup whose stages each run in a sandbox of their own keeps most of them off the disk in the first place.

How long does a box from an image take to start?

The docs promise no time. On AWS a box's first boot from an image loads its disk's blocks from the snapshot as they're first read, so how long it takes depends on what the setup reads, and only a run on a real account says.

Setting a box up

Is a box's setup a file of stages, or a script?

A script, yours, in plain sh, each stage a command in it. A spec describes only what its plugin owns and runs nothing, so things are combined by commands, never inside a file: a box's spec has no setup, and a sandbox's spec no script. A file of [[stage]] tables would be shorter to write, but it would be a file combining specs, owned by a plugin of its own. lnk box setup is a convenience over the script, which runs as well by hand.

Where does a setup run?

On boxes only. The script runs on this computer, where the cloud account and Link's SSH key are, and drives the box with lnk box commands, so it can make the box, set it up and save it as an image in one go. Setting up this computer's own folders in stages is the same stages run with lnk sandbox run here, with no image to save, so lnk box setup has no counterpart for it.

How does lnk box setup say which stage failed?

From sh's own trace: it runs the script with sh -e -x and a PS4 of its own, reads the trace from the script's error output, passes everything else through, and names the last command traced when the script fails. The script stays plain sh, with no markers to write and nothing Link must parse, and the trace is never shown, since it holds each command expanded, secrets in variables included.

How do files reach a box?

lnk box copy, over the box's SSH as lnk box ssh reaches it: Link's key, the host keys pinned for the box, the options that keep this computer's SSH agent and forwards off it, and the shared connection. Never through the relay, so a file is only ever between your computer and your box. Each file's bytes go on ssh's input, so no scp or sftp is needed on either end.

Why does installing system packages run outside the sandbox?

It needs root, which no sandbox gives. So a setup writes it as a stage of its own, lnk box ssh <box> -- sudo apt-get install -y ..., where it can be read for what it is, before the sandboxed stages.

How does a setup reuse an image?

By a hash the script makes of what made the image, such as the setup script, its specs and the lockfile its dependencies come from. lnk box image save --hash tags the image with it, and lnk box image list --hash and lnk box start --image-hash find it, so a script saves the image when nothing has that hash and starts from it when one does. Link never computes the hash: only the script knows which files make a stage. The hash is a tag, so the cloud keeps it, and any of your computers finds it; an adapter that doesn't say an image's hash finds none, and the script makes the image again.

Spot boxes, GPUs and limits

How does a spot box stay a box?

By being stopped, never deleted, when the cloud takes the machine back, so its disk, with its login and an agent's state, stays. On AWS that takes a persistent spot request, which would make the box again if only its machine went, so a request is cancelled with its box.

Which GPUs does a Google Cloud box get?

Those built into its machine type (G2, A2, A3, A4), which machine-types list gives with their counts, and T4s attached to an N1 machine, the only way Google Cloud has a T4. An N1 machine and its T4s are one size to the box plugin, named n1-standard-4.nvidia-tesla-t4.1. The box plugin matches and pins sizes by name and knows no cloud, so the name carries everything. The adapter lists each N1 machine again with every count of T4s the zone's accelerator-types list allows and Google permits with that many CPUs. It splits the name back when it makes the box, and names a box that way when Google says it has T4s, so lnk box spec makes the same box again. Only the T4 is attached: the other GPUs Google attaches to N1 (P4, P100, V100) are older than any Link names.

Does a box cap memory or disk on your own computer?

No: on a Mac only a VM can.