Decisions

Why buckets work the way they do.

The engine

Why is rclone the engine?

rclone already reaches many clouds, and its check and cryptcheck compare checksums, through encryption too. The plugin pins rclone 1.75.1, downloads it on first use, and installs it only if its SHA-256 matches the one in the plugin's source. It runs with its own config file and no RCLONE_* variables.

How does a bucket sign in to AWS or Google Cloud?

As its lnk cloud account, the one your boxes there use: the cloud's adapter hands rclone the account's sign-in for each command, and rclone.conf keeps none of it. Buckets had signed in their own way, which meant a second set of bugs and a service account that worked for boxes but not buckets. Why, and how a Google Cloud account signs in: clouds and boxes' decisions.

Why is everything else rclone does passed through, not wrapped?

Link's commands stay the safe path, and rclone's power is one command away. create takes any rclone storage type and lets rclone ask its own questions, lnk bucket rclone runs rclone itself on the connected clouds, and push and pull pass flags after --.

Why rclone's crypt for encryption?

It encrypts file names as well as contents, and cryptcheck keeps clean possible through the encryption. Encrypting with age before upload would do neither. It's set per cloud at create, and always on for an agent's files.

Pushing and cleaning

Why does push never sync deletes or renames?

A sync that deletes is how backups get lost. push only adds, and replaces only with --replace; remove deletes from clouds on purpose.

How are good copies protected from bad local files?

push refuses to replace a copy that differs from the file here. Link doesn't turn on bucket versioning instead: rclone can't set it up the same way on every cloud, and it keeps paying for old versions.

When is a copy "checked"?

When its checksum matches the file here. Where the cloud has no checksum, Link reads it byte for byte; through encryption, it uses cryptcheck. The manifest records these checks as a cache for tree: clean checks again every time, but for an S3 copy it read back or proved before, unchanged here and there since.

Why does clean read some S3 copies back?

On S3, rclone takes the MD5 of an object uploaded in parts from metadata its uploader set, which S3 never checks, so anyone who can write to the bucket could forge it. Reading every copy back would cost a full download on every clean, and would never work on archived copies. Link instead asks for each object's metadata and reads back only those carrying an MD5 there; every other S3 MD5 is the ETag, which S3 computed. rclone hides the ETag and drops that metadata from its own listing, so Link lists with rclone's --s3-directory-bucket, under which an object's MD5 shows only when it comes from the metadata. push keeps comparing checksums: it is not what deletes files.

What a clean proved is kept, so a dry run and then a clean, or a clean tried again, don't download every large copy each time. The copy is known by the time S3 wrote it, from a listing that costs one request per thousand objects: S3 alone sets that time, so a copy replaced since shows another, and is checked again. The same listing, taken again after the check, catches a copy replaced while it ran.

Why does push skip files unchanged since their check?

Reading every file to hash it, once to send and again to check, on every push and for every cloud, makes a push of a large folder take as long as the first one. push skips a file whose size and modification time match the manifest and whose copy there was checked, and checks only what it sent, like rsync without --checksum. push --all reads everything, and clean checks every file before deleting any, taking an S3 copy as proved only by the time S3 wrote it.

Why does push send to every cloud at once?

A push to three clouds shouldn't take three times as long. Each cloud is sent and checked in a thread of its own, with rclone's progress as a line every few seconds after the cloud's name, since its redrawn progress would mix. What only happens once, like a sign-in or the Keychain, is asked first, one at a time. The manifest is written after, one cloud after another, so no cloud's record overwrites another's.

Why is there no copy of the manifest in each cloud?

The bucket itself is the truth. A new machine lists it (connect --bucket, tree --refresh) and loses only the "checked" times, which clean redoes regardless.

Buckets

Why a bucket per folder, not one bucket with a folder per part?

Every cloud can limit a key to one bucket simply, but not to a prefix. A bucket per folder (~/Link/Personal, and each agent's files folder) also lets each have its own settings, so an agent's is always encrypted, and yours holds exactly your folder. The path picks the bucket. Your files' bucket is made at create; an agent's on its first push. Before each agent had a folder of its own, a single agent folder had a bucket; that bucket is now main's.

Why are bucket names random?

Bucket names are global and anyone can probe whether one exists. A new bucket is local-link- plus 12 random letters and digits, with no username or account name. On storage with folders, it's a local-link folder holding personal and one folder per agent ID.

Why are create and connect two commands?

A command named connect that also makes a bucket surprises people. And because connect stops when the bucket isn't there, a typo in --bucket never makes a new one. lnk cloud connect makes no bucket either; it prints the create command.

Why does delete delete everything in the buckets?

So that what Link made can be removed with Link. It still refuses, without --everywhere, when a file's only copy is there. disconnect stays the way to forget a cloud and keep its files.

Agents' files

How does a box back up an agent's files?

With a key that reaches only its agent's bucket: an IAM user or a service account per bucket, made through your account when the agent arrives, and replaced or revoked when it leaves. A box holds no account of yours, so a box broken into reaches that one bucket, still encrypted with the agent's password.

How does another computer find an agent's bucket?

By labels on the bucket: the agent's ID and name, and your GitHub id, found through lnk cloud. No list is kept anywhere; the bucket is its own record, as a box is by its tags.

How does an agent's password move with it?

It's carried in the move's stream, as rclone keeps it, and removed from the machine the agent left. There's one password per agent, whichever cloud or machine, and it never goes on a command line.

Why does the bucket plugin restore an agent's buckets?

Finding an agent's buckets by their labels, asking their password, making a key and putting them in place is bucket knowledge, so lnk bucket restore does all of it and pulls the files, and lnk agent restore runs that one command before unpacking the harness's state. The harness never builds what lnk bucket adopt reads, and a change to how buckets are found or adopted stays in the bucket plugin.

Straight from your cloud, through URLs it signs on your machine (S3 presigned URLs, through rclone's link). The relay holds only the URLs, so files never pass through Local Link and cloud keys never leave the machine. That's also why a link lasts at most 7 days, S3's longest signature.

Why can only S3 storage share?

Share links need signed URLs that expire. rclone can't sign Google Cloud Storage URLs, which need a service account's private key, and other storage's links, such as Dropbox or Drive, don't expire. So only aws and s3 clouds share. Encrypted buckets are refused, an agent's always, because a link would hand out ciphertext.

On your own host, https://<you>.local.link/b/<id>, with an id of 24 random letters and digits. Who shared it is in the address, and one host per user holds their links; it gets a certificate only while they have one.

Why do some share pages show the visitor warning?

A share page is the relay's, under its name, and links to whatever URLs the sharer sent. The relay takes only URLs shaped like S3's presigned ones, but S3-compatible storage can be any host, so it can't tell storage from a site dressed as it. So an open-signup user's share page shows the visitor warning, as their tunnels do.

Through lnk auth token --json, a token for the share API that expires in 10 minutes. It doesn't read the tunnel plugin's saved login token: plugins never read each other's files, and the bucket plugin never holds a token that runs tunnels.

As for tunnels: --auth user:password, checked by the relay with HTTP Basic, which browsers and lnk bucket pull both speak. The relay keeps only its hash.

What does unshare end?

The link on the relay. URLs already handed out can't be recalled, since a presigned URL is valid until it expires. So the page says how long it lasts, and --expires can make a share short.