dedup
Content-address object bodies by SHA-256 hash so identical content is stored only once - re-uploads skip the byte upload and copies share one blob.
The built-in dedup() plugin stores each distinct body once. On upload it hashes the bytes (SHA-256), writes them a single time to a content-addressed blob under a store prefix (.dedup/ by default), and leaves a tiny pointer at your logical key. Upload the same content again — under any key — and the byte upload is skipped; only the pointer is written.
import { createFiles } from "files-sdk";
import { s3 } from "files-sdk/s3";
import { dedup } from "files-sdk/dedup";
const files = createFiles({
adapter: s3({ bucket: "uploads" }),
plugins: [dedup()],
});
await files.upload("a.png", bytes);
await files.upload("b.png", bytes); // same content — no second byte upload
await files.copy("a.png", "c.png"); // shares the one stored blob
How it works
The logical key (a.png) becomes an empty object whose metadata records the content hash; the bytes live at .dedup/<sha256>. Two keys with identical content point at the same blob, so the content is stored once no matter how many keys reference it.
uploadhashes the body, writes the blob only if that hash isn’t already stored (exists()), then writes the pointer.downloadfollows the pointer to the blob and returns it under your key — ranges included, because blobs are stored verbatim (unlikecompression()).head/listreport the logical content size with the internal fields stripped, without fetching the blob. Their body accessors (text(),stream(), …) are lazy: reading one fetches the blob, so it returns the content, never the empty pointer. List hides the blob store from normal listings. On adapters whose listing carries no metadata (S3 and the S3-compatible adapters), alist()item can’t be recognized as a pointer, so it reports the pointer’s own size (0) and ETag; its body still follows the pointer, andhead()/download()report the content size and hash as usual.copy/moverelocate the small pointer, so duplicating a de-duplicated file is near-free and the copy shares the original blob.
Every result reports the content hash as its etag (upload, download, head, and list alike). The pointer’s own ETag is the same for every key — its body is always empty — so the hash is what changes exactly when the content does. That keeps sync()’s default etag comparison and versioning() ids correct. The exception is list() on the S3 family (above): those items carry the pointer’s shared ETag and size 0, so a sync() between two de-duplicated instances there can’t see a content change and skips the key. Use transfer() to copy between them instead.
Bulk upload([...]) / download([...]) are de-duplicated per item. Objects without this plugin’s marker — pre-existing, or written by another tool — pass straight through on read, so it’s safe to enable on a bucket that already has data.
Options
| Option | Default | What it does |
|---|---|---|
prefix |
".dedup" |
Where the content-addressed blobs live. Hidden from list(), and write-protected. |
Objects under prefix are never themselves de-duplicated and are hidden from list() (unless you list within the prefix). You can read them, but writes into the store through the instance are refused: an upload, a signedUploadUrl(), or a copy / move whose destination is inside prefix throws a permanent FilesError. Otherwise anyone who can write a key could overwrite .dedup/<sha256> and change what every pointer to that content returns. The check resolves the key the way a filesystem would, so /.dedup/x, a/../.dedup/x, and .DEDUP/x are refused too.
Ordering
Put dedup() first, before any body-transforming plugin. Encrypted bytes don’t de-duplicate — a random per-object key makes identical inputs encrypt to different bytes — so de-duplication has to see the original content:
plugins: [dedup(), compression(), encryption(key)];
With this order the one stored blob is itself compressed and encrypted, and reads unwind the onion automatically (decrypt → decompress → follow the pointer).
Things to keep in mind
- It buffers the whole body to hash it, so — like
compression()— it’s unsuitable for very large or unknown-length streaming uploads.multipartand a resumablecontrolapply to the blob write only; the pointer is always one small write. When the content is already stored, no bytes move and thecontrolis left undriven. In front ofcompression()orencryption(), which refuse resumable uploads, acontrolon new content throws. - Reads cost a second fetch. A download reads the pointer, then the blob; a ranged download does a
headfirst.headandlistadd nothing — they only read the pointer. - It needs adapter metadata support. The hash round-trips through
metadata, the same gate a directmetadataupload hits. url()andsignedUploadUrl()fail closed. A presigned GET would hand out the empty pointer, not the content, and a presigned PUT would write directly and bypass content-addressing — so both throw. Download through theFilesinstance instead.- Conditional operations fail closed. A pointer’s native ETag is the same for every key, so no provider compare-and-set can guard its content; every conditional mode throws.
- It narrows
files.capabilities.signedUrl.supportedand everyconditionalflag readfalsewhile the plugin is installed;rangeReadstays whatever the adapter reports, since ranges apply to the verbatim blob. The gateway branches on these, so with the defaultdownloadMode: "auto"it streams downloads through the instance (following the pointer, with fullRange/206support) instead of redirecting to a URL the plugin refuses, and it uploads through the instance too. - Blobs aren’t garbage-collected.
delete(and an overwrite) drop the pointer but leave the content addressed — so it’s reused if the same bytes reappear. Reclaim unreferenced blobs with a storage lifecycle rule or a periodic sweep. Deleting a blob through the instance is allowed (that’s how a sweep works), but only delete blobs no pointer references: a pointer to a missing blob reads asNotFound.