Skip to content
Files SDK
Esc
↑↓navigate↵open⌘Jpreview
On this page

Migrate S3 files to R2 and verify the result

Copy an S3 bucket to Cloudflare R2 with a script that records failures, retries only those keys, checks every object by checksum, and plans a rollback.

Try Cloudflare’s own migration services first. Super Slurper copies a bucket in bulk from Cloudflare’s side, and Sippy copies objects into R2 as your app reads them. A Files SDK script earns its place when you need something they don’t give you: keys rewritten on the way, Cache-Control carried across explicitly, a cutover your deploy controls, or a verification report you produce yourself.

The script here uses transfer() to stream every object from S3 to R2, writes the keys that failed to a manifest, retries only those, and then checks every object on both sides by size, type, metadata, Cache-Control, and a SHA-256 it computes itself. It doesn’t compare ETags, because they differ across providers and part sizes even when the bytes match. It also copies a snapshot, so freeze writes to the bucket before you start. Nothing here makes the move zero-downtime or guarantees it loses nothing; the verification step is how you find out.

Before you start

  • AWS credentials that can list and read the source bucket (s3:ListBucket, s3:GetObject), and an R2 API token with Object Read & Write on the destination bucket, set as R2_ACCOUNT_ID, R2_ACCESS_KEY_ID, and R2_SECRET_ACCESS_KEY. This guide calls both buckets uploads.
  • Node 22.18 or later, or Bun, on a machine with good bandwidth to both providers. Every byte streams through it.
  • Written against files-sdk 3.0 and @aws-sdk/client-s3 3.1148.
npm install files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner
pnpm add files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner
yarn add files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner
bun add files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner
nub add files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner
aube add files-sdk @aws-sdk/client-s3 @aws-sdk/lib-storage @aws-sdk/s3-presigned-post @aws-sdk/s3-request-presigner

@aws-sdk/lib-storage matters here. transfer() hands each body to the destination as a stream of unknown length, and the AWS SDK engine uploads such streams as multipart through lib-storage. R2’s fetch engine would buffer each object in memory before a single PUT, so the script sets client: "aws-sdk".

Cloudflare’s tools or a script

Super Slurper Sippy Files SDK script
How it copies A bulk job that Cloudflare runs On demand, when a GetObject misses in R2 A process you run, streaming through your machine
Metadata Custom metadata is preserved Only x-amz-meta-* user metadata Content type and user metadata; Cache-Control with the plugin below
Existing objects in R2 Overwrite (default) or skip Served from R2 once copied Overwrite (default) or skip with overwrite: false
Limits Cloudflare documents Objects over 1 TB, and objects in Glacier tiers other than Instant Retrieval, are skipped Changes to a source object after its first copy aren’t reflected Whatever your machine and network sustain
Key rewrites Prefix selection only No transformKey

Cloudflare’s migration strategies combine the two for buckets that are still being written to: enable Sippy, point the app at R2, then run Super Slurper with “skip existing”. Both pages say ETags aren’t guaranteed to match after migration. Super Slurper’s page doesn’t say whether it copies Cache-Control, so check before you rely on it.

For a straight copy of a bucket you can pause, Super Slurper is less work than anything below. You can still run this guide’s verification script afterwards; it doesn’t care how the objects got there.

Freeze writes first

transfer() lists the whole source before it copies anything. An object written after that listing isn’t copied, and an object deleted after it fails with NotFound. Put the app into a state where nothing writes to the bucket for the length of the copy. If the app talks to storage through Files SDK, constructing its instance with readonly: true makes every write throw a ReadOnly FilesError instead of landing in a bucket you’re about to leave. See Read-only.

Configure both sides

These instances belong to the migration scripts only. Your app keeps its own Files instance.

import { HeadObjectCommand, S3Client } from "@aws-sdk/client-s3";
import { Files, handlers } from "files-sdk";
import { r2 } from "files-sdk/r2";
import { mapS3Error, s3 } from "files-sdk/s3";

export const SOURCE_BUCKET = "uploads";
export const DEST_BUCKET = "uploads";

// Credentials come from the AWS credential chain.
export const source = new Files({
  adapter: s3({ bucket: SOURCE_BUCKET, region: "us-east-1" }),
  retries: 3,
});

// The S3 client under a Files instance. r2() on its aws-sdk engine loads
// the client lazily, so call any method on that instance first.
export const s3ClientOf = (files: Files): S3Client => {
  if (!(files.raw instanceof S3Client)) {
    throw new Error("Expected an adapter backed by @aws-sdk/client-s3");
  }
  return files.raw;
};

// FileInfo has no Cache-Control field, so read it from the native client.
export const cacheControlOf = async (
  files: Files,
  bucket: string,
  key: string
): Promise<string | undefined> => {
  try {
    const head = await s3ClientOf(files).send(
      new HeadObjectCommand({ Bucket: bucket, Key: key })
    );
    return head.CacheControl;
  } catch (error) {
    // Raw client errors aren't FilesErrors; normalize before branching.
    const err = mapS3Error(error);
    if (err.code === "NotFound") {
      return undefined;
    }
    throw err;
  }
};

export const dest = new Files({
  // Reads R2_ACCOUNT_ID, R2_ACCESS_KEY_ID, and R2_SECRET_ACCESS_KEY.
  adapter: r2({ bucket: DEST_BUCKET, client: "aws-sdk" }),
  retries: 3,
  plugins: [
    {
      // transfer() doesn't carry Cache-Control. Copy it from the source
      // object onto each upload. Assumes source and destination keys match.
      name: "carry-cache-control",
      wrap: handlers({
        upload: async (op, next) => {
          const cacheControl = await cacheControlOf(
            source,
            SOURCE_BUCKET,
            op.key
          );
          return next(
            cacheControl
              ? { ...op, options: { ...op.options, cacheControl } }
              : op
          );
        },
      }),
    },
  ],
});

Why the plugin. transfer() carries the body, content type, and user metadata. It doesn’t carry Cache-Control, because FileInfo has no field for it; neither head() nor download() returns it. If your objects never set Cache-Control, delete the plugin. If they do, the plugin wraps every upload on the destination, reads the source object’s Cache-Control with one HeadObject through files.raw, and passes it as the upload’s cacheControl option. That’s one extra request per key. If you remap keys with transformKey, map op.key back to the source key before the lookup.

Retries. retries: 3 covers listing and opening each download. The plugin’s HeadObject goes through the native client, which applies the AWS SDK’s own retries. The upload leg is a stream, and Files SDK never replays a stream, so a failed upload lands in the failure manifest instead of retrying.

Copy with a failure manifest

import { mkdir, writeFile } from "node:fs/promises";

import { transfer } from "files-sdk";

import { dest, source } from "./storage";

const result = await transfer(source, dest, {
  concurrency: 16,
  onProgress: ({ done, total, key, status }) => {
    if (status === "failed") {
      console.log(`fail  ${key}`);
    }
    if (done % 1000 === 0 || done === total) {
      console.log(`${done}/${total}`);
    }
  },
});

const failed = (result.errors ?? []).map(({ key, error }) => ({
  key,
  code: error.code,
  message: error.message,
}));

await mkdir("migration", { recursive: true });
await writeFile("migration/failed.json", JSON.stringify(failed, null, 2));
console.log(
  `${result.transferred.length} transferred, ${failed.length} failed`
);

transfer() doesn’t throw when some keys fail. Successes come back in transferred and failures in errors, each with its key and a normalized FilesError. onProgress fires once per key, failures included (with status: "failed"), so the last progress line reads total/total even when keys fail. The event carries only the key; the error itself is in errors, which is what the script writes to the manifest.

concurrency is how many objects stream at once. A large object’s upload holds up to four 5 MiB parts in memory (lib-storage’s defaults, which Files SDK keeps), so 16 large objects in flight can hold around 320 MiB. Lower it on a small machine.

Retry only the failed keys

import { readFile, writeFile } from "node:fs/promises";

import { FilesError } from "files-sdk";

import { dest, source } from "./storage";

interface Failure {
  key: string;
  code: string;
  message: string;
}

const failed: Failure[] = JSON.parse(
  await readFile("migration/failed.json", "utf8")
);
const stillFailing: Failure[] = [];

for (const { key } of failed) {
  try {
    // The same three things transfer() carries: body, type, metadata.
    const file = await source.download(key, { as: "stream" });
    await dest.upload(key, file.stream(), {
      contentType: file.contentType,
      ...(file.metadata && { metadata: file.metadata }),
    });
    console.log(`ok    ${key}`);
  } catch (error) {
    const err = FilesError.wrap(error);
    // NotFound here means the key was deleted from the source after the walk.
    stillFailing.push({ key, code: err.code, message: err.message });
    console.log(`fail  ${key} (${err.code})`);
  }
}

await writeFile("migration/failed.json", JSON.stringify(stillFailing, null, 2));
console.log(
  `${failed.length - stillFailing.length} recovered, ${stillFailing.length} still failing`
);

The retry goes through the same dest instance, so the Cache-Control plugin applies to it too. Run it until the manifest is empty or holds only keys you’ve decided about. A NotFound that persists means the object no longer exists at the source, which is a question for the app, not the script. Re-running transfer() with overwrite: false also fills the gaps, but it lists the whole bucket and sends one exists() per key to get there.

Verify every object

import { createHash } from "node:crypto";
import { writeFile } from "node:fs/promises";

import type { FileInfo, Files } from "files-sdk";

import {
  DEST_BUCKET,
  SOURCE_BUCKET,
  cacheControlOf,
  dest,
  source,
} from "./storage";

const listKeys = async (files: Files): Promise<Map<string, FileInfo>> => {
  const all = new Map<string, FileInfo>();
  for await (const file of files.listAll()) {
    all.set(file.key, file);
  }
  return all;
};

// Stream the object and hash it. ETags can't be compared across providers.
const inspect = async (files: Files, bucket: string, key: string) => {
  const file = await files.download(key, { as: "stream" });
  const hash = createHash("sha256");
  for await (const chunk of file.stream()) {
    hash.update(chunk);
  }
  return {
    size: file.size,
    contentType: file.contentType,
    metadata: JSON.stringify(Object.entries(file.metadata ?? {}).sort()),
    cacheControl: await cacheControlOf(files, bucket, key),
    sha256: hash.digest("hex"),
  };
};

const [from, to] = await Promise.all([listKeys(source), listKeys(dest)]);
const missing = [...from.keys()].filter((key) => !to.has(key));
const extra = [...to.keys()].filter((key) => !from.has(key));
const common = [...from.keys()].filter((key) => to.has(key));

const mismatches: {
  key: string;
  field: string;
  source: unknown;
  dest: unknown;
}[] = [];
let next = 0;
const worker = async () => {
  while (next < common.length) {
    const key = common[next++]!;
    const [a, b] = await Promise.all([
      inspect(source, SOURCE_BUCKET, key),
      inspect(dest, DEST_BUCKET, key),
    ]);
    for (const field of Object.keys(a) as (keyof typeof a)[]) {
      if (a[field] !== b[field]) {
        mismatches.push({ key, field, source: a[field], dest: b[field] });
      }
    }
  }
};
await Promise.all(Array.from({ length: 8 }, worker));

const report = { checked: common.length, missing, extra, mismatches };
await writeFile("migration/verify.json", JSON.stringify(report, null, 2));
console.log(
  `${common.length} checked, ${missing.length} missing, ${extra.length} extra, ${mismatches.length} mismatches`
);
process.exitCode = missing.length || mismatches.length ? 1 : 0;

What it compares, and why each check reads where it does:

  • Key sets come from listAll() on both sides: missing is in S3 but not R2, extra the reverse.
  • Size, type, and metadata come from download(), not list(). S3’s ListObjectsV2 returns no content type or metadata, so a listed FileInfo guesses its contentType from the key’s extension and has no metadata at all.
  • SHA-256 is computed from the bytes on both sides. ETags aren’t compared. A multipart object’s ETag depends on how it was split into parts, and R2 documents its own multipart ETag format, so identical bytes can carry different ETags.

The cost is real: the script downloads every object from both buckets once more. AWS bills that as S3 data transfer out, and both providers count the reads as requests. For a large bucket, compare sizes, types, and metadata everywhere and hash a random sample, or hash only the keys that matter most.

sync() with dryRun: true is a cheap drift check that lists both sides without downloading. Use compare: "size": the default "etag" comparison flags multipart objects as changed even when their bytes match.

const plan = await sync(source, dest, { dryRun: true, compare: "size" });
// plan.uploaded lists keys that are new or a different size in S3.

What a local run showed

The scripts above ran against two buckets on a local MinIO server, with r2() pointed at MinIO through its endpoint option as a stand-in for R2. Nothing here says how R2 itself behaves. The fixture was 206 objects: text and JSON with user metadata, an image with Cache-Control: public, max-age=31536000, immutable, an empty file, a key with spaces and accents, a 20 MiB object uploaded in 8 MiB parts, and 200 small files. Two test-only plugins, not part of the scripts above, made one destination upload fail and deleted one source key after the listing.

  • migrate.ts reported 204 transferred and 2 failed: the injected failure as Provider, the deleted key as NotFound.
  • verify.ts then found 1 missing key and no mismatches across the other 204.
  • retry.ts recovered the injected failure and kept the deleted key as NotFound. A second verify.ts found 205 checked, nothing missing, and no mismatches.
  • The 20 MiB object’s ETag ended in -3 in the source and -4 in the destination, because the destination upload used lib-storage’s 5 MiB parts. Both copies hashed to the same SHA-256.
  • A plain transfer() without the plugin left the image with no Cache-Control. With the plugin, it arrived intact.
  • sync() dry runs reported 0 keys to upload with compare: "size", and 1 (the 20 MiB object) with the default "etag".

Cut over, and keep a way back

  1. Freeze writes as above, and note the time.
  2. Copy, retry, verify until verify.ts exits with code 0, or until every remaining entry is one you’ve accepted.
  3. Check for drift right before switching, with the sync() dry run. It should plan no uploads.
  4. Switch the app to R2 and lift the freeze.
  5. Leave the S3 bucket alone for a rollback window. Signed S3 URLs the app issued before the switch keep working until they expire, and any absolute S3 URLs stored in your database need rewriting before the bucket goes away.
  6. Delete the S3 bucket only after the window closes.

A flag in the app’s storage module turns the freeze, the switch, and the switch back into environment changes:

import { Files } from "files-sdk";
import { r2 } from "files-sdk/r2";
import { s3 } from "files-sdk/s3";

// STORAGE=r2 after the cutover; STORAGE_READONLY=1 during the freeze.
export const files = new Files({
  adapter:
    process.env.STORAGE === "r2"
      ? r2({ bucket: "uploads" })
      : s3({ bucket: "uploads", region: "us-east-1" }),
  readonly: process.env.STORAGE_READONLY === "1",
});

To roll back, writes made to R2 after the cutover have to go back to S3 first. Run this with the cutover time, check the dry run, then run it for real and switch the flag back:

import { sync } from "files-sdk";

import { dest, source } from "./storage";

// When the app switched to R2, as an ISO timestamp.
const cutoverAt = Date.parse(process.env.CUTOVER_AT ?? "");
if (Number.isNaN(cutoverAt)) {
  throw new Error(
    "Set CUTOVER_AT to the cutover time, e.g. 2026-10-08T22:00:00Z"
  );
}

// Copy what the app wrote to R2 after the cutover back to S3. Objects R2
// hasn't touched since the cutover are skipped.
const plan = await sync(dest, source, {
  dryRun: process.argv.includes("--dry-run"),
  compare: (r2File, s3File) =>
    r2File.size === s3File.size && (r2File.lastModified ?? 0) < cutoverAt,
});
console.log(`upload ${plan.uploaded.length}, skip ${plan.skipped.length}`);
for (const { key, error } of plan.errors ?? []) {
  console.log(`fail  ${key} (${error.code})`);
}

The custom compare skips an object only when it’s the same size and R2 last modified it before the cutover, so a same-size edit made on R2 afterwards still goes back. In the local run, after one new key and one same-size edit on the R2 stand-in, the dry run planned exactly those two uploads and skipped the other 204. This direction doesn’t carry Cache-Control, and it doesn’t delete anything: objects deleted on R2 after the cutover stay in S3 unless you add prune: true, which deletes every S3 key that R2 doesn’t have. Dry-run that before you trust it.

Troubleshooting

Multipart, progress, and unknown-length stream uploads on S3 require the optional peer dependency '@aws-sdk/lib-storage'. Every transfer fails with this when lib-storage isn’t installed, because each body arrives as a stream. Install @aws-sdk/lib-storage.

Many NotFound entries in failed.json. Objects were deleted from the source after the listing, so writes weren’t frozen. Freeze them, then decide whether those keys should exist at all. retry.ts keeps reporting them until they do.

Expected an adapter backed by @aws-sdk/client-s3. cacheControlOf ran against an R2 instance on the fetch engine, or before the lazy S3Client loaded. Keep client: "aws-sdk" on the destination and call another method on it first, as verify.ts does with download().

cacheControl mismatches on every key that has one. The destination was written without the plugin, by an earlier run or by another tool. Re-upload those keys through the dest instance, for example by listing them in failed.json and running retry.ts.

Last updated on

Was this page helpful?