RE:NODE

Operations11 min read

S3 storage sizing and the 3-2-1 backup rule

How much S3 storage you need for backups and uploads, what one copy versus replicated storage protects against, and building a 3-2-1 setup that restores.

0 readers

Size S3 storage from three numbers: how much data you protect, how many past versions you keep, and how well your backup tool avoids storing the same bytes twice. A 5 GB website kept as fourteen daily full copies needs 70 GB; the same history in a deduplicating tool like restic often fits in 10 to 15 GB. Then decide where the bucket sits in the 3-2-1 rule - three copies of the data, on two different kinds of storage, one of them somewhere else. A single-copy bucket in one location is a good second or third copy and a poor only copy, and replicated storage is not the same thing as a backup either: replication faithfully copies your mistakes. This post works through both questions with real numbers, so you buy the right tier and know exactly what it does and does not protect.

The 3-2-1 rule, and what each number is for#

The rule is older than cloud storage and has survived because each number answers a different failure.

NumberMeansProtects against
3 copiesThe live data plus two backupsAny single copy being lost or corrupt
2 kinds of storageNot all copies on the same type of systemOne system-wide fault taking out every copy
1 off-siteAt least one copy in a different placeFire, theft, a provider problem, losing an account

The common modern extension is 3-2-1-1-0: one copy offline or immutable, so that someone with your credentials cannot delete it, and zero errors when you verify a restore. The last digit is the one people skip and the one that matters most. Backups that actually restore makes the case at length.

What counts as "a different kind of storage" is looser today than when the rule meant tape versus disk. The point is independence: a panel backup and an S3 bucket at a different service are different systems with different credentials and different failure modes, and that is what the "2" is reaching for. Two folders on the same server are not two kinds of storage. Two buckets under the same access key are barely two copies.

One copy, replicated, and what each protects against#

Storage providers describe durability in different ways, and it is worth being precise about what is being promised.

Replicated storage keeps several copies of every object, on different disks, machines, or sites. Large providers quote durability figures with many nines, which means: the chance that the hardware loses an object you stored is extremely small. It is a real and valuable property.

Single-copy storage keeps your object once. It still sits on redundant disks in most setups, but there is one system, in one place, holding it.

Now look at what actually destroys backups in practice:

Cause of lossReplication helps?A separate copy helps?
A disk or server failsYesYes
The whole site has an outage or disasterOnly multi-site replicationYes, if elsewhere
You delete the wrong prefixNo - the delete is replicatedYes
A sync job copies a corrupted file over a good oneNoYes, if it has history
Ransomware or an intruder with your keysNoOnly if the keys differ
The account is suspended or unpaidNoYes, if a different account

The top two rows are what replication is for. The bottom four are, in most post-mortems, what actually happened. Replication protects the provider's hardware from failing you; it does nothing about you, your scripts or your credentials failing you. That is why a single-copy bucket can be a perfectly good part of a backup plan, and a replicated bucket on its own is not a backup plan.

RE:NODE's S3 storage is the single-copy kind, and says so: one copy of your data, on NVMe, in one location in Germany, not replicated. The storage plans also carry no panel backup slots - the bucket is the backup, not something that is itself backed up. That makes it a good place for a second copy of your backups and for application uploads that also exist somewhere else. It should not be the only place anything lives.

What belongs in object storage#

Object storage is good at holding files that are written once and read many times, over HTTP, from anywhere. It is bad at files that change a little at a time.

Good fits:

  • Backup copies - restic repositories, database dumps, archived game worlds, server snapshots exported as files.
  • User uploads and media - avatars, attachments, product images, generated PDFs, as long as the application or a backup job also keeps them elsewhere.
  • Build and release artefacts - versioned files you want to fetch from several servers.
  • Log archives - compressed daily logs moved off the server that wrote them.
  • Data exports and handoffs - files you share with a presigned link instead of an email attachment.

Poor fits:

  • A live database. Databases rewrite small blocks constantly; S3 replaces whole objects. Store dumps, not data directories.
  • A mounted "drive" for an application that writes heavily. It works, slowly, with lots of traffic and odd consistency.
  • The only copy of anything. That rule applies to every storage system, but bears repeating for a single-copy one.

How much space: the arithmetic#

Three inputs decide the size:

  1. Source size - what you protect, after compression. Databases compress well (often 5-10x for text-heavy dumps); images, video and game worlds hardly at all.
  2. Retention - how many restore points you keep.
  3. Change rate and deduplication - how much of each restore point is new data, and whether the tool stores unchanged data again.
MethodSpace usedExample: 5 GB source, 14 daily points
Full copy every timesource x points70 GB
One current copy plus changed filessource + changes x days5 GB + 14 x 0.3 GB = about 9 GB
Deduplicating tool (restic, borg-style)source + unique changesOften 6-10 GB
Single current sync, no historysource5 GB, but no protection against bad syncs

The last row is a trap in disguise: a sync with no history is a mirror, and a mirror is a copy of whatever went wrong. Do not count it as a backup.

Add headroom on top of the result. Two things consume space unexpectedly: abandoned multipart uploads (a large upload killed halfway leaves its parts behind until someone aborts them) and the growth of the source itself. A sensible rule is to buy roughly 1.5 times what the arithmetic says, and to look at actual usage monthly.

Retention: how far back you need to go#

Retention is the multiplier in every sizing sum, so it deserves a decision rather than a default. The question is not "how many backups feel safe" but "how long could a problem go unnoticed".

Some problems announce themselves: the server will not start, the site shows an error, players complain within the hour. A day or two of dailies covers those. Others are quiet. A plugin that corrupts one table, a cron job that has been deleting the wrong folder, a griefer who hid their damage in a corner of the map nobody visits, an intruder who changed a file and waited. These are found days or weeks later, and the only backup that helps is one from before the problem began. If your oldest restore point is seven days old and the damage is ten days old, every backup you own contains it.

That is why tiered retention - grandfather, father, son - is the usual answer. Keep many recent points close together, and fewer older ones further apart:

TierKeepCovers
Daily7Loud problems, noticed quickly
Weekly4-5Problems found within a month
Monthly6-12Slow corruption, audits, "what did this look like in spring"

Twenty-odd restore points reach back a year, and with a deduplicating tool the older ones cost very little because most of their data is shared with the newer ones. With full copies, the monthly tier is where the space goes, so price it before you promise yourself a year of history.

Worked examples#

A Minecraft server with a 3 GB world. Worlds barely compress. Fourteen dated full copies would take 42 GB; restic, where most region files do not change from night to night, typically stores the first copy plus a few hundred megabytes per day - call it 6-9 GB for two weeks of dailies and a few monthlies. The 10 GB tier is enough with restic; the 50 GB tier for dated full copies. Game server backups to S3 builds the job.

A web app with a 2 GB PostgreSQL database and 20 GB of uploads. A custom-format dump compresses to perhaps 400 MB. Keeping 7 daily, 5 weekly and 12 monthly dumps is 24 files, about 10 GB. Uploads, synced with history via restic, are 20 GB plus their growth. Together around 35 GB today, so the 50 GB tier with room to grow. Database dumps to S3 on a schedule has the retention script.

A small office with 60 GB of shared documents. Office files compress moderately and change slowly. restic with 30 dailies and 12 monthlies might need 75-90 GB. That is the 100 GB tier - and for a business, the third copy at home or at a second provider is not optional.

RE:NODE's storage tiers run from 10 GB to 100 GB, starting from $1.5 a month. Moving up a tier changes the limit on the server you already have rather than rebuilding it, so starting small and growing is a reasonable strategy, as long as you watch usage before you hit the ceiling rather than after.

Checking what you actually use#

Measure the bucket, not your estimate. Two commands give the total:

bash
$ aws s3 ls s3://backups --recursive --summarize --human-readable --profile store | tail -n 2$ rclone size store:backups

Both list every object, so on a bucket with millions of keys they take a while; for normal backup buckets they finish in seconds. Look also for forgotten multipart uploads, which do not appear in a normal listing but do use space:

bash
$ aws s3api list-multipart-uploads --bucket backups --profile store$ aws s3api abort-multipart-upload --bucket backups --key big.tar --upload-id <id> --profile store

For restic repositories, restic stats --mode raw-data reports how much space the repository really uses after deduplication, which is the number to compare with your tier. And if usage grows faster than your data, the usual cause is a retention job that stopped running - check that forget --prune or your delete script is still executing.

Building a 3-2-1 setup with a single-copy bucket#

A single-copy S3 bucket fits naturally as one of the three. Some layouts that satisfy the rule:

scheduled backupnightly jobmonthly copyApp and databasecopy 1, livePanel backupscopy 2, off the machineS3 bucketcopy 3, restic + dumpsHome diskmonthly download
A 3-2-1 layout for a small web app

Each arrow is a separate job with its own schedule, and each box can be lost without taking the others with it. Reading the layout box by box:

  • Copy 1 is the live data on the server.
  • Copy 2 is the panel's backups, stored off the machine they protect and restorable with one button - the fastest recovery for everyday mistakes.
  • Copy 3 is the S3 bucket, filled by a job with its own credentials and its own retention. It survives the server being deleted, which takes the panel backups with it.
  • The off-site copy is what the bucket alone cannot give you when it is in the same location as the server. A monthly download to a disk at home, or a second bucket at another provider filled by rclone sync from the first, closes that gap. For a hobby project, a monthly download is plenty; for a business, automate it.

Keep the credentials separate. The application should not hold the keys that can delete its own backups. If the server running the backup job is compromised, the attacker gets whatever keys it holds, so a copy that the job cannot delete - the download at home - is your offline "1" in 3-2-1-1-0.

Testing that it restores#

Every layer of this setup is worth nothing until a restore from it has worked. Restore quarterly from each copy, not only the most convenient one:

  1. From the bucket: download a dump or restic restore latest --target /tmp/restore-test and load it into a scratch database or test server.
  2. From the off-site copy: open the file from the home disk and check it is the date you expect.
  3. Write down how long each took. Downloading 50 GB takes real time even on a fast line, and that time is part of your outage.

Testing a restore before you need it turns this into a checklist.

FAQ#

Is a single-copy bucket safe for backups?

As one copy among several, yes. Its risk is concentrated in one system in one place, which is exactly what the other copies in 3-2-1 are for. As the only backup, no - but that is true of any single storage system, replicated or not.

Does replication replace backups?

No. Replication protects against hardware failure and copies every delete, overwrite and corruption instantly. Backups protect against mistakes because they keep the past. You want history, not just redundancy.

How many days of backups should I keep?

At least as long as it might take you to notice a problem. Silent corruption or a deleted folder nobody looks at can go unnoticed for weeks, so a week of dailies plus a few monthlies is a sensible minimum for most sites and servers.

Do backups count against my storage if I delete them?

Not once they are deleted - storage without versioning frees the space immediately. Abandoned multipart uploads are the exception: they use space until aborted, and do not show in a normal listing.

Should I encrypt backups in S3?

Yes, for anything containing personal data or credentials. restic encrypts by default; for dumps, encrypt with a tool like age before upload, and keep the decryption key somewhere that does not depend on the server being alive.


Comments

Completely anonymous: no account, no email, no cookie. We store the name you type, the text and the time - nothing else. Links are limited and markup is not rendered.

0/2000