infrastructure//storage//backup
A backup is a copy of data kept separately from the original so that the data can be restored after it is lost, corrupted or deleted: a failed disk, a bad migration, ransomware, an operator's mistaken `DELETE`. Every system that holds data someone cannot recreate (a customer database, a plant historian, a project's repository) needs one, and needs to have tried restoring it.
A backup is a copy of data kept separately from the original so that the data can be restored after it is lost, corrupted or deleted: a failed disk, a bad migration, ransomware, an operator's mistaken DELETE. Every system that holds data someone cannot recreate (a customer database, a plant historian, a project's repository) needs one, and needs to have tried restoring it.
Two numbers define what a backup is worth. How much recent data may be lost (taken every 24 hours, a crash at 23:00 loses a day of orders) and how long the restore takes (minutes from a hot copy, a day from tape). Both are business decisions before they are technical ones.
A snapshot is a point-in-time image of a disk or database, usually taken in seconds by the storage system itself. It is excellent for rolling back a mistake made an hour ago, and weak as a backup when it lives on the same storage, account or provider as the original: whatever destroys the disk, or deletes the account, takes the snapshots with it.
The classic rule is 3-2-1: three copies of the data, on two different kinds of media, one of them off-site. Its modern addition is one copy that cannot be altered or deleted for a while (immutable storage), because ransomware and compromised credentials go after the backups first.
Point-in-time recovery replays a database's transaction log on top of the last full copy, so it can be restored to any second (just before the faulty migration ran), at the cost of keeping the log.
A backup that has never been restored is a hope. Formats change, keys get lost, a restore takes ten times longer than expected; a periodic restore test into a scratch environment is the only proof.
Redundancy protects against hardware failing; backups protect against the data itself going wrong.
A mirrored disk faithfully mirrors the accidental deletion.
?Is a provider's automatic snapshot enough as a backup?
Only against small mistakes. A snapshot on the same platform shares its failure modes (the account, the region, the billing, the credentials), so a copy that survives the loss of all of them has to live somewhere else, under other credentials.
Cloud databases include automated backups and snapshots (managed database), which cover the first layer; the off-site, independent copy remains the owner's job.