The problem with full snapshots
A full snapshot copies the whole disk. On an 80 GB volume that is 80 GB of reads, 80 GB of writes and 80 GB of storage, every time — and on a typical server, well over 95% of those blocks are byte-for-byte identical to the previous snapshot.
Do that daily and you are spending most of your I/O budget and most of your storage on copies of an operating system that has not changed since install.
Tracking what changed
The fix is to know which blocks were written since last time. QEMU maintains a dirty bitmap: one bit per block, set on write, cleared when a snapshot consumes it.
# disk as blocks; 1 = written since last snapshot
block: 0 1 2 3 4 5 6 7 8 9 ...
bitmap: 0 0 1 0 0 1 1 0 0 0
# the snapshot copies blocks 2, 5 and 6. Nothing else.
An 80 GB disk with 400 MB of daily change copies 400 MB rather than 80 GB. The snapshot takes seconds instead of minutes, and the storage cost of retention collapses.
The bitmap lives with the volume and survives a reboot. It does not survive a live migration to another host, which is why a migration is followed by one full snapshot to re-establish the baseline.
Consistency: the part that matters
Copying a disk that is being written to gives you a torn image — some blocks from before your write, some from after. For a database that is a corrupt file rather than an old one.
So the instance is paused for a few milliseconds while a consistent point is established, then resumed. Copying happens afterwards, in the background, against that frozen point. The pause is short enough that a TCP connection does not notice, and it is what makes the snapshot a coherent moment rather than a smear across several seconds.
This gives crash consistency — equivalent to pulling the power. Journalled filesystems and modern databases recover from that cleanly, because it is the case they are built for. What it is not is application consistency: a database mid-transaction will roll that transaction back on restore, correctly. If you need the transaction, quiesce the application before the snapshot. For most workloads, crash consistency is genuinely enough.
The new problem: chains
Incrementals depend on what came before. A week of daily snapshots is a chain:
full ← inc¹ ← inc² ← inc³ ← inc⁴ ← inc⁵ ← inc⁶
Two consequences, and the second is the one people miss.
Restoring the newest point means reading the whole chain. A long chain restores slowly, and restore speed is the number that matters — it is the one you find out about during an incident.
A corrupt link breaks everything after it. Lose inc³ and points 4, 5 and 6 are unrecoverable, because each is defined in terms of its predecessor.
Which is why chains get collapsed
Both problems have the same answer: merge periodically. A background job folds older incrementals into the base, so the chain never grows past a fixed length:
# before
full ← inc¹ ← inc² ← inc³ ← inc⁴ ← inc⁵ ← inc⁶
# after merging the oldest three into the base
full' ← inc⁴ ← inc⁵ ← inc⁶
Restore time stays bounded, and the window in which a single corrupt link can hurt you stays short. The merge runs off-peak and does not touch the running instance.
Every restore point is also verified on write — checksummed, and the chain walked to confirm it resolves. A backup you have never tested is a hypothesis, and discovering that during an incident is the worst possible time.
Snapshots are not backups
Worth stating plainly, because the words get used interchangeably and the difference is the whole thing.
A snapshot lives on the same storage as the volume. It protects you from a bad deploy, a dropped table, a plugin update that ate the site. It does not protect you from that storage array failing, because it fails with it.
A backup lives somewhere else. It protects you from the storage failing, the data centre burning down, and the account being compromised.
You want both. We take snapshots on every plan and sell off-server backup as an add-on, and we will keep saying — in the terms, on the phone, and here — that you should keep your own independent copy of anything you cannot afford to lose. That is not a disclaimer to cover ourselves. It is what we would tell a friend.
Using them
Take one before anything risky. It costs seconds and it turns "I hope this works" into "I can undo this". Snapshots are included on every VPS and dedicated server, and if a restore ever does not behave, +91 75994 50220 is answered at any hour.