Talk to a real engineer in Agra, 24×7 — +91 75994 50220 support@bigdomainhost.com
Product

How incremental snapshots work

A full snapshot of an 80 GB disk copies 80 GB every time, most of it identical to last time. Incremental snapshots copy what changed — and introduce one new problem worth understanding before you rely on them.

The problem with full snapshots

A full snapshot copies the whole disk. On an 80 GB volume that is 80 GB of reads, 80 GB of writes and 80 GB of storage, every time — and on a typical server, well over 95% of those blocks are byte-for-byte identical to the previous snapshot.

Do that daily and you are spending most of your I/O budget and most of your storage on copies of an operating system that has not changed since install.

Tracking what changed

The fix is to know which blocks were written since last time. QEMU maintains a dirty bitmap: one bit per block, set on write, cleared when a snapshot consumes it.

# disk as blocks; 1 = written since last snapshot
block:   0  1  2  3  4  5  6  7  8  9 ...
bitmap:  0  0  1  0  0  1  1  0  0  0

# the snapshot copies blocks 2, 5 and 6. Nothing else.

An 80 GB disk with 400 MB of daily change copies 400 MB rather than 80 GB. The snapshot takes seconds instead of minutes, and the storage cost of retention collapses.

The bitmap lives with the volume and survives a reboot. It does not survive a live migration to another host, which is why a migration is followed by one full snapshot to re-establish the baseline.

Consistency: the part that matters

Copying a disk that is being written to gives you a torn image — some blocks from before your write, some from after. For a database that is a corrupt file rather than an old one.

So the instance is paused for a few milliseconds while a consistent point is established, then resumed. Copying happens afterwards, in the background, against that frozen point. The pause is short enough that a TCP connection does not notice, and it is what makes the snapshot a coherent moment rather than a smear across several seconds.

This gives crash consistency — equivalent to pulling the power. Journalled filesystems and modern databases recover from that cleanly, because it is the case they are built for. What it is not is application consistency: a database mid-transaction will roll that transaction back on restore, correctly. If you need the transaction, quiesce the application before the snapshot. For most workloads, crash consistency is genuinely enough.

The new problem: chains

Incrementals depend on what came before. A week of daily snapshots is a chain:

full ← inc¹ ← inc² ← inc³ ← inc⁴ ← inc⁵ ← inc⁶

Two consequences, and the second is the one people miss.

Restoring the newest point means reading the whole chain. A long chain restores slowly, and restore speed is the number that matters — it is the one you find out about during an incident.

A corrupt link breaks everything after it. Lose inc³ and points 4, 5 and 6 are unrecoverable, because each is defined in terms of its predecessor.

Which is why chains get collapsed

Both problems have the same answer: merge periodically. A background job folds older incrementals into the base, so the chain never grows past a fixed length:

# before
full ← inc¹ ← inc² ← inc³ ← inc⁴ ← inc⁵ ← inc⁶

# after merging the oldest three into the base
full' ← inc⁴ ← inc⁵ ← inc⁶

Restore time stays bounded, and the window in which a single corrupt link can hurt you stays short. The merge runs off-peak and does not touch the running instance.

Every restore point is also verified on write — checksummed, and the chain walked to confirm it resolves. A backup you have never tested is a hypothesis, and discovering that during an incident is the worst possible time.

Snapshots are not backups

Worth stating plainly, because the words get used interchangeably and the difference is the whole thing.

A snapshot lives on the same storage as the volume. It protects you from a bad deploy, a dropped table, a plugin update that ate the site. It does not protect you from that storage array failing, because it fails with it.

A backup lives somewhere else. It protects you from the storage failing, the data centre burning down, and the account being compromised.

You want both. We take snapshots on every plan and sell off-server backup as an add-on, and we will keep saying — in the terms, on the phone, and here — that you should keep your own independent copy of anything you cannot afford to lose. That is not a disclaimer to cover ourselves. It is what we would tell a friend.

Using them

Take one before anything risky. It costs seconds and it turns "I hope this works" into "I can undo this". Snapshots are included on every VPS and dedicated server, and if a restore ever does not behave, +91 75994 50220 is answered at any hour.

Answers

Frequently asked questions

Still unsure? Call +91 75994 50220 or write to support@bigdomainhost.com — a human replies, 24×7.

What is the difference between a snapshot and a backup?

A snapshot usually lives on the same storage as the volume it came from, so it protects against your mistakes but not against that storage failing. A backup lives somewhere else. You want both, and conflating them is how people discover they had neither.

How long does an incremental snapshot take?

Seconds, because it copies only the blocks that changed since the last one. A first, full snapshot of an 80 GB disk takes minutes; the incrementals after it are typically a few seconds.

Does taking a snapshot slow my server down?

Briefly, and by very little. The instance pauses for a few milliseconds to establish a consistent point, then copying happens in the background while the server runs normally.

How many snapshots should I keep?

Enough to cover the time it takes you to notice a problem. Seven daily points suits most people; if a corrupted database might go unnoticed for a fortnight, you need longer retention, not more frequent snapshots.

Support that picks up the phone.

24×7, from our office in Agra, in IST — Hindi or English. Sales, migration and emergencies all reach the same engineers. No offshore queue, no 48-hour first reply.

Questions about any of this?

Call +91 75994 50220. The people who wrote this are the people who answer.