Skip to content
TaeyoungKim.dev

EBS Snapshot Consistency: Why Recent Writes May Be Missing

CloudWritten 3 min readTaeyoungKim
LinkedInX

It is easy to assume an EBS snapshot safely contains every write made at the moment you create it. But if an application or operating system still holds a write in memory, the point recorded on the volume can differ from the point the application considers complete. A snapshot is useful for recovery, yet creating one alone does not guarantee application-consistent data.

Which data does an EBS snapshot contain?

The diagram separates buffered writes from volume data. It also shows why coordinating writes before the snapshot and checking restored data afterward are different steps.

An EBS snapshot preserves a point-in-time view of a volume. Imagine an order write beginning at 10:00 while part of the data is still buffered as the app prepares its success response. You cannot assume a snapshot includes that buffered part. The time and order are illustrative, not a report of a real system.

Point in the flowApplication sideEBS volume side
Just after a write requestChanges may remain in a bufferThey may not yet be recorded
After coordinating writesRequired changes are sent toward storageThe volume is closer to the intended snapshot state
After restorationThe app's data structures need checkingThe mere presence of files is not proof of recovery

AWS's EBS snapshot guidance explains that application- or OS-cached data may be absent and recommends pausing writes to improve consistency and completeness. Learning that a snapshot copies a volume at a point in time is only the beginning; application consistency needs its own procedure.

What should a service with ongoing writes prepare?

A read-only file volume and an actively written database volume need different consistency guarantees. For a database, check its supported backup and checkpoint procedure and the recovery point it can provide. Decide whether to pause writes, flush buffers, or freeze a filesystem according to the application and operating procedure. Do not run an untested freeze command on a production server. If you freeze a filesystem, plan how to thaw it, including when a step fails.

If one dataset spans several volumes, snapshots taken independently at arbitrary times may not form one consistent application state. A snapshot marked “successful” and a database that opens correctly after restoration are different checks.

What should a restore test verify?

Create a new volume from the snapshot in an isolated environment. Check whether the application can read the data, not just whether files exist. For order data, compare expected counts and key records at the reference point and inspect the state of writes that were in progress. If a restore exercise uses real customer data, determine access and privacy handling first.

Choose retention, cost, and the recovery point you need together. Many old snapshots do not help if none can restore the required state. This article does not claim a test in a particular AWS account or a successful production recovery.

Key takeaways: snapshot success is not recovery success

An EBS snapshot captures volume state, but it does not automatically include writes still held in application or OS buffers. Coordinate writes for active services, then restore into an isolated environment and verify that the application can read the data.

Author

TaeyoungKim

Connecting technical foundations with implementation, verification, and production decisions.

#AWS#EBS#snapshots#backup#data consistency

Read next