Asia/Kolkata
Posts

Instance Store vs EBS vs EFS: Understanding AWS Storage for EC2

A practical breakdown of EC2 storage options — Instance Store, EBS, and EFS — what actually differs between them and when to reach for each one.
January 20, 2023
Share
Instance Store vs EBS vs EFS: Understanding AWS Storage for EC2
On this page
When you launch an EC2 instance, you immediately have to think about storage. AWS gives you three fundamentally different options, and picking the wrong one is the kind of mistake that shows up later as data loss or a painful migration. Here's how they actually work.
Instance Store is physical storage attached directly to the host machine running your EC2 instance — NVMe SSDs in most modern instance types. No network hop, so the throughput and IOPS numbers are significantly higher than anything network-attached. The trade-off is absolute: when the instance stops, the data is gone. Not soft-deleted, not recoverable — gone. This includes host hardware failures. A few things people miss:
  • Not all instance types have Instance Store. You need to specifically pick families that include it — look for a d in the name: m5d, c5d, i3, i4i
  • The volumes show up as raw block devices on boot. You format and mount them yourself
  • You can't snapshot them
Good for: Kafka broker disk buffers, Redis AOF on nodes where you're already replicating, Spark shuffle data, any cache where losing the node means rebuilding the cache anyway. Never use it as primary storage for anything you can't recreate.
Amazon EBS is the standard persistent storage you get with EC2. It persists independently of the instance lifecycle — if the instance dies, you detach the volume and remount it elsewhere. The AZ trap: EBS volumes are locked to a single Availability Zone. A volume in us-east-1a cannot be attached to an instance in us-east-1b. For multi-AZ architectures you either replicate at the application level or go through snapshot → restore to move data between zones. This catches people off guard when they design for high availability without accounting for it. Always use gp3 over gp2. gp2 ties IOPS to volume size at 3 IOPS per GB, which pushes teams to overprovision disk size just to get more IOPS. gp3 lets you configure storage, IOPS, and throughput independently — and it's cheaper:
Bash
aws ec2 create-volume \
  --availability-zone us-east-1a \
  --volume-type gp3 \
  --size 100 \
  --iops 6000 \
  --throughput 250
io2 Block Express is for serious database workloads — up to 256,000 IOPS, 4,000 MB/s throughput, sub-millisecond latency. If you're running RDS with Provisioned IOPS, this is what's running under the hood. Multi-Attach exists (io1/io2 only, same AZ, cluster-aware filesystem required) but it's narrow enough that it's rarely the right tool for shared storage.
Volume TypeMax IOPSWhen to use
gp316,000Default for almost everything
io2 Block Express256,000High-performance databases
st1500Sequential reads — Kafka, logs, Hadoop
sc1250Cold data, infrequent access
Amazon EFS is managed NFS. The core difference from EBS: multiple instances can mount it at the same time. EFS latency is higher than EBS — you're going over the network to a managed service. It's not designed for high-IOPS or latency-sensitive workloads. Where it actually makes sense:
  • Multiple EC2 instances or containers need to read and write the same files
  • You don't want to think about capacity (EFS scales automatically, you pay per GB used)
  • Kubernetes workloads that need ReadWriteMany PVCs
That last point matters in EKS. EBS only supports ReadWriteOnce — one node at a time. If you have pods across different nodes that need shared writable storage (shared ML model store, CMS uploads directory), EFS via the EFS CSI driver is your managed option in AWS. On throughput modes: the default Bursting mode scales throughput with filesystem size, but small file systems drain their burst credit bucket fast under sustained load. Switch to Elastic throughput for anything where you can't predict the access pattern — it auto-scales and you pay for actual usage.
Can you lose this data if the host dies?
  └─ Yes → Instance Store (fastest, no extra cost)
  └─ No:
       Do multiple instances/pods need to write to it?
         └─ No → EBS gp3
         └─ Yes → EFS with Elastic throughput
Storage decisions compound. Getting this wrong and migrating a stateful workload later — especially in Kubernetes — means downtime and a data migration. Figure out the access pattern before you provision. One cost note: EFS looks expensive at ~$0.30/GB-month vs EBS gp3 at ~$0.08/GB-month, but EFS charges for actual usage. EBS charges for provisioned size. For large, sparsely-used filesystems the math flips in EFS's favour.