IOps (I/O Operations Per Second)

Performance

IOPS (I/O operations per second) is the rate at which a storage device completes read and write operations. It counts operations, not bytes: throughput equals IOPS times the I/O size. It binds small random access, where a 7,200 rpm HDD is rated in the hundreds and datacentre NVMe above a million.

Also known as IOPS, I/O operations per second, input/output operations per second.

12 min read Updated Aug 30, 2026

Full Explanation

IOPS (I/O operations per second) is the rate at which a storage device completes read and write operations. It is a storage metric, not a network one. It is not a byte rate: throughput counts bytes moved per second. IOPS counts operations completed. The size of each operation joins the two. IOPS is also not a fixed property of a device. The same disk reports wildly different IOPS for large sequential reads and small random ones. It reports different figures again for reads and for writes. IOPS is the ceiling that binds when a workload issues many small, scattered operations. That is exactly what a CDN edge does when it serves small cached objects.

How it works

SNIA's Solid State Storage glossary defines IOPS as the number of inputs/outputs per second. It says IOPS "provides an indication of the performance of a device in applications generating random reads and/or writes". The same glossary defines throughput separately, as the amount of data transferred in a period. Throughput is "typically measured in MegaBytes per second (MB/s)". It is indicative of performance under sequential reads or writes. Two metrics, two questions: how many operations, and how many bytes. SNIA's current online dictionary carries the term as IOPS / IOPs / iops, "shorthand for I/O Operations per second".

One relationship joins them: bytes per second equals IOPS times the I/O size. Every device therefore has both an operation ceiling and a byte ceiling. Small operations run into the first ceiling, while large ones run into the second. AWS documents the arithmetic with a worked example. A gp2 volume under 1,000 GiB, with burst credits available, has an IOPS limit of 3,000 and a throughput limit of 250 MiB/s. At a 256 KiB I/O size, "your volume reaches its throughput limit at 1000 IOPS (1000 x 256 KiB = 250 MiB)". At 16 KiB, the same volume sustains all 3,000 IOPS, because the bytes stay well under the cap (Amazon EBS I/O characteristics and monitoring).

What counts as "one operation" is not simply "one request". That is the first thing to get straight. Amazon EBS caps a single I/O at 256 KiB on SSD volumes and 1,024 KiB on HDD volumes. It merges physically sequential smaller requests up to that cap, and it splits larger ones. Eight sequential 32 KiB requests count as one IOPS. The same eight requests, issued randomly, count as eight. A single 1,024 KiB request on an SSD volume counts as four. The HDD-backed st1 and sc1 volumes are different. AWS says these volumes "define performance in terms of throughput rather than IOPS". Their accounting is blunter still: "an I/O request of 1 MiB or less counts as a 1 MiB I/O credit" (Amazon EBS HDD volume types). A 4 KiB random read there costs the same credit as a 1 MiB one.

Device class sets the scale. It does so far more sharply for operations than for bytes. SNIA notes that a solid state storage device "will typically exhibit higher IOPS than HDD". That is because semiconductor storage has none of the electromechanical latencies of a spinning drive. Current vendor datasheets show the size of the gap:

  • One 7,200 rpm enterprise HDD is Seagate's Exos X18. It is rated at 170 random-read and 550 random-write IOPS at 4 K, queue depth 16, with 4.16 ms average latency, against a maximum sustained sequential transfer rate of 270 MB/s.
  • One datacentre SATA SSD is Solidigm's D3-S4520. It is rated up to 92,000 4 KB random-read and up to 48,000 4 KB random-write IOPS.
  • One datacentre PCIe 4.0 NVMe SSD is Solidigm's D7-P5520. It is rated up to 1.1 million 4 KB random-read and up to 220,000 4 KB random-write IOPS.

Read those figures against each other and the purpose of the metric appears. On the Exos, 170 random 4 K reads per second is well under 1 MB/s of useful data. That is against 270 MB/s sequential from the same drive. Access pattern alone decides a difference of more than two orders of magnitude in bytes delivered. Across the three parts, the random-read rating spans 170 to 1.1 million. Sequential byte rates, by contrast, differ by roughly one order of magnitude. Flash reads and writes are not symmetric, either. The D7-P5520's write rating is about a fifth of its read rating. A write-heavy workload therefore meets a different ceiling than a read-heavy one, on the same device.

Every one of those numbers is a vendor rating at a stated block size, access pattern and queue depth. It also assumes enough work is in flight to reach it. That last condition is easy to miss. AWS warns that "if your workload is not delivering enough I/O requests to fully use the performance available to your EBS volume, then your volume might not deliver the IOPS or throughput that you have provisioned". AWS recommends an average queue depth of one for every 1,000 provisioned IOPS. Push past the available rate instead, and the cost lands on latency. Consistently driving more IOPS to a volume than it has available "can cause increased I/O latency".

Why it matters for a CDN

A CDN edge server serves enormous numbers of small objects per second: images, JavaScript bundles, stylesheets, API responses. Every object not already in memory is a separate storage operation. That makes the operation rate, rather than the byte rate, the quantity most likely to run out first. A node can sit comfortably below its bandwidth ceiling and still be at its operation ceiling. That is why a capacity plan expressed only in gigabits per second misses the real constraint.

IOPS is spent on both sides of a cache lookup. A hit served from RAM costs no disk I/O at all. So a high cache hit ratio in the memory tier is directly an IOPS saving. A hit that has fallen to the disk tier costs a read instead. A miss costs more than a read. That is because the cache fill that follows a miss writes the object into the cache, and on flash the random-write ceiling is materially lower than the random-read ceiling. Churn therefore spends the operation budget twice: once evicting, once refilling. Eviction rules such as LRU decide which objects keep their place in the fastest tier. So cache policy and IOPS pressure are one problem, seen from two sides.

When the operation queue saturates, requests wait. The wait reaches users as latency on the miss path, rather than as a bandwidth shortfall. The architectural answer is a tier per order of magnitude: hottest objects in memory, warm working set on flash, coldest and largest content on cheap bulk storage. This way, the fastest and most operation-constrained tier only ever handles objects that are genuinely in demand.

What CDNs do

  • Fastly describes the tier ladder from the client's side. In a February 2017 engineering post, Hooman Beheshti, its VP of Technology, writes that a cached object "can be served from the memory of that machine (best case scenario), disk storage on that machine (let's hope that means an SSD because if it doesn't, there's an extra performance hit there too), a local peer/parent, or a not-so-local peer/parent". He writes that performance degrades along that chain. Memory, then local disk, then a peer, then a more distant peer: each step down costs either operations or a network hop. Reference: Fastly, The truth about cache hit ratios.
  • Cloudflare puts the ladder between data centres rather than inside one. With Tiered Cache, "if content is not cached in lower-tier data centers (generally the ones closest to a visitor), the lower-tier must ask an upper-tier to see if it has the content". Only the upper tier may go to the origin. Above that sits Cache Reserve, "a large, persistent data store implemented on top of R2". It is documented as "the ultimate upper-tier cache" and is now offered as part of Cloudflare's Smart Shield. It is a paid product with its own admission rules. An asset needs a freshness TTL of at least 10 hours to be admitted. It is removed if it goes unrequested for the retention period, which starts at 30 days and resets on each request. Long-lived content therefore sits in object storage instead of occupying the tiers where operations are scarce. References: Cloudflare Tiered Cache and Cloudflare Cache Reserve.
  • AWS is where IOPS is most visibly a purchasable quantity rather than a property of hardware. Provisioned IOPS SSD volumes "deliver their provisioned IOPS performance 99.9 percent of the time". The io1 volumes "can range in size from 4 GiB to 16 TiB and you can provision from 100 IOPS up to 64,000 IOPS per volume". The io2 Block Express volumes support "Provisioned IOPS up to 256,000". Those ceilings are instance-dependent, not universal. AWS states that Nitro-based instances support volumes provisioned with up to 256,000 IOPS. Other instance types can attach volumes provisioned up to 64,000 IOPS, but they achieve up to 32,000. At the other end of the catalogue, HDD-backed volumes are capped at 500 IOPS for st1 and 250 for sc1, counted in 1 MiB units. References: AWS EBS Provisioned IOPS SSD volumes and Amazon EBS volume types.

All three converge on the same shape. Keep the fastest, most operation-constrained tier small and hot. Push everything else down to storage, where bytes are cheap and operations are not scarce.

Watch out for

  • A bare IOPS number is not a fact. It needs a block size, a read-to-write ratio, an access pattern and a queue depth. Those are the variables that define the workload. SNIA's own definition of a workload names them: "the ratio of reads to writes, and the block size (Access Pattern), and the data pattern". "170 IOPS" and "270 MB/s" describe the same drive.
  • Do not read %util as saturation on an SSD or a RAID array. Note which way the error runs. iostat reports %util as the "percentage of elapsed time during which I/O requests were issued to the device". So a device that serves requests in parallel can show a figure close to 100 percent while still being far from its limit. The man page is explicit: saturation near 100 percent holds "for devices serving requests serially". But "for devices serving requests in parallel, such as RAID arrays and modern SSDs, this number does not reflect their performance limits". Judge those devices by request rate and queue length instead.
  • Reads and writes have separate ceilings. On the flash parts above, random-write ratings sit well below random-read ratings. Sizing a write-heavy cache-fill path against a read benchmark overstates the headroom you have.
  • A fresh device flatters itself. SNIA notes that devices straight out of the box, or purged to a near-new state, "exhibit a short period of higher performance which then levels off to a relatively sustained level of performance called Steady State". A benchmark that does not run long enough to reach steady state reports a number the device will not hold in production.
  • IOPS and throughput fail independently. A device can be under its operation cap and at its byte cap, or the reverse. Sizing on one number silently ignores the other.
  • In the cloud, operations are a line item. On Amazon EBS "you pay only for what you provision". Provisioned IOPS is billed by the amount provisioned per month, whether or not you drive it. A gp3 volume includes a free baseline of 3,000 IOPS and 125 MB/s, and anything above that is charged (Amazon EBS pricing). The same bytes, served as many small objects rather than few large ones, cost more operations. Measure the real object mix before committing to a number.

Best practice

  • Tier by order of magnitude: hottest objects in RAM, warm working set on NVMe or SATA flash, cold and large content on high-capacity disks or object storage. Spend the scarce operations on objects that are actually requested.
  • Benchmark with your own workload, not a synthetic 4 KiB-only run. fio "spawns a number of threads or processes doing a particular type of I/O action as specified by the user". Its job file takes the I/O type, block size, I/O size and I/O engine as explicit parameters. So describe the real block sizes, read/write mix, access pattern and queue depth, and run long enough to reach steady state.
  • Monitor the operation rate next to the byte rate. In iostat, r/s is "the number (after merges) of read requests completed per second for the device". The w/s column is the write equivalent. Watch both alongside aqu-sz, "the average queue length of the requests that were issued to the device". A queue that grows while bandwidth stays low means IOPS is the bottleneck. Buying bandwidth will not help.
  • Remember that "after merges" matters. The kernel, and in the cloud the storage layer, coalesce adjacent operations. So the count the device sees is lower than the count your application issued. Compare like with like when reconciling application metrics against device metrics.
  • Keep headroom on the flash tier. Budget operations explicitly wherever they are priced. Do not treat IOPS as a free property of the hardware.

Examples

Benchmark IOPS on a cache server:

# Check current disk IOPS with iostat (Linux)
iostat -x 1 5
# Look at the r/s (reads/sec) and w/s (writes/sec) columns

# Benchmark random read IOPS with fio
fio --name=random-read \
    --ioengine=libaio \
    --rw=randread \
    --bs=4k \
    --numjobs=4 \
    --size=1G \
    --runtime=30 \
    --direct=1

# Typical results:
# HDD:  ~150 IOPS
# SATA SSD: ~75,000 IOPS
# NVMe SSD: ~400,000 IOPS
Storage Tier     Random IOPS    Access Time    CDN Use Case
HDD              100-200        8-14ms         Cold archive storage
SATA SSD         10K-100K       ~0.1ms         Warm cache tier
NVMe SSD         100K-500K+     ~0.05ms        Hot cache tier
RAM              Millions+      ~0.0001ms      Hottest content cache

Frequently Asked Questions

IOPS (I/O operations per second) is the rate at which a storage device completes read and write operations. It counts operations, not bytes: throughput equals IOPS times the I/O size. It binds small random access, where a 7,200 rpm HDD is rated in the hundreds and datacentre NVMe above a million.

Benchmark IOPS on a cache server:

# Check current disk IOPS with iostat (Linux)
iostat -x 1 5
# Look at the r/s (reads/sec) and w/s (writes/sec) columns

# Benchmark random read IOPS with fio
fio --name=random-read \
    --ioengine=libaio \
    --rw=randread \
    --bs=4k \
    --numjobs=4 \
    --size=1G \
    --runtime=30 \
    --direct=1

# Typical results:
# HDD:  ~150 IOPS
# SATA SSD: ~75,000 IOPS
# NVMe SSD: ~400,000 IOPS
Storage Tier     Random IOPS    Access Time    CDN Use Case
HDD              100-200        8-14ms         Cold archive storage
SATA SSD         10K-100K       ~0.1ms         Warm cache tier
NVMe SSD         100K-500K+     ~0.05ms        Hot cache tier
RAM              Millions+      ~0.0001ms      Hottest content cache

Yes. IOps (I/O Operations Per Second) is also known as IOPS, I/O operations per second, input/output operations per second. IOPS (I/O operations per second) is the rate at which a storage device completes read and write operations. It counts operations, not bytes: throughput equals IOPS times the I/O size. It binds small random access, where a 7,200 rpm HDD is rated in the hundreds and datacentre NVMe above a million.

Related CDN concepts include:

  • Throughput — Throughput is the rate at which data actually crosses a link in a measured interval …