FSx for NetApp ONTAP Tiering, Measured: What Moves, When, and What It Saves
We ran the four FSx for ONTAP tiering policies side by side on a live file system and measured where the data actually went. Here is what moved, how fast, what it does to the bill, and the one cost line tiering never touches.
Introduction: The Promise, Tested
Amazon FSx for NetApp ONTAP makes a simple promise: keep the data people are using on SSD, move the rest to a much cheaper capacity pool, and do it automatically. The documentation describes four tiering policies that decide what moves and when. What it cannot tell you is how that behaves on your data, how quickly it happens, and what it does to the monthly bill.
So we measured it. We built a file system in AWS Canada Central with four identical volumes, one per tiering policy, wrote the same data to each, and read ONTAP's own footprint counters to see where every gigabyte landed. The whole lab is Terraform, and it is public: you can run the same experiment in your own account with fsx-ontap-eks-lab.
The results below are from our run. Prices are AWS on-demand rates for ca-central-1 as of September 2026; other regions differ, but the ratios are what matter.
How FSx for ONTAP Bills You
Three lines dominate the bill, and they behave differently. SSD storage is provisioned: you pay for the size you chose, used or not. The capacity pool is usage-based: you pay for what is stored there. Throughput capacity is provisioned separately, in MBps. Each gigabyte of SSD also comes with 3 SSD IOPS included.
| Line item (Single-AZ, ca-central-1) | Price | Billed on |
|---|---|---|
| SSD storage | $0.138 per GB-month | Provisioned size |
| Capacity pool storage | $0.0238 per GB-month | Data stored |
| Throughput capacity | $0.788 per MBps-month | Provisioned throughput |
| Capacity pool requests | $0.0055 per 1,000 writes, $0.00044 per 1,000 reads | Requests |
Data in the capacity pool costs about a sixth as much as data on SSD (a 5.8 to 1 ratio). Two details from the FSx for ONTAP guide shape how much SSD you actually need. Up to 16% of the SSD tier is reserved for ONTAP overhead. And AWS recommends not running the SSD tier above 80% utilization, partly because it also stages writes to, and caches random reads from, the capacity pool. Above 90%, reads from the pool stop being cached on SSD at all. You size SSD for your hot data plus headroom, not for everything you store.
The Four Tiering Policies
Each volume has exactly one policy. File metadata always stays on SSD, whatever the policy. The behavior below is from the AWS documentation:
| Policy | What moves to the pool | Cooling period | When tiered data is read |
|---|---|---|---|
NONE | Nothing | n/a | n/a |
SNAPSHOT_ONLY | Snapshot data only | 2 days by default (2 to 183) | Brought back to SSD |
AUTO | All cold data: files and snapshots | 31 days by default (2 to 183) | Random reads bring it back to SSD; sequential reads, such as an antivirus scan, leave it in the pool |
ALL | All data, marked cold immediately | None | Served from the pool; stays cold |
The defaults are worth knowing because they differ by tool. A volume created in the FSx console defaults to AUTO. One created through the AWS CLI, the API (and therefore Terraform), or the ONTAP CLI defaults to SNAPSHOT_ONLY. The same team can end up with different behavior depending on how a volume was made.
The Threshold the Policy Table Leaves Out
The policy decides what is eligible to move. Whether anything actually moves depends on how full the SSD tier is. The tiering thresholds in the AWS guide are easy to miss and change the picture completely:
| SSD tier utilization | What happens |
|---|---|
| 50% or below | Only ALL volumes tier. AUTO and SNAPSHOT_ONLY move nothing, however long their data has been cold. |
| Above 50% | AUTO and SNAPSHOT_ONLY tier data that has passed its cooling period. The move happens 24 to 48 hours after the period expires, as a low-priority background job. |
| 90% or above | Reads from the capacity pool are no longer cached on SSD, so repeat reads keep coming from the pool. |
| 98% or above | All tiering stops. Reads continue, writes fail. |
The first row is the one that matters for cost. A file system that is generously provisioned, which is what you get when SSD is sized for peak or for "everything, to be safe", will never tier an AUTO volume, and the bill will look exactly like NONE. Tiering only starts paying once the tier is more than half full. That is a design constraint, not a bug: ONTAP has no reason to shuffle data when there is no pressure on the expensive tier. But it means "we turned on AUTO" and "we are saving money" are two different statements, and only the footprint counters tell you which one is true.
The Experiment
One Single-AZ file system at the smallest size AWS offers (1024 GiB of SSD, 128 MBps), one storage VM, and four 10 GiB volumes that differ only in their tiering policy. The two policies with a cooling period use the minimum, two days, so the lab finishes in days rather than a month.
We ran it in two rounds. In the first, the four volumes put under 1% of data on the SSD tier, which is how a fresh or over-provisioned file system looks in practice. In the second, we added a fifth volume, tier_filler with the NONE policy, and wrote about 456 GiB of random data to it, taking the SSD tier from 1.2% to 54.6% of its 862 GiB usable (1024 GiB provisioned, less ONTAP's overhead). Then we waited again. The difference between the two rounds is the threshold at work.
locals {
tiering_volumes = {
tier_none = { policy = "NONE", cooling = null }
tier_snapshot_only = { policy = "SNAPSHOT_ONLY", cooling = 2 }
tier_auto = { policy = "AUTO", cooling = 2 }
tier_all = { policy = "ALL", cooling = null }
}
}
resource "aws_fsx_ontap_volume" "tiering" {
for_each = local.tiering_volumes
name = each.key
junction_path = "/${each.key}"
size_in_megabytes = 10240
storage_virtual_machine_id = aws_fsx_ontap_storage_virtual_machine.this.id
storage_efficiency_enabled = true
tiering_policy {
name = each.value.policy
cooling_period = each.value.cooling
}
}We wrote 2 GiB of random data to each volume over NFS. Random data matters: deduplication and compression would otherwise shrink the footprint and blur the result. Then we read each volume's SSD and capacity-pool footprint from the ONTAP REST API. The lab wraps that in one command, which runs on an instance inside the VPC so the storage admin password never leaves AWS:
$ make footprint
volume policy used_gib snapshot_gib ssd_gib pool_gib
tier_auto auto 2.04 0 2.05 0
tier_none none 2.04 0 2.04 0
tier_all all 2.04 0 0.05 2
tier_snapshot_only snapshot_only 3.58 2.03 4.09 0For SNAPSHOT_ONLY we took a snapshot and then overwrote the file, so the volume holds 2 GiB of old blocks that exist only in the snapshot. Those are exactly the blocks that policy is meant to move.
What Actually Moved
| Volume | Right after writing (SSD / pool) | 90 seconds later | Two days later, SSD under 50% | Cooling period after SSD pushed past 50% |
|---|---|---|---|---|
NONE | 2.04 / 0 GiB | 2.04 / 0 GiB | 2.12 / 0 GiB | 2.08 / 0 GiB |
ALL | 2.04 / 0 GiB | 0.05 / 2.00 GiB | 0.14 / 2.00 GiB | 0.09 / 2.00 GiB |
AUTO | 2.05 / 0 GiB | 2.05 / 0 GiB | 2.13 / 0 GiB | 0.09 / 2.00 GiB |
SNAPSHOT_ONLY | 2.05 / 0 GiB | 4.09 / 0 GiB (2.03 GiB in the snapshot) | 4.18 / 0 GiB | 2.12 / 2.00 GiB (the snapshot moved, the live data stayed) |
tier_filler (NONE, round 2 only) | — | — | — | 456.74 / 0 GiB |
ALL moved in about 90 seconds. New writes land on SSD first, as AWS documents, and a background process then moves them. On our small file system that took about a minute and a half, leaving 0.05 GiB of metadata on SSD. If you are using ALL for a backup or DR copy, the SSD you need is for metadata and in-flight writes, not for the data set.
NONE stayed put, as it should. That is the baseline every other policy is measured against: everything on the most expensive tier.
Snapshots cost SSD until they tier. After the overwrite, the SNAPSHOT_ONLY volume held 4.09 GiB on SSD: the new data plus 2.03 GiB of blocks kept only for the snapshot. Snapshot retention on a volume that never tiers is paid for at SSD prices, which is easy to miss when snapshot schedules are set once and forgotten.
Round one: nothing moved. Two full days after the write, with the cooling period long expired and the SSD tier at 1.2%, AUTO still held all 2.13 GiB on SSD and SNAPSHOT_ONLY still held 4.18 GiB, snapshot blocks included. Not a policy failure: the threshold. Below 50% utilization there is no pressure on the expensive tier, so ONTAP does not spend effort relieving it. A file system that was sized generously behaves exactly like this in production, and the tiering line on the bill never appears.
Round two: both moved, together, 41 hours after the tier crossed 50%. The capacity pool grew in a single five-minute window between 19:05 and 19:15 UTC on 28 September, inside the 24-to-48-hour lag AWS documents. AUTO went from 2.13 GiB on SSD to 0.09 GiB, with 2.00 GiB in the pool: the same shape as ALL, just later. SNAPSHOT_ONLY moved exactly the 2.03 GiB held only by the snapshot and kept its 2.12 GiB of live data on SSD, which is the whole point of that policy. Two things follow. The cooling period is a floor, not a schedule: this data had been cold for four days before anything happened, and what finally triggered the move was the tier filling up, then the daily scanner reaching it. And when the scanner ran, it took everything eligible in one pass rather than trickling.
The Line Tiering Never Touches
Throughput capacity is billed on what you provision, and tiering does nothing to it. At the smallest size, 128 MBps, it costs about $101 a month in this region, against about $141 for the minimum 1024 GiB of SSD. On a minimum-size file system, throughput is over 40% of the bill before a single byte is stored.
The practical consequence: tiering and throughput are two separate decisions. Tiering lowers the storage line. Right-sizing throughput, based on what the workload actually pulls rather than a guess made at provisioning time, lowers the other. Cost reviews that only look at capacity leave the second one untouched.
What It Means at Real Scale
As an illustration, take a 10 TiB file share where about 20% of the data is in active use. Using the prices above and AWS's 80% SSD guidance:
| Policy | SSD provisioned | Capacity pool | Storage per month |
|---|---|---|---|
NONE | 12,800 GiB (all 10 TiB at 80%) | 0 | $1,766 |
AUTO | 2,560 GiB (2 TiB hot at 80%) | 8,192 GiB | $548 ($353 SSD + $195 pool) |
That is about 69% off the storage line, roughly $1,200 a month on one share. Throughput is identical in both cases and sits on top. The arithmetic is deliberately simple: it leaves out ONTAP's overhead (which adds SSD to both, so it widens the gap in absolute terms), snapshots, capacity-pool request charges, and whatever deduplication and compression save. It also assumes the cold data really is cold. A workload that randomly reads across its whole data set keeps pulling blocks back to SSD, and the savings shrink.
Choosing a Policy
NONE— Latency-sensitive data that must never wait on the pool: databases, build caches, anything with a strict performance floor. Pay SSD prices knowingly, and only for these volumes.AUTO— General file shares, home directories, and project data where most files go quiet after a few weeks. The right default for most volumes. Tune the cooling period to how long your data actually stays warm, not the 31-day default, and check that the SSD tier actually runs above 50%: below that,AUTOdoes nothing and you are payingNONEprices.SNAPSHOT_ONLY— Active data that must stay on SSD but with long snapshot retention. It keeps the working set fast and moves the history to the cheaper tier.ALL— Backup, archive, and disaster-recovery copies, such as the destination of a SnapMirror relationship. Reads are served from the pool, so it is not for anything users open day to day.
If You Run Kubernetes on It
FSx for ONTAP can also back Kubernetes storage on Amazon EKS through NetApp Trident, NetApp's CSI driver. Each persistent volume Trident creates with the ontap-nas driver is its own ONTAP volume, so everything above applies per volume. One setting deserves attention: Trident's documented default for tieringPolicy is none. Unless you change it, every persistent volume in the cluster sits on SSD for its whole life.
It is a backend default in the Trident backend configuration, applied to volumes provisioned after it is set:
apiVersion: trident.netapp.io/v1
kind: TridentBackendConfig
metadata:
name: fsx-ontap-nas
namespace: trident
spec:
version: 1
storageDriverName: ontap-nas
svm: svm1
aws:
fsxFilesystemID: fs-0123456789abcdef0
credentials:
name: arn:aws:secretsmanager:ca-central-1:123456789012:secret:vsadmin
type: awsarn
defaults:
tieringPolicy: autoUse separate backends, or separate storage classes pointing at them, when some workloads need none and others can tier. The lab repository includes an EKS and Trident setup to try this against the same file system.
Run It Yourself
Everything above comes from fsx-ontap-eks-lab, an Apache-2.0 Terraform lab we maintain. It builds the file system inside a private VPC with its own encryption key, a client reachable only through Systems Manager, and a budget alarm, and it tears everything down in one command. Budget about $9.50 a day in this region while it runs.
git clone https://github.com/nubiscore/fsx-ontap-eks-lab
cd fsx-ontap-eks-lab
cp terraform/terraform.tfvars.example terraform/terraform.tfvars # region, budget email
# for round 2, also set: tiering_filler_gib = 480
make up # about 20 minutes
make footprint # run now, and again after the cooling period
make downTiering Is a Design Decision, Not a Checkbox
The capacity pool is several times cheaper than SSD, and FSx for ONTAP moves data there reliably. The savings are real. They depend on choices made per volume and per file system: which policy, which cooling period, whether the SSD tier is full enough for tiering to run at all, how much throughput, and whether Kubernetes volumes were ever told they are allowed to tier.
Running or planning FSx for ONTAP? NubisCore designs and implements the cloud around it: landing zones and networking, tiering and throughput sizing, SnapMirror migrations, and Kubernetes storage with Trident. We will show you where your storage bill actually comes from and what changes it.
References
- AWS — FSx for ONTAP: Volume storage capacity. Tiering policies, cooling periods, defaults, and read behavior.
- AWS — FSx for ONTAP: Managing storage capacity. ONTAP overhead, the 80% SSD recommendation, and the 90% caching threshold.
- AWS — FSx for NetApp ONTAP pricing. Prices in this article are from the AWS Price List API for ca-central-1, September 2026.
- NetApp — Trident: Configure the storage backend for FSx for ONTAP. Backend options, including
tieringPolicy. - NubisCore — fsx-ontap-eks-lab. The Terraform, the measurement script, and the raw results.