Back to Articles
CloudSep 29, 2026

FSx for NetApp ONTAP Tiering, Measured: What Moves, When, and What It Saves

We ran the four FSx for ONTAP tiering policies side by side on a live file system and measured where the data actually went. Here is what moved, how fast, what it does to the bill, and the one cost line tiering never touches.

Share

Introduction: The Promise, Tested

Amazon FSx for NetApp ONTAP makes a simple promise: keep the data people are using on SSD, move the rest to a much cheaper capacity pool, and do it automatically. The documentation describes four tiering policies that decide what moves and when. What it cannot tell you is how that behaves on your data, how quickly it happens, and what it does to the monthly bill.

So we measured it. We built a file system in AWS Canada Central with four identical volumes, one per tiering policy, wrote the same data to each, and read ONTAP's own footprint counters to see where every gigabyte landed. The whole lab is Terraform, and it is public: you can run the same experiment in your own account with fsx-ontap-eks-lab.

The results below are from our run. Prices are AWS on-demand rates for ca-central-1 as of September 2026; other regions differ, but the ratios are what matter.

How FSx for ONTAP Bills You

Three lines dominate the bill, and they behave differently. SSD storage is provisioned: you pay for the size you chose, used or not. The capacity pool is usage-based: you pay for what is stored there. Throughput capacity is provisioned separately, in MBps. Each gigabyte of SSD also comes with 3 SSD IOPS included.

Line item (Single-AZ, ca-central-1)PriceBilled on
SSD storage$0.138 per GB-monthProvisioned size
Capacity pool storage$0.0238 per GB-monthData stored
Throughput capacity$0.788 per MBps-monthProvisioned throughput
Capacity pool requests$0.0055 per 1,000 writes, $0.00044 per 1,000 readsRequests

Data in the capacity pool costs about a sixth as much as data on SSD (a 5.8 to 1 ratio). Two details from the FSx for ONTAP guide shape how much SSD you actually need. Up to 16% of the SSD tier is reserved for ONTAP overhead. And AWS recommends not running the SSD tier above 80% utilization, partly because it also stages writes to, and caches random reads from, the capacity pool. Above 90%, reads from the pool stop being cached on SSD at all. You size SSD for your hot data plus headroom, not for everything you store.

The Four Tiering Policies

Each volume has exactly one policy. File metadata always stays on SSD, whatever the policy. The behavior below is from the AWS documentation:

PolicyWhat moves to the poolCooling periodWhen tiered data is read
NONENothingn/an/a
SNAPSHOT_ONLYSnapshot data only2 days by default (2 to 183)Brought back to SSD
AUTOAll cold data: files and snapshots31 days by default (2 to 183)Random reads bring it back to SSD; sequential reads, such as an antivirus scan, leave it in the pool
ALLAll data, marked cold immediatelyNoneServed from the pool; stays cold

The defaults are worth knowing because they differ by tool. A volume created in the FSx console defaults to AUTO. One created through the AWS CLI, the API (and therefore Terraform), or the ONTAP CLI defaults to SNAPSHOT_ONLY. The same team can end up with different behavior depending on how a volume was made.

The Threshold the Policy Table Leaves Out

The policy decides what is eligible to move. Whether anything actually moves depends on how full the SSD tier is. The tiering thresholds in the AWS guide are easy to miss and change the picture completely:

SSD tier utilizationWhat happens
50% or belowOnly ALL volumes tier. AUTO and SNAPSHOT_ONLY move nothing, however long their data has been cold.
Above 50%AUTO and SNAPSHOT_ONLY tier data that has passed its cooling period. The move happens 24 to 48 hours after the period expires, as a low-priority background job.
90% or aboveReads from the capacity pool are no longer cached on SSD, so repeat reads keep coming from the pool.
98% or aboveAll tiering stops. Reads continue, writes fail.

The first row is the one that matters for cost. A file system that is generously provisioned, which is what you get when SSD is sized for peak or for "everything, to be safe", will never tier an AUTO volume, and the bill will look exactly like NONE. Tiering only starts paying once the tier is more than half full. That is a design constraint, not a bug: ONTAP has no reason to shuffle data when there is no pressure on the expensive tier. But it means "we turned on AUTO" and "we are saving money" are two different statements, and only the footprint counters tell you which one is true.

The Experiment

One Single-AZ file system at the smallest size AWS offers (1024 GiB of SSD, 128 MBps), one storage VM, and four 10 GiB volumes that differ only in their tiering policy. The two policies with a cooling period use the minimum, two days, so the lab finishes in days rather than a month.

We ran it in two rounds. In the first, the four volumes put under 1% of data on the SSD tier, which is how a fresh or over-provisioned file system looks in practice. In the second, we added a fifth volume, tier_filler with the NONE policy, and wrote about 456 GiB of random data to it, taking the SSD tier from 1.2% to 54.6% of its 862 GiB usable (1024 GiB provisioned, less ONTAP's overhead). Then we waited again. The difference between the two rounds is the threshold at work.

terraform/modules/fsx-ontap/main.tfhcl
locals {
  tiering_volumes = {
    tier_none          = { policy = "NONE", cooling = null }
    tier_snapshot_only = { policy = "SNAPSHOT_ONLY", cooling = 2 }
    tier_auto          = { policy = "AUTO", cooling = 2 }
    tier_all           = { policy = "ALL", cooling = null }
  }
}

resource "aws_fsx_ontap_volume" "tiering" {
  for_each = local.tiering_volumes

  name                       = each.key
  junction_path              = "/${each.key}"
  size_in_megabytes          = 10240
  storage_virtual_machine_id = aws_fsx_ontap_storage_virtual_machine.this.id
  storage_efficiency_enabled = true

  tiering_policy {
    name           = each.value.policy
    cooling_period = each.value.cooling
  }
}

We wrote 2 GiB of random data to each volume over NFS. Random data matters: deduplication and compression would otherwise shrink the footprint and blur the result. Then we read each volume's SSD and capacity-pool footprint from the ONTAP REST API. The lab wraps that in one command, which runs on an instance inside the VPC so the storage admin password never leaves AWS:

bash
$ make footprint
volume              policy         used_gib  snapshot_gib  ssd_gib  pool_gib
tier_auto           auto           2.04      0             2.05     0
tier_none           none           2.04      0             2.04     0
tier_all            all            2.04      0             0.05     2
tier_snapshot_only  snapshot_only  3.58      2.03          4.09     0

For SNAPSHOT_ONLY we took a snapshot and then overwrote the file, so the volume holds 2 GiB of old blocks that exist only in the snapshot. Those are exactly the blocks that policy is meant to move.

What Actually Moved

VolumeRight after writing (SSD / pool)90 seconds laterTwo days later, SSD under 50%Cooling period after SSD pushed past 50%
NONE2.04 / 0 GiB2.04 / 0 GiB2.12 / 0 GiB2.08 / 0 GiB
ALL2.04 / 0 GiB0.05 / 2.00 GiB0.14 / 2.00 GiB0.09 / 2.00 GiB
AUTO2.05 / 0 GiB2.05 / 0 GiB2.13 / 0 GiB0.09 / 2.00 GiB
SNAPSHOT_ONLY2.05 / 0 GiB4.09 / 0 GiB (2.03 GiB in the snapshot)4.18 / 0 GiB2.12 / 2.00 GiB (the snapshot moved, the live data stayed)
tier_filler (NONE, round 2 only)———456.74 / 0 GiB

ALL moved in about 90 seconds. New writes land on SSD first, as AWS documents, and a background process then moves them. On our small file system that took about a minute and a half, leaving 0.05 GiB of metadata on SSD. If you are using ALL for a backup or DR copy, the SSD you need is for metadata and in-flight writes, not for the data set.

NONE stayed put, as it should. That is the baseline every other policy is measured against: everything on the most expensive tier.

Snapshots cost SSD until they tier. After the overwrite, the SNAPSHOT_ONLY volume held 4.09 GiB on SSD: the new data plus 2.03 GiB of blocks kept only for the snapshot. Snapshot retention on a volume that never tiers is paid for at SSD prices, which is easy to miss when snapshot schedules are set once and forgotten.

Round one: nothing moved. Two full days after the write, with the cooling period long expired and the SSD tier at 1.2%, AUTO still held all 2.13 GiB on SSD and SNAPSHOT_ONLY still held 4.18 GiB, snapshot blocks included. Not a policy failure: the threshold. Below 50% utilization there is no pressure on the expensive tier, so ONTAP does not spend effort relieving it. A file system that was sized generously behaves exactly like this in production, and the tiering line on the bill never appears.

Round two: both moved, together, 41 hours after the tier crossed 50%. The capacity pool grew in a single five-minute window between 19:05 and 19:15 UTC on 28 September, inside the 24-to-48-hour lag AWS documents. AUTO went from 2.13 GiB on SSD to 0.09 GiB, with 2.00 GiB in the pool: the same shape as ALL, just later. SNAPSHOT_ONLY moved exactly the 2.03 GiB held only by the snapshot and kept its 2.12 GiB of live data on SSD, which is the whole point of that policy. Two things follow. The cooling period is a floor, not a schedule: this data had been cold for four days before anything happened, and what finally triggered the move was the tier filling up, then the daily scanner reaching it. And when the scanner ran, it took everything eligible in one pass rather than trickling.

The Line Tiering Never Touches

Throughput capacity is billed on what you provision, and tiering does nothing to it. At the smallest size, 128 MBps, it costs about $101 a month in this region, against about $141 for the minimum 1024 GiB of SSD. On a minimum-size file system, throughput is over 40% of the bill before a single byte is stored.

The practical consequence: tiering and throughput are two separate decisions. Tiering lowers the storage line. Right-sizing throughput, based on what the workload actually pulls rather than a guess made at provisioning time, lowers the other. Cost reviews that only look at capacity leave the second one untouched.

What It Means at Real Scale

As an illustration, take a 10 TiB file share where about 20% of the data is in active use. Using the prices above and AWS's 80% SSD guidance:

PolicySSD provisionedCapacity poolStorage per month
NONE12,800 GiB (all 10 TiB at 80%)0$1,766
AUTO2,560 GiB (2 TiB hot at 80%)8,192 GiB$548 ($353 SSD + $195 pool)

That is about 69% off the storage line, roughly $1,200 a month on one share. Throughput is identical in both cases and sits on top. The arithmetic is deliberately simple: it leaves out ONTAP's overhead (which adds SSD to both, so it widens the gap in absolute terms), snapshots, capacity-pool request charges, and whatever deduplication and compression save. It also assumes the cold data really is cold. A workload that randomly reads across its whole data set keeps pulling blocks back to SSD, and the savings shrink.

Choosing a Policy

  • NONE — Latency-sensitive data that must never wait on the pool: databases, build caches, anything with a strict performance floor. Pay SSD prices knowingly, and only for these volumes.
  • AUTO — General file shares, home directories, and project data where most files go quiet after a few weeks. The right default for most volumes. Tune the cooling period to how long your data actually stays warm, not the 31-day default, and check that the SSD tier actually runs above 50%: below that, AUTO does nothing and you are paying NONE prices.
  • SNAPSHOT_ONLY — Active data that must stay on SSD but with long snapshot retention. It keeps the working set fast and moves the history to the cheaper tier.
  • ALL — Backup, archive, and disaster-recovery copies, such as the destination of a SnapMirror relationship. Reads are served from the pool, so it is not for anything users open day to day.

If You Run Kubernetes on It

FSx for ONTAP can also back Kubernetes storage on Amazon EKS through NetApp Trident, NetApp's CSI driver. Each persistent volume Trident creates with the ontap-nas driver is its own ONTAP volume, so everything above applies per volume. One setting deserves attention: Trident's documented default for tieringPolicy is none. Unless you change it, every persistent volume in the cluster sits on SSD for its whole life.

It is a backend default in the Trident backend configuration, applied to volumes provisioned after it is set:

backend-ontap-nas.yamlyaml
apiVersion: trident.netapp.io/v1
kind: TridentBackendConfig
metadata:
  name: fsx-ontap-nas
  namespace: trident
spec:
  version: 1
  storageDriverName: ontap-nas
  svm: svm1
  aws:
    fsxFilesystemID: fs-0123456789abcdef0
  credentials:
    name: arn:aws:secretsmanager:ca-central-1:123456789012:secret:vsadmin
    type: awsarn
  defaults:
    tieringPolicy: auto

Use separate backends, or separate storage classes pointing at them, when some workloads need none and others can tier. The lab repository includes an EKS and Trident setup to try this against the same file system.

Run It Yourself

Everything above comes from fsx-ontap-eks-lab, an Apache-2.0 Terraform lab we maintain. It builds the file system inside a private VPC with its own encryption key, a client reachable only through Systems Manager, and a budget alarm, and it tears everything down in one command. Budget about $9.50 a day in this region while it runs.

bash
git clone https://github.com/nubiscore/fsx-ontap-eks-lab
cd fsx-ontap-eks-lab
cp terraform/terraform.tfvars.example terraform/terraform.tfvars   # region, budget email
# for round 2, also set: tiering_filler_gib = 480
make up          # about 20 minutes
make footprint   # run now, and again after the cooling period
make down

Tiering Is a Design Decision, Not a Checkbox

The capacity pool is several times cheaper than SSD, and FSx for ONTAP moves data there reliably. The savings are real. They depend on choices made per volume and per file system: which policy, which cooling period, whether the SSD tier is full enough for tiering to run at all, how much throughput, and whether Kubernetes volumes were ever told they are allowed to tier.

Running or planning FSx for ONTAP? NubisCore designs and implements the cloud around it: landing zones and networking, tiering and throughput sizing, SnapMirror migrations, and Kubernetes storage with Trident. We will show you where your storage bill actually comes from and what changes it.

References