Articles

System Design – How SSDs Work Internally: NAND, FTL & Write Amplification

Dive into the inner mechanics of solid‑state drives, from floating‑gate NAND cells and multi‑level cell types to the Flash Translation Layer, garbage collection, write amplification and wear‑leveling. Understand why these details matter for database engineers and performance tuning.

Written by:
APin

Senior Technology Analyst • Verified Expert

More from this author →
System Design – How SSDs Work Internally: NAND, FTL & Write Amplification

Dive into the inner mechanics of solid‑state drives, from floating‑gate NAND cells and multi‑level cell types to the Flash Translation Layer, garbage collection, write amplification and wear‑leveling. Understand why these details matter for database engineers and performance tuning.

Context and Why SSD Internals Matter

Database engines are built around assumptions about the storage substrate. On a hard‑disk drive (HDD) the medium consists of rotating magnetic platters and a mechanical arm that seeks to a track; latency is measured in milliseconds and the cost of a random read or write is dominated by seek time. Solid‑state drives (SSDs) replace that mechanism with billions of floating‑gate transistors (NAND flash). Reads complete in microseconds because there are no moving parts, but the physical medium imposes a strict “erase‑before‑write” rule: a page can be programmed only after the entire block that contains it has been erased.

This constraint reshapes the entire I/O path that a database sees:

  • Granularity mismatch. The smallest readable and writable unit is a page (4–16 KB), while the smallest erasable unit is a block (256–512 pages, roughly 1–8 MiB). Overwriting a single 4 KB row therefore forces the SSD controller to write the new page elsewhere, mark the old page invalid, and later reclaim the block via garbage collection.
  • Write amplification. The ratio of NAND writes to host‑requested writes (WAF) rises because valid pages must be copied out of a victim block before the block can be erased. Enterprise SSDs typically achieve a steady‑state WAF of 1.5–3, whereas a naïve random workload can push the factor much higher.
  • Wear leveling. Each block endures a limited number of program/erase cycles. The controller distributes erasures evenly, preventing hot blocks from wearing out prematurely.

For a database engineer, these behaviors explain why sequential writes are far cheaper than random writes, and why tuning parameters such as RocksDB’s write_buffer_size or enabling TRIM/UNMAP can dramatically affect throughput and SSD lifespan. A practical example: inserting 1 MiB of rows in a random order may cause the SSD to relocate dozens of valid pages during garbage collection, inflating I/O volume and latency. In contrast, bulk‑loading data in sorted order lets the controller fill whole blocks, allowing a single erase operation to reclaim the entire region.

Understanding the erase‑before‑write constraint therefore lets engineers design write patterns, choose appropriate over‑provisioning levels, and align database log structures with the SSD’s page/block hierarchy, ultimately achieving predictable performance and durability.

NAND Flash Basics: Floating‑Gate Transistor and Cell Types

The fundamental storage element in NAND flash is the floating‑gate transistor. It consists of a control gate (word line) above a thin tunnel‑oxide, a conductive floating gate completely surrounded by insulating oxide, and a channel that connects source and drain. When a high voltage (≈20 V) is applied to the control gate, electrons quantum‑tunnel through the tunnel‑oxide onto the floating gate, raising its threshold voltage; this operation programs a logical “0”. Erasing applies a high voltage to the source, pulling the electrons off the floating gate and restoring the threshold to its low state, representing a logical “1”. During a read, a moderate voltage is placed on the control gate; if the transistor conducts, the cell is interpreted as “1”, otherwise as “0”. The trapped charge can remain for years without power, providing non‑volatile storage.

Cell types differ by how many distinct threshold‑voltage windows the controller can discriminate, which determines the number of bits stored per transistor:

  • SLC (single‑level cell): 1 bit per cell, two voltage levels, ~100 k program/erase (P/E) cycles, lowest error rate.
  • MLC (multi‑level cell): 2 bits per cell, four voltage levels, typically 10‑20 k P/E cycles, moderate error rate.
  • TLC (triple‑level cell): 3 bits per cell, eight voltage levels, ~3 k P/E cycles, higher raw bit‑error rate (RBER).
  • QLC (quad‑level cell): 4 bits per cell, sixteen voltage levels, ~1 k P/E cycles, highest RBER and most stringent voltage margins.

Increasing bits per cell raises capacity per die but introduces two trade‑offs:

  1. Program/erase endurance: More voltage windows require finer charge placement, accelerating wear. Enterprise SSDs therefore favor TLC (or SLC for mission‑critical caches) and employ aggressive over‑provisioning to spread P/E cycles evenly.
  2. Error susceptibility: Tighter voltage margins increase the likelihood of mis‑interpreting a level, so error‑correction code (ECC) strength must grow from simple BCH for SLC/MLC to LDPC for TLC/QLC. This adds latency and controller complexity.

Practical example: an enterprise database server may use a TLC‑based SSD with 28 % over‑provisioning, allowing the controller to maintain a write‑amplification factor near 1.5 while still meeting durability targets. A consumer laptop, prioritizing cost over endurance, might select a QLC drive, accepting a lower P/E budget and relying on the host OS’s TRIM command to keep garbage‑collection overhead manageable.

Page, Block, and Die Hierarchy

The NAND flash substrate is arranged in a strict hierarchy that determines the smallest units the SSD controller can address for different operations. At the bottom, a page is the minimum readable and writable entity, typically ranging from 4 KB to 16 KB. Pages are grouped into a block, which contains 256 – 512 pages; a block is the smallest region that can be erased, usually 1 – 8 MB in size. Several blocks form a plane (approximately one thousand blocks per plane). A die comprises 2 – 4 planes, and an SSD package contains 2 – 16 dies.

  • Page: 4 KB – 16 KB (read/write granularity)
  • Block: 256 – 512 pages (erase granularity)
  • Plane: ~1 000 blocks
  • Die: 2 – 4 planes
  • Package: 2 – 16 dies

Because NAND obeys the “erase‑before‑write” rule, the controller cannot overwrite a page in place. When the host issues an overwrite, the Flash Translation Layer (FTL) allocates a fresh, erased page, writes the new data, updates the logical‑to‑physical mapping, and marks the old page as invalid. This copy‑on‑write behavior is analogous to a notebook: you may write on any blank page, but to reuse a page that already contains ink you must tear out the entire chapter (the block) and replace it with a fresh one, even if only a single page in that chapter is stale.

Practical example: a database writes a 8 KB record. The SSD programs two 4 KB pages in an already‑erased block. Later the record is updated; the FTL writes the new 8 KB to two new pages in a different block and flags the original pages as invalid. The original block remains occupied until the garbage‑collection process selects it, copies any still‑valid pages elsewhere, and erases the whole block to make it reusable.

This hierarchy creates an inherent asymmetry: read/write operations work at page granularity, while erase operations work at block granularity. Understanding this asymmetry is essential for designing write‑intensive workloads, tuning TRIM/UNMAP commands, and minimizing write amplification in enterprise SSDs.

Flash Translation Layer (FTL) and Mapping Strategies

The Flash Translation Layer (FTL) sits between the host’s logical block address (LBA) space and the NAND’s physical page address (PPA) space. For every write the operating system issues, the FTL looks up the current mapping entry, allocates a fresh, erased physical page, programs the data there, and then updates the mapping so the LBA points to the new PPA. The original page is marked invalid but is not erased immediately; it becomes a candidate for later garbage‑collection. This log‑structured, copy‑on‑write behavior is analogous to the way LSM‑tree databases handle updates.

Mapping granularity

  • Page‑level mapping: each LBA maps to a single physical page. A 1 TB SSD with 4 KB pages contains roughly 250 million pages, requiring about 1 GB of DRAM for a full mapping table (250 M × 4 bytes per PPA). Enterprise SSDs typically provision 1–4 GB of DRAM to hold this table and to cache hot mapping entries.
  • Block‑level mapping: groups many LBAs into a single block‑level entry. The table size shrinks dramatically (e.g., a 1 TB drive with 256‑page blocks needs only ~1 MB of DRAM), but the FTL must maintain a separate “log” for recent writes. This hybrid approach is used in DRAM‑less or low‑cost SSDs, trading write performance for smaller memory footprints.

Copy‑on‑write overwrite process

  1. Host issues a write to LBA 42.
  2. FTL allocates a new, erased page (e.g., Die0, Block17, Page3).
  3. Data is programmed to the new page.
  4. The mapping table entry for LBA 42 is updated to point to the new PPA.
  5. The previous page (Die0, Block17, Page0) is marked invalid.

Invalid pages accumulate until the background garbage‑collection (GC) routine selects a victim block, copies any remaining valid pages to a new block, erases the victim, and returns it to the free‑page pool. Because the FTL never erases on overwrite, the drive can serve random writes without the latency of a block erase, but the trade‑off is increased write amplification when GC moves many valid pages.

Practical implications for engineers:

  • Provision sufficient DRAM to hold a full page‑level map when low latency and predictable performance are required.
  • Consider hybrid mapping only for workloads dominated by sequential writes or where cost constraints outweigh random‑write performance.
  • Design host‑side buffers (e.g., using TRIM/UNMAP) to help the FTL identify invalid pages early, reducing GC overhead.

Garbage Collection, Write Amplification, and Mitigation

Garbage collection (GC) is a critical background firmware process required by the erase-before-write constraint of NAND flash. Because data cannot be overwritten in place, the Flash Translation Layer (FTL) must reclaim blocks—the smallest unit of erasure—by migrating valid pages to new locations and clearing stale, invalid pages. The efficiency of this process is governed by victim block selection strategies:

  • Greedy: Selects the block with the highest number of invalid pages. This minimizes immediate migration overhead but may ignore the age or frequency of the valid data remaining.
  • Cost-Benefit: Evaluates blocks based on both the invalid page ratio and the data age. This strategy avoids relocating "hot" data, effectively reducing the frequency of future migrations.
  • FIFO (First-In-First-Out): Selects blocks based on the oldest write time. This is computationally inexpensive but often suboptimal, as it may reclaim blocks with high volumes of valid data.

The operational cost of these background movements is measured by the Write Amplification Factor (WAF), defined as the ratio of actual NAND writes to host-requested writes. A WAF of 1.0 represents perfect efficiency, whereas higher values indicate excessive internal writes that consume P/E cycles and degrade performance. Typical steady-state WAF ranges are 2–5 for consumer-grade drives and 1.5–3 for enterprise-grade drives, the latter benefiting from higher hardware-level overprovisioning.

Engineers can mitigate write amplification through several architectural strategies:

  • Overprovisioning (OP): Reserving additional physical NAND capacity (typically 7–28%) beyond the reported user capacity. This provides the GC process more "breathing room" to find blocks with higher invalid ratios, drastically reducing the number of valid pages that must be relocated.
  • TRIM/UNMAP: Utilizing OS-level commands to explicitly inform the controller that specific LBAs are no longer required, enabling the firmware to immediately mark those pages as invalid rather than waiting for them to be overwritten by the host.
  • Sequential Write Patterns: Designing database storage engines (e.g., LSM trees) to output data sequentially. Sequential writes fill entire blocks, ensuring that when data becomes stale, entire blocks are invalidated simultaneously, which renders GC near-instantaneous.
  • Write Coalescing: Buffering smaller host writes in DRAM to flush them as full, aligned pages, preventing unnecessary partial-page programming.

Wear Leveling and Longevity

Each NAND flash block can endure only a finite number of program/erase (P/E) cycles before the cells become unreliable. The endurance varies with cell type: single‑level cells (SLC) survive roughly 100 k cycles, triple‑level cells (TLC) about 3 k cycles, and quad‑level cells (QLC) around 1 k cycles. Enterprise SSDs typically employ TLC and allocate a large over‑provisioned area to keep the average P/E count well below the raw limit.

Because the SSD controller cannot overwrite a page in place, it writes new data to a fresh, erased page and marks the old page as invalid. Without intervention, “hot” logical blocks that are updated frequently would cause the same physical blocks to be erased repeatedly, exhausting their P/E budget while the rest of the drive remains relatively unused. This uneven wear shortens the usable lifespan of the device.

Wear‑leveling algorithms redistribute write activity across the entire NAND pool, ensuring that every block experiences a similar number of erase cycles. The two common strategies are:

  • Dynamic wear leveling: When a write targets a hot logical block, the controller allocates a previously less‑used physical block for the new page, thereby moving the hot data to a fresh block.
  • Static wear leveling: Periodically the firmware selects blocks that contain mostly static (unchanged) data and copies that data to a different block, freeing the original block to be used for future writes. This prevents long‑standing cold blocks from remaining idle while other blocks wear out.

Practical example: a database log file that receives continuous appends will cause the logical pages covering the log to be rewritten thousands of times per day. The SSD’s wear‑leveling routine will map each new log page to a different physical block, and after a configurable interval it will relocate older, immutable log segments to a separate block, allowing the original blocks to re‑enter the free pool.

By spreading P/E cycles evenly, wear leveling reduces the probability that any single block reaches its endurance limit early, thereby extending the overall drive lifespan and maintaining predictable performance throughout the SSD’s operational life.

APPWORKS ENGINEERING

Looking for Custom Software or AI Solutions?

Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.

Editorial Policy & Research Methodology

Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.

Have an Idea? we offer services in Lucknow, Bangalore, Delhi NCR and other locations