Where the bytes actually live
Everything in the lab is replaceable except one thing. The nodes are second-hand desktops; if one dies I buy another. The containers rebuild from configuration. The VPS rebuilds from a compose file. But the data, the photos and the documents and the years of accumulated everything, exists nowhere else in the world. Storage is the one part of a homelab where a mistake is not a learning experience, it is a loss.
So it gets treated differently from everything else. This is how.
§ One machine owns the bytes
All storage lives on a single NAS running TrueNAS, and every other machine borrows from it. The nodes keep their guests' operating systems on local disks for speed, but anything measured in terabytes, and anything that would hurt to lose, lives on the NAS.
Centralising it is a deliberate trade. It means one box to protect properly instead of protection smeared thinly across five, and one place where "is the data okay" is a question with an answer. The cost is obvious: that box matters more than any other machine I own. Which is why the rest of this post exists.
§ The pool
The disks are arranged as a single ZFS pool: a dozen spinning drives in RAIDZ2, which means any two of them can fail outright and nothing is lost. Alongside them sits a fast NVMe device that absorbs synchronous writes, so the guests waiting on write confirmations get answered at NVMe speed while the spinning disks catch up behind it.
Two-disk tolerance sounds generous until you have lived through a rebuild. When a big disk fails and gets replaced, the pool spends a long time re-reading everything to rebuild it, and that is precisely when the surviving disks work hardest. RAIDZ2 exists for the second failure that happens during that window.
§ Disks lie, so make them prove it
The property that made me choose ZFS is not the pooling or the snapshots. It is that ZFS refuses to trust the disks.
Every block written gets a checksum, and every read verifies it. When a disk quietly returns something different from what was written, and at sufficient scale a disk eventually will, ZFS notices, serves the correct copy from redundancy, and repairs the bad one. Silent corruption stops being silent.
Trust also gets audited on a schedule. Every Sunday at midnight a scrub walks the entire pool and re-reads every allocated byte, forcing each disk to prove it can still return what it accepted. The scrub reports a number, and the number I care about is zero errors. Not "probably fine". Zero, measured weekly.
§ Traffic gets its own wire
Storage traffic between the NAS and the nodes runs over a dedicated network segment that carries nothing else. Two protocols share it, for two different jobs.
NFS carries the shared datasets: the media library and the backup target, things several guests read at once and where a filesystem is the right shape. NVMe-oF carries block storage: raw disks served over the network for the one VM that wants a real disk rather than a file share, because for that workload latency matters more than shareability.
Keeping this on its own segment means a heavy transfer never competes with everything else on the network, and it means the storage protocols, which were never designed to meet the internet, structurally cannot.
§ Snapshots, and the trap in their names
The dataset that matters, the one holding the guest backups and VM disks, gets snapshotted on two schedules: hourly, kept for two days, and daily, kept for a month. A snapshot is ZFS remembering what the data looked like at a moment, and because of how ZFS writes, taking one is instant and keeping one costs only the space of what changed since.
The daily snapshot is timed thirty minutes after the nightly guest backups finish, so every daily pins a complete, fresh backup set. Retention can prune what it likes afterwards; the pinned copies stay for a month.
One trap worth passing on: TrueNAS matches snapshots to their schedule by the snapshot's naming pattern. Give two schedules the same naming pattern and the short retention will happily prune the long schedule's snapshots, and nothing warns you. The two schedules here have deliberately distinct names, and I verified the behaviour by running the tasks and reading the snapshot list, not by assuming.
§ What deliberately gets less protection
The media library, tens of terabytes of it, is not snapshotted. That is a decision, written down, not an oversight.
Snapshots protect against mistakes, and the mistake that matters for media is a misconfigured automation quietly mass-deleting things. I have accepted that risk, because the library is re-acquirable and the terabytes of snapshot churn are not worth it. Protection budgets are finite; spending them evenly across data of unequal value protects the wrong things.
If you take one idea from this post, take that one: decide per dataset, write the decision down, and let some data be less protected on purpose.
§ Leaving the house, encrypted first
The important data also syncs to object storage far away from here, at a different company in a different city. It is encrypted before it leaves, and the keys stay home. The provider stores ciphertext; what they can read of my files is nothing.
That covers the failure modes a pool cannot: fire, theft, flood, or me doing something catastrophic to the NAS itself. The pool protects against disks; distance protects against the room the disks are in.
§ And every guest, every night
Separately from all of the above, every container and VM in the cluster gets backed up nightly to the pool, half past two, compressed, all of them, no exceptions list to forget to update. Guest restore is therefore boring: pick the archive, restore it, done. The snapshot schedule pinning each night's set is what lets those archives be pruned aggressively without fear.
§ The point
Nothing in this post is clever, and that is the point. A pool that tolerates two failures, checksums that catch lying disks, a weekly audit that proves it, snapshots for mistakes, encryption before anything leaves, and a nightly sweep of every guest. Each piece is dull. The property they add up to is the only one that matters: the data outlives any single bad day.