Allow and prefer special vdevs as ZIL #17505

amotin · 2025-07-02T17:23:10Z

Before this change ZIL blocks were allocated only from normal or SLOG vdevs. In typical situation when special vdevs are SSDs and normal are HDDs it could cause weird inversions when data blocks are written to SSDs, but ZIL referencing them to HDDs.

This change assumes that special vdevs typically have much better (or at least not worse) latency than normal, and so in absence of SLOGs should store ZIL blocks. It means similar to normal vdevs introduction of special embedded log allocation class and updating the allocation fallback order to: SLOG -> special embedded log -> special -> normal embedded log -> normal.

The code tries to guess whether data block is going to be written to normal or special vdev (it can not be done precisely before compression) and prefer indirect writes for blocks written to a special vdev to avoid double-write. For blocks that are going to be written to normal vdev, special vdev by default plays as SLOG, reducing write latency by the cost of higher special vdev wear, but it is tunable via module parameter.

This should allow HDD pools with decent SSD as special vdev to work under synchronous workloads without requiring additional SLOG SSD, impractical in many scenarios.

Types of changes

Bug fix (non-breaking change which fixes an issue)
New feature (non-breaking change which adds functionality)
Performance enhancement (non-breaking change which improves efficiency)
Code cleanup (non-breaking change which makes code smaller or more readable)
Quality assurance (non-breaking change which makes the code more robust against bugs)
Breaking change (fix or feature that would cause existing functionality to change)
Library ABI change (libzfs, libzfs_core, libnvpair, libuutil and libzfsbootenv)
Documentation (a change to man pages or other documentation)

Checklist:

My code follows the OpenZFS code style requirements.
I have updated the documentation accordingly.
I have read the contributing document.
I have added tests to cover my changes.
I have run the ZFS Test Suite with this change applied.
All commit messages are properly formatted and contain Signed-off-by.

pcd1193182

Conceptually, I think this makes sense. The special class is likely going to have better performance, and using that to store ZIL writes makes a lot of sense. The two parts that give me pause are:

The special embedded log; this is a permanent sacrifice of some of the high-performing storage dedicated to storing metadata. For users with zil-heavy workloads, who don't also have a SLOG, it's probably worth it; but for users who don't fit that bill, they're losing some of their special space unconditionally. On the other hand, it's not that much space, all things considered. One metaslab (currently at most 16GiB) is not the end of the world.
The attempt to automatically detect which class the write will eventually end up in. This sort of auto-sensing is very tricky, and can result in weird an unpredictable behavior once deployed. That said, for this case it's probably not too problematic if it gets it wrong; the data will still end up in the right place eventually.

module/zfs/vdev.c

module/zfs/zio.c

Before this change ZIL blocks were allocated only from normal or SLOG vdevs. In typical situation when special vdevs are SSDs and normal are HDDs it could cause weird inversions when data blocks are written to SSDs, but ZIL referencing them to HDDs. This change assumes that special vdevs typically have much better (or at least not worse) latency than normal, and so in absence of SLOGs should store ZIL blocks. It means similar to normal vdevs introduction of special embedded log allocation class and updating the allocation fallback order to: SLOG -> special embedded log -> special -> normal embedded log -> normal. The code tries to guess whether data block is going to be written to normal or special vdev (it can not be done precisely before compression) and prefer indirect writes for blocks written to a special vdev to avoid double-write. For blocks that are going to be written to normal vdev, special vdev by default plays as SLOG, reducing write latency by the cost of higher special vdev wear, but it is tunable via module parameter. This should allow HDD pools with decent SSD as special vdev to work under synchronous workloads without requiring additional SLOG SSD, impractical in many scenarios. Signed-off-by: Alexander Motin <[email protected]> Sponsored by: iXsystems, Inc.

amotin · 2025-07-11T18:35:02Z

@pcd1193182 :

but for users who don't fit that bill, they're losing some of their special space unconditionally.

True. I decided that it does not worth complexity of changing it in run time when SLOG added/removed, and may be not good for fragmentation to not reserve the log space. With my recent change we may allow fallback from special class to special_embedded_log to allow use of that space if have to, but with the same consequences for fragmentation once it is used.

That said, for this case it's probably not too problematic if it gets it wrong; the data will still end up in the right place eventually.

True. The only difference is whether we double-write when we should not or not when we should. The difference should be only in performance.

robn · 2025-07-15T10:00:41Z

include/sys/spa_impl.h

 	metaslab_class_t *spa_log_class;	/* intent log data class */
 	metaslab_class_t *spa_embedded_log_class; /* log on normal vdevs */
 	metaslab_class_t *spa_special_class;	/* special allocation class */
+	metaslab_class_t *spa_special_embedded_log_class; /* log on special */
 	metaslab_class_t *spa_dedup_class;	/* dedup allocation class */


Future cleanup opportunity: there's possibly enough of these now to make them an array indexed on an enum, and then we don't have to repeat ourselves over and over when doing something to all of them (init/fini, stats, etc).

amotin force-pushed the special_zil branch from 6a5e50b to 193af7f Compare July 2, 2025 17:33

amotin added the Status: Code Review Needed Ready for review and testing label Jul 2, 2025

behlendorf self-requested a review July 3, 2025 17:39

amotin mentioned this pull request Jul 9, 2025

PANIC at zil.c:872:zil_free_lwb() #17509

Open

pcd1193182 approved these changes Jul 11, 2025

View reviewed changes

module/zfs/vdev.c Outdated Show resolved Hide resolved

module/zfs/zio.c Show resolved Hide resolved

amotin force-pushed the special_zil branch from 193af7f to 4ecd3f0 Compare July 11, 2025 18:27

robn approved these changes Jul 15, 2025

View reviewed changes

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Allow and prefer special vdevs as ZIL #17505

Allow and prefer special vdevs as ZIL #17505

amotin commented Jul 2, 2025 •

edited

Loading

Uh oh!

pcd1193182 left a comment

Uh oh!

Uh oh!

Uh oh!

amotin commented Jul 11, 2025 •

edited

Loading

Uh oh!

robn Jul 15, 2025

Uh oh!

Uh oh!

Allow and prefer special vdevs as ZIL #17505

Are you sure you want to change the base?

Allow and prefer special vdevs as ZIL #17505

Conversation

amotin commented Jul 2, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Types of changes

Checklist:

Uh oh!

pcd1193182 left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

amotin commented Jul 11, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

robn Jul 15, 2025

Choose a reason for hiding this comment

Uh oh!

Uh oh!

amotin commented Jul 2, 2025 •

edited

Loading

amotin commented Jul 11, 2025 •

edited

Loading