Pre-RFC: Gradual Transition of NixOS x86_64 Baseline to x86-64-v3 with an Intermediate Step to x86-64-v2

I’d like to start a discussion around eventually raising the default x86_64 baseline used by NixOS/nixpkgs.

The rough proposal is to do this in two stages:

  • move from the current x86-64 baseline to x86-64-v2
  • later move to x86-64-v3, tentatively around 2027

The main reason for doing this incrementally rather than jumping directly to v3 is to give us time to find compatibility problems, work out the nixpkgs/toolchain implications, and decide what we want long-term support for older x86_64 hardware to look like.

There are still some fairly fundamental questions around implementation and whether microarchitecture levels should be represented as separate systems, system features, or something else.

Background info

The current x86_64 baseline lets NixOS run on a very wide range of hardware, which is obviously useful. The downside is that the compiler has to assume a pretty old CPU and can’t make use of instructions that have been common for years.

Other distributions have started looking at, or already shipping, newer x86-64 microarchitecture levels. I think it’s worth discussing whether NixOS should eventually do the same.

The x86-64 psABI defines several commonly discussed levels:

  • x86-64-v1: effectively the original x86_64 baseline
  • x86-64-v2: adds instructions such as SSSE3, SSE4.1, SSE4.2, POPCNT, etc.
  • x86-64-v3: adds AVX, AVX2, BMI1/2, FMA, and related features
  • x86-64-v4: primarily AVX-512-era extensions

I don’t think v4 makes sense as a general-purpose NixOS baseline in the foreseeable future since it’s so new and would cut off a lot of people systems out, so this proposal is specifically about v2 and v3.

Phase 1: x86-64-v2

The first step would be switching the default x86_64 build baseline to something equivalent to:

-march=x86-64-v2

It cuts off some older x86_64 CPUs, but still supports considerably more hardware than v3. It would mostly serve as an intermediate baseline that lets us exercise the infrastructure required for changing the architecture level without immediately dropping everything that lacks AVX2.

During this period we could evaluate:

  • build failures caused by the new baseline
  • packages that make incorrect CPU-feature assumptions
  • Hydra/ofborg implications
  • binary cache compatibility
  • how architecture levels should be represented inside nixpkgs
  • how well the performance gains justify the compatibility cost

It would also give users and maintainers some advance warning before any eventual move to v3.

Phase 2: x86-64-v3

The longer-term goal would be moving the default baseline to:

-march=x86-64-v3

A possible target would be around 2027, although I don’t think the date should be considered fixed before we have actual data and agreement on the migration path. v3 is considerably more interesting from a performance perspective because it makes AVX2, BMI1/2, FMA and several other instruction set extensions universally available to the compiler.

At the same time there is still perfectly usable hardware which does not support v3. Before making such a transition, we would therefore need a reasonably clear answer to the question of what will happen to machines that don’t support v3

That could mean retaining a legacy package set, supporting multiple architecture levels in Hydra, relying on community-maintained configurations, or some other solution. I don’t think this pre-RFC needs to decide that yet, but we probably need to solve it before v3 can realistically become the default.

Why bother?

The obvious argument is performance.

If every package has to assume an original x86_64-era CPU, compilers cannot freely generate instructions which are available on almost all modern systems. Individual projects can and sometimes do perform runtime CPU dispatch, but that isn’t free and many packages simply don’t bother.

Moving the baseline allows those optimizations to happen across the distribution rather than only in specially optimized packages.

There is also an ecosystem argument. Several Linux distributions and distribution projects have been experimenting with newer x86-64 architecture levels, including Arch-related projects, RHEL/CentOS, openSUSE and others. CachyOS in particular has useful experience distributing packages optimized for newer architecture levels.

Compatibility

Biggest downside to this change by far is comp. Every increase in the baseline makes some hardware unable to run binaries from the normal NixOS binary cache. The v2 transition would affect older systems, while v3 would exclude a substantially larger group of machines.

NixOS has users running servers, workstations and unusual systems for much longer than the typical consumer replacement cycle. f we raise the baseline, we need to support these folks rather than abandoning those users.

Another complication is determining whether architecture levels should actually be exposed as different Nix system values.

For example, should we eventually have something conceptually like:

x86_64-linux
x86_64-v2-linux
x86_64-v3-linux

or should the architecture remain x86_64-linux with the microarchitecture expressed through features or another mechanism (If so what exactly)?

There are implications either way for package evaluation, binary substitutes, Hydra, flakes, cross compilation and ofborg.

Build infrastructure

things we’d need to look at:

  • stdenv and compiler configuration
  • bootstrap binaries
  • Hydra build platforms
  • binary caches
  • ofborg
  • package tests
  • cross compilation
  • packages with handwritten assembly
  • packages performing their own CPU detection
  • reproducibility between builders

ofborg in particular would need to understand whichever model we adopt so that PR evaluation and tests happen against the correct architecture level.

If multiple levels remain supported simultaneously, that also increases build and maintenance costs.

Benchmarking

One weakness of this proposal right now is that there isn’t enough NixOS-specific benchmark data.

I’ve seen claims and measurements from other distributions showing meaningful gains from v3 builds, including results in roughly the 10-20% range for some workloads, but that shouldn’t be interpreted as translating to actual NixOS performance since well distros can be vastly different from one another.

The effect is extremely workload-dependent.

Some applications will benefit substantially. Others will barely change at all. Projects which already perform runtime CPU dispatch may show almost no improvement from changing the distribution baseline.

Before making any decision I’d like to see repeatable NixOS/nixpkgs benchmarks comparing:

x86-64
x86-64-v2
x86-64-v3

on identical hardware.

Ideally the test set would contain a mixture of:

  • compilation
  • compression/decompression
  • cryptography
  • scientific/numerical workloads
  • multimedia encoding
  • databases
  • interpreters
  • common desktop applications
  • representative server workloads

CachyOS and similar projects could also be useful references, but they obviously shouldn’t substitute for measurements done with nixpkgs.

Migration strategy

A very rough migration could look something like this:

  1. Establish proper support for representing/building multiple x86-64 architecture levels.
  2. Add CI and Hydra coverage for those levels.
  3. Benchmark v1/v2/v3 using representative nixpkgs workloads.
  4. Document an easy way for users to check which level their CPU supports.
  5. Move the default to x86-64-v2.
  6. Keep collecting compatibility and performance data.
  7. Decide on a sustainable legacy-support mechanism.
  8. Consider moving the default to x86-64-v3 once the ecosystem and hardware distribution make that reasonable.

Probably the most important part is that each transition should be reversible during testing, rather than just plainly committing to v3.

Alternatives

Jump directly to x86-64-v3

This avoids spending time on an intermediate baseline and gets the largest potential optimization benefit immediately.

The downside is that the compatibility break is much larger and it gives us less opportunity to test the infrastructure incrementally.

Switch to x86-64-v2 only

Pretty much what it says on the tin. Move to v2 and leave v3 as an optional target instead of making it the future default.

Stay on the existing baseline

This gives us maximum compatibility and avoids additional build infrastructure.

The cost is that the default nixpkgs package set remains constrained by an architecture baseline dating back to the beginning of x86_64.

Maintain multiple baselines indefinitely

We could provide v1/v2/v3 package sets in parallel.

This would offer the best compatibility and optimization options for users, but would also have by far the largest Hydra, cache, CI and maintenance cost, which makes this option nonviable unless there some sort of miracle sponsor.

There may be a compromise where one architecture level is officially built by Hydra while older levels remain buildable but are not necessarily provided by the main binary cache.

Keep v1 and selectively build performance-sensitive packages for v3

Another option would be to keep the general nixpkgs baseline at x86-64-v1 while building selected performance-sensitive packages for x86-64-v3.

That could include things like compilers, compression tools, crypto libraries, multimedia codecs, numerical/scientific software, databases, and other packages where AVX2/FMA/BMI actually produce measurable gains.

This would preserve compatibility for most of nixpkgs while still getting some of the practical benefit of v3 where it matters.

The main problem is dependency handling. If a supposedly v1-compatible package ends up depending on a library built for v3, then that whole closure effectively requires v3. Maybe we can cache a v1 and v3 version of these packages?

We’d also need to decide how optimized variants are exposed and cached. This could be done through separate package sets, overlays, or some other mechanism rather than changing the entire system .

This approach would be interesting to benchmark, since it might capture a large part of the real-world performance benefit without requiring the entire distribution to move to v3.

Questions

The questions I’d like feedback on are:

  • Does moving the default baseline make sense at all?
  • Is x86-64-v2 a useful transition step, or should we eventually go directly from the current baseline to v3?
  • How should microarchitecture levels be represented in Nixpkgs: separate systems, system features, or something else?
  • What level of legacy hardware support should NixOS commit to?
  • Would building multiple architecture levels in Hydra be practical?
  • What should the benchmark suite look like?
  • What existing work in nixpkgs could be reused for this?
  • What would need to change in ofborg?
  • What is the best way to let users determine whether their system supports v2 or v3?

I’m particularly interested in input from people familiar with stdenv/bootstrap, Hydra, ofborg, cross compilation, and the existing gcc.arch / platform CPU-feature machinery.

41 Likes

It would be nice if anyone wanting to work on this would take the time to read systems/architecture: bump default architecture to x86-64-v2 by SuperSandro2000 · Pull Request #202526 · NixOS/nixpkgs · GitHub and summarize it and put the important objectives that we higlighted in that PR to make this pre-RFC realistic.

As of now, it contains no useful information for NixOS stakeholders.

6 Likes
  1. Arch Linux Community Discussions: On the Arch Linux mailing list, there’s a discussion about the shift to x86_64-v3 microarchitectures. See the Benchmark here.

  2. Sunnyflunk’s Analysis: A GitHub user named Sunnyflunk provides a comprehensive analysis of the x86-64-v3’s performance, revealing a varied/mixed bag impact across different applications. Refer to the analysis here.

  3. CachyOS Performance Insights: Phoronix tested CachyOS, an Arch Linux-based distribution with v3 support, and reported some notable performance improvements. You can go through the data here.

  4. CentOS ISA Performance Investigation: The CentOS ISA Special Interest Group conducted an extensive review of different ISA levels, including x86-64-v3, offering valuable insights into performance changes. Their detailed findings can be found here.

  5. Red Hat’s Strategy with RHEL 9: Red Hat discusses their decision to build Red Hat Enterprise Linux 9 for the x86-64-v2 microarchitecture level, considering various CPU incompatibilities and performance. The blog post is available here.

3 Likes

First of all, thank you for looking into this topic. Although I am highly sceptical of the benefits, I also think we should take advantage of them should they actually exist, so I support any efforts towards clearing up this matter.

It’s been reasonably well shown that generic compiler optimisations can provide significant benefit for many applications. Clear Linux significantly outperforms most generic Linux distros across a wide range of tasks by using package-specific optimisations. These often include march flags aswell but I don’t think it’s clear whether these drive the performance benefit.

Benchmarks I have seen I’ve seen on raising march have not been very convincing so far. They usually have massive biases (i.e. selecting only packages which are known to benefit from generic compiler optimisations), do not actually test µArch optimisations in isolation but in combination with unsafe optimisations (-O3, which we will not use) and none of them demonstrate benefits for users, only higher (or lower) numbers in a collection of semi-synthetic benchmarks.

Based on this rather poor quality data, you can already tell that the benefit is highly dependant on the specific package. The amount of packages that benefit significantly appear to be rather low, possibly less than 50%.

This is conjecture but I additionally do not believe that most of these synthetic benchmarks necessarily reflect a better user-experience.
I think we should instead focus on applications which users actually need to be performant. On the desktop, this would include commonly interacted tools such as coreutils, browsers, text editors, word processors and the like.

A note on hardware:

The thought that compiler optimisations such as AVX only exclude “older systems” (as in: decades old) is wrong.

There is hardware as recently released as this year which does not support AVX of any kind: Tremont (microarchitecture) - Wikipedia. Moving to v3 in 2027 would exclude this hardware 4 years after release.

Such low power Celerons are somewhat popular in the homelab scene for their extreme power efficiency and low prices (commonly available in used thin clients).

My NAS uses a Celeron J4105 from 2017 (Goldmont Plus) and I know that @musicmatze uses a similar chip aswell.


I am of the opinion that, if compiler optimisations only really help a small group (or category) of packages, they should be applied to those specific packages only. Ideally by switching the code paths at runtime which many packages already do.

If you can show data which shows a more wide-spread significant increase in performance (let’s say a median improvement >5% across 10^2-10^3 “desirable” packages), I’d revise that opinion as that’s too many to reasonably “optimise” by hand.

All in all, this whole endeavour gets a big rejection from me until there is clearer data showing the benefits.
As a general purpose distro, we should not take excluding hardware lightly. Even if it actually is very old, there might still be uses left for it. Not to mention people who aren’t quite as socioeconomically privileged as most of us whose only access to anything resembling a PC is ancient hardware we threw away.

It is also worth mentioning that the people who really need it (i.e. HPC people or misguided Gentoomen) can and could always apply these generic tree-wide optimisations themselves for their environments.

20 Likes

Many of these analysis are also mentioned in the issue @RaitoBezarius posted.
Unfortunately, many of them have problems that make it hard to know if x86-64-v3 is beneficial for NixOS:

  1. Sunnyflunk’s Analysis: As mentioned at the end of the blog post, it didn’t compare x86-64 to x86-64-v3, but it also change other compiler flags, so the results are useless if we want to show that x86-64-v3 is worth it.

However, this post was intended to be more about x86-64-v3, but some quick tests (which requires further analysis) suggest that CachyOS using -O3 is what’s actually responsible for some of the larger gains rather than x86-64-v3.

  1. CachyOS Performance Insights: Same comparison as 2., same problems.

  2. CentOS ISA Performance Investigation: As mentioned in the blog post, it didn’t compare x86-64 to x86-64-v3, but it also changed the compiler version, so the results are useless if we want to show that x86-64-v3 is worth it.

Given that we changed both the compiler version and the baseline, we dug into which of those variables contributed the most impactful change to the results. For the latter two benchmarks, we saw a 2.2x speed up. Mocassin seems to benefit the most from the auto-vectorization that GCC12 does.

As also mentioned in the a fore mentioned issue, I think we need benchmarks of a proposed NixOS change, so we can see what performance uplift we really get.

5 Likes

I feel like waiting till 2027 to migrate to x86_64-v3 puts Nix quite behind the industry. Does NixOS Foundation have enough resources to run v3 or maybe even v4 builds in addition to standard x86_64? If that’s the case, it might be the right way to go about it purely from marketing perspective

1 Like

Nope, we don’t have the resources to do so.

1 Like

I don’t believe we should follow “industry trends” just for the sake of following them.

In the current moment, there are very clear downsides to moving to new march targets and little to no good data on the benefits.

“But everyone is doing it” does not count as a benefit IMHO.

18 Likes

what resources do you need?

maybe the companies that are interested in these optimisations can help with those resources ?

~100-400TB extra of S3 storage and maybe something like 1000ish cores of compute time to mass rebuild all of that when needed?

Given that no one proved they bring anything to the table, that’s highly uncertain.

1 Like

whoa…

ok, that’s a big ask.

Agreed.

(Playing devil’s advocate, I’m not personally convinced that -v2 does much performance-wise.)

FWIW, x86_64-v2 doesn’t mandate AVX. Steam’s hardware survey (one of the best public data source to look at, IMO) shows SSE4.2 at 99.52% availability, vs. AVX at 97.28%.

I don’t know if I generally agree with that. We probably exclude more users and interesting use cases by not supporting ARMv7 than we’d do by moving the x86_64 baseline to -v2. NixOS makes it trivial for users to build from source if they have specific architecture constraints, so “someone might need it” doesn’t seem like a super strong argument to me.

9 Likes

For machines out there that are used for gaming. Skips over lots of machines out there: servers, routers, workstations.

I’m pretty sure that due to growing system requirements, the average gaming PC is more modern, than everything else out there.

12 Likes

How do enthusiastic NixOS users go about testing the impact of these flags themselves? The last time I tried to set build flags to optimize for a specific x86-64 psABI level for my whole system following the guide on the wiki, it simply didn’t seem to work: Nix CPU global CPU flags - #2 by pauldoo

2 Likes

I would argue the only real reason we don’t support ARMv7 is that because it’s hard to have it in CI, we have a… surprisingly good and active maintenance of ARMv7 in NixOS (yes, people are running systemd with it and what not.)

3 Likes

In the past, I did the work to look into this, you can use two of my branches towards this:

They simulate what would be the changes to nixpkgs if we bumped the minimal baseline.

I built them over https://hydra.newtype.fr/jobset/nixos/trunk-combined-x86_64-v2 and https://hydra.newtype.fr/jobset/nixos/trunk-combined-x86_64-v3, but I think I removed recently the binaries because I wanted to bump with a recent unstable and use the new timeout features for tests because my Hydra often ended up stuck in NixOS tests for no reason.

If people are interested, I can rebase, clean up and ask Hydra to re-evaluate.

(The Hydra links are IPv6-only, I am sorry for people who may not have IPv6, I do not have money to spare on IPv4.)

3 Likes

Also, with glibc-hwcaps, shouldn’t it be possible to provide multiple compiled libraries in a single package. One could be for x86-64 and another for x86-64-v2, and a third possibly for x86-64-v3.

It would also allow for this to be enabled or disabled at a package level. Some performance sensitive packages could build for multiple levels (media codecs, compression libraries, etc), while others might opt to build only for the baseline (a basic text editor, mkfs, lots of other examples).

3 Likes

This doesn’t change the storage costs.

If only someone can come up with a list of package that benefit from it.

1 Like

Surely it must. There is more to a compiled package than the binaries and libraries. There are all sorts of other assets. Using glibc hwcaps only the libraries and binaries are duplicated, not the entire package.

4 Likes

The note was pertaining x86_64-v3.

v2 is a much easier pill to swallow as hardware without SSE4 really is getting to the point of not being useful anymore as even basic ARM SoCs outperform the best CPUs from that era nowadays. Even there I’d err on the side of caution though.

With v2 however, the benefits are even more questionable than with v3.

I’m all for supporting armv7l-linux too. I’ve got two older RPIs that I’d like to put NixOS on.

Difference is that we never supported armv7l-linux to any decent capacity while x86_64-v1 has pretty much always been supported.

The problem is that we don’t know who might need it. It could be literally noone or thousands; we’re blind here.

That could probably happen organically.

For example, let’s say someone wants to compress their music library to a higher FLAC level to save on storage. Being a typical NixOS user, they might spend an unreasonable time optimising the re-encode to be a few minutes faster. Assuming such a flag optimising for separate HWCAPS was already proliferated in Nixpkgs, they might try it out to see whether it makes a difference and whip up a quick PR if it shows a significant benefit.

What I also like about the glibc HWCAPS approach is that we could optimise packages for even higher levels (i.e. x86_64-v4 with AVX512) where I wouldn’t be surprised if gains were quite significant without breaking the other >90% of users’ systems.

3 Likes