LLD prefers system LLVM libraries to user provided ones on Windows

We’re having an issue where lld-link.exe on Windows seems to prefer libraries that come from its own LLVM installation over ones provided by the user. This is an issue for us specifically where we have newer (and patched) versions of the Fortran runtime libraries from the flang project that we need to link our library against, however LLD will pick the Fortran runtimes from its own LLVM installation and not the library path we provide on the link line.

I have also observed similar when trying to link against a newly built MLIR from git using a clang+lld from a released LLVM install, where I see linker errors because lld tries to choose the MLIR libraries from its own install path.

I believe this behaviour was introduced by this patch (⚙ D151188 [LLD][COFF] Add LLVM toolchain library paths by default.) in an attempt to make sure that lld can find compiler-rt when linking on Windows. I understand the motivation there but I think the behaviour as regards the Flang runtimes and MLIR libraries I have described above is quite surprising to the end user, as it shouldn’t matter that the linker has those libraries in its install path if the user is explicitly asking for them from somewhere else.

Does anyone have any ideas how we can improve this behaviour so that user provided library paths are considered first for these libraries?

Thanks!

IMHO: this is the intended behaviour and how it works if you have libc++ in your toolchain path on Linux as well. The way it’s solved there is to pass -nostdlib I think (I can’t check the source right now) and it won’t add those paths.

I think this behaviour is correct, we should not pick libraries in your libpath that collides with what the toolchain ships, that can lead to other unintended consequences, a flag to disable this behaviour seems correct to me, but open to other ideas.

I think the case is different for libc++ as clang and libc++ expect to be the same version and the user doesn’t explicitly ask to link against libc++ in that case. Similarly if linking with flang-new I would expect the Fortran runtime libraries for the flang-new you’re linking with to be selected.

Neither of these are what actually happens with the current behaviour though. If you imagine a situation where someone has LLVM installed, and then builds clang;flang;libc++ but not lld, when the user tries to link with that clang and/or flang-new they will get the runtime libraries from lld’s LLVM install, not from the LLVM they have just built, most likely leading to a (possibly silent!) error down the line. Again I think this is surprising behaviour.

A flag would at least allow us to work around this issue but I think we should do something about the underlying behaviour, even if that something is to explicitly say you must build lld when building clang/flang on Windows and you must use that new lld for linking and not any existing one you have. We would in that case also have to add a warning/error to the cmake scripts to inform the user of this, in my opinion.

Cc @hansw2000 @petrhosek @mstorsjo @rnk that was all in the original discussion about these changes.

I will reply more in depth when I have time to think about it a bit.

As a side note - we already ran into other surprises due to this directory ordering, fixed in [LLD] [COFF] Restore the current dir as the first entry in the search… · llvm/llvm-project@f906fd5 · GitHub.

After having another think about this, I think the solution we chose, to force these include directories to be first, probably was a bit overzealous.

I think (although I have no data to back it up), that the number of cases where it fixes a case where we’d accidentally use a different version (people having a separate copy of e.g. libc++ in their explicitly passed libdirs, but not wanting to use it), is quite low - while it does block people from getting what they are explicitly requesting.

One of the issues we IIRC tried to fix, was that we should prefer the compiler-rt libraries shipped along with LLVM, over an older copy that might have shipped with MSVC (which they do for x86 and x64, but not for arm/arm64) - but we can probably still achieve this.

In llvm-project/lld/COFF/Driver.cpp at 4e8986fc58dd88cbef9089a9b2841e0a87cbb481 · llvm/llvm-project · GitHub, we could flip the order of addClangLibSearchPaths and adding directories from OPT_libpath, while still keeping addClangLibSearchPaths before detectWinSysRoot and addLibSearchPaths.

1 Like

Moving the user provided include directories would fix our issue so I’d support that change.

I’m not sure if it fixes the more general issue I described though, where one has built clang, libc++ and compiler-rt but not lld, will lld in that case choose the libc++ and compiler-rt from its own LLVM installation or the newly built one?

I’m not sure I fully understand why the clang driver can’t provide the link paths to these libraries like it does on Linux rather than expecting them to be found magically? Presumably we wouldn’t see any of these issues in that case? It’d be good for me to understand this so we can match the correct behaviour in the flang driver.

The main reason for this, is that within the MSVC ecosystem, the norm is that one doesn’t do linking by invoking the compiler driver, but one invokes link.exe or lld-link directly, so the compiler can’t pass in the relevant paths. We don’t want to embed library paths to use for lookup within the object files themselves, thus the idea is to have lld-link automatically add these paths.

I don’t think the case about having compiler-rt, libc++ and lld built and installed in separate roots is one that has been considered, and I don’t think there’s much interest to handle that in any extra graceful way - I guess the user (or the build system) would need to specify those directories manually in that case, if the lld install tree doesn’t have them.

I think that’s fine but in this case I think we should issue a warning in the CMake configure step that if the user has lld installed already and is building some LLVM runtimes (libc++/compiler-rt/openmp) that they also should build lld as well to get correct linking behaviour. That shouldn’t be too difficult to do and should avoid anyone getting confused by this. Do you think that would make sense?

Just to clarify I still think, in addition to that warning, we should make the change that user specified paths come first before the LLVM install directory

I don’t think we should add such a warning. I get your frustration, but the specific case where you get hit by this is fairly specific, so I think the warning would add more confusion than what value it adds:

  • The issue only occurs if you build lld and the runtimes separately, into different target roots. In most toolchains I assemble, I build them in separate steps, but into the same root directory.
  • The issue only occurs if the lld root directory also has a separate copy of the runtime you’re overriding. If the directory where lld is installed doesn’t have a copy of the runtime you’re, this shouldn’t be an issue.
  • If we fix lld according to the suggestion, to prefer user provided paths above the ones next to the lld binary, the warning wouldn’t hold up anymore either

I think what you’re getting at, is that since lld has got this logic, to automatically pick up libraries in the lld installation directories - that would work better than libraries installed elsewhere. But if you install libraries elsewhere, you’re expected to point the linker to them somehow anyway - and then they should work equally well. The confusion only arises when the linker prefers the implicitly detected paths over the ones provided on the command line, no?

I’m imagining an issue that I think isn’t as specific as you are describing. If the user takes the following steps they will be hit by this issue:

  1. Install a release version LLVM using the Windows installer
  2. Build LLVM from git with -DLLVM_ENABLE_RUNTIMES=libc++;compiler-rt but without lld.
  3. Compile something with clang++ -stdlib=libc++

In this case, the version of libc++ from the newly built LLVM will be used for header files, but the version of libc++ from the release LLVM from the installer will be used to link against.

It’s not necessarily about building LLD and the runtimes separately, but rather building a new version of LLVM on a system with an older LLVM already installed, but not building LLD in the new LLVM. This is why I’d prefer a warning when building the new LLVM that the correct libraries will not be selected to link in this case. This case won’t be fixed by changing the order to prefer user provided paths, because clang doesn’t pass the paths to the linker.

Ok, fair enough - that’s arguably not all that obscure. I’m still not very keen on adding this warning (I mostly fear that it would be giving false positives in a large number of cases), but I hear you that this quite valid setup can lead to surprises.

In MSVC style environments, the most common usage pattern would be to first compile an object file, clang++ -c -stdlib=libc++ and then later link it by directly invoking lld-link. In that case, I think it’s perhaps marginally more understandable why that second step won’t end up using libraries colocated with the clang++ that was used in a different install directory - do you agree?

But for a plain clang++ -stdlib=libc++ main.cpp -o main.exe, that’s indeed less obvious.

(Side note - overall, I think -stdlib=libc++ isn’t entirely fully set up for MSVC environments yet, there has been some progress on that on and off during the last year.)

As such an invocation, where the clang driver invokes the linker, that does pass the relevant library paths to the linker in non-MSVC environments, so I guess one could consider having it pass it here as well. I’m a bit undecided on whether it’s a good thing or not, though. It makes these cases (with multiple mixed tool installations) less ambiguous, that’s good. But on the other hand, since the primary use case is to invoke lld-link directly, it would lead to a bit more inconsistency between how things work, when doing linking via the compiler driver vs when linking by invoking lld-link directly.

I perhaps shouldn’t have mentioned libc++ as actually, even in the non-libc++ case we depend on compiler-rt and that has the same issue. Admittedly in that case there isn’t the “header vs library” case, but there’s still a “what clang/crt expect vs what they actually get” case.

Is there a reason we can’t embed the library paths clang wants to search for crt/libc++ for in the object file as linker directives instead? That would seem to solve the issue for both lld-link.exe and indeed MSVC’s link.exe as they’d find the correct path in each case, but I admit that I could easily be missing something obvious.

In principle I’d expect lld-link.exe and link.exe to do exactly the same thing with the same invocation and that certainly isn’t the case at the moment as lld-link.exe adds paths that link.exe wouldn’t search. Having clang add those extra directories to the object file as directives would make the two commands behave the same, as I understand it

I’m pretty sure we considered that alternative as well, but didn’t go with that as it felt like a worse trade-off. Off the top of my head, drawbacks with that approach:

  • lld-link currently doesn’t accept -libpath: options in embedded object file directives, which presumably means that link.exe doesn’t either. So one can’t embed -libpath:/path/to/clang/install/lib -defaultlib:libc++.lib, but would have to embed -defaultlib:/path/to/clang/install/lib/libc++.lib instead, which is arguably more brittle. (The former could find things in a different libpath if one is specified manually, if the embedded path doesn’t exist, the latter form can’t.)
  • It makes compiler outputs differ based on the filesystem layout of the build machine
  • It ties object files to exactly the toolchain install on the machine where they were built
  • Unclear how it would work with e.g. distributed compilation

It does look like link.exe doesn’t accept libpath directives in object files, that’s unfortunate.
I would point out for the latter points though that compiler outputs already differ based on the filesystem layout of the build machine due to the way we’re currently doing things.

I think we’re getting slightly sidetracked from the core issue I posted above though (and that’s mostly my fault!)
Would we be able as at least a first step to swap the order of lld’s library path search to search user provided directories first over the LLVM install’s directories? That would at least fix the issue we’re seeing in production so it would be good to have that in before the LLVM 18 fork.

In general, that shouldn’t be the case (we do have options for e.g. remapping paths to be agnostic of the absolute path, for debug info and for __FILE__ preprocessor directives and such). But anyway, that’s a sidetrack and irrelevant to the rest of this.

Yeah, I think this would be straightforward to do. I can post a patch to do this, but I was kind of waiting for an agreement about that direction, from any of the others who were involved in setting this up before doing that.

I’m pressed for time so I didn’t read closely, but reordering the search paths sounds good to me. It seems better to give users control in this case.

Thanks - I got this reviewed in [LLD] [COFF] Prefer paths specified with -libpath: over toolchain paths by mstorsjo · Pull Request #78039 · llvm/llvm-project · GitHub and merged in [LLD] [COFF] Prefer paths specified with -libpath: over toolchain pat… · llvm/llvm-project@92126ca · GitHub.