I compiled Clang with C++20 header-units and achieved 25% speed-up.
https://claude.ai/share/c459bad8-06b5-45c0-946c-1d05603f8777
It only works on Linux and there are a few warts currently.
To reproduce,
- you need to build my patch. Name the build-dir as
my-fork. It is hard reference in hmake.cpp.
- Modify the last line in
ToolsCache::detectAndInitialize to point to the custom clang
- Build my project
- add release-build-dir in path
- Run CMake target
htools. It creates /home/toolsCache.json.
- Run CMake target
Dir-Mapping in llvm-project. It adjusts the include-names in the llvm and clang source-code. A single header-unit or header-file can be included by more than one include-names in current target and its dependents. It will output lines like Replaced #include "CXXRecordDeclDefinitionBits.def" to #include "clang/AST/CXXRecordDeclDefinitionBits.def" in file /home/hassan/Projects/llvm-project/clang/include/clang/AST/DeclCXX.h. Build will fail without this.
- Create
LLVMBuild in llvm-project.
- cd to
LLVMBuild
- Run
hhelper
- Run
hhelper
Now you are ready to build by running hbuild. But there are a few considerations.
Running hbuild in LLVMBuild will start building both configurations. Running it in
standard or hu will build only the respective one.
First run ulimit -n 4096. As multiple simultaneous processes are launched while HMake waits for a
dependency header-unit or module. You might also run into OOM Killer, especially if you have less than
2GB/thread. In that case, you can run hbuild -hu in the hu configuration. This will take approx 4s
and will build all the hu processes only. Build algorithm will be enhanced to not have indefinite compilation launches if these all are waiting on a single one. I assure you this would be supported without any performance penalty.
Another wart is that you have to build support/LLVMSupportC first
before building hu configuration.
LLVMSupport target has some C and assembly files.
Building these files with header-units IPC does not work yet.
So, I created a phony target LLVMSupportC just with C and assembly files.
In standard configuration, this target is build
and LLVMSupport has a dependency on it.
While in hu configuration LLVMSupport has dependency
on the prebuilt static library of standard/LLVMSupportC.
You might think that these are
lots of warts for such small improvement.
But all of these are can easily be permanently addressed.
While I previously estimated >10x speed-up, I still expect >2.5x for complete LLVM build with hu. Only, 7-8 targets were compiled as header-units. Compilation failed with others. Compilation failed with big-hu as well. Most of this is fixable by small changes. Like header-units can not have non-inline variable definitions. Or, if using big header-units, there could be macro collision. Or it is failing because of a compiler bug. You can modify the following lines or place these somewhere else to try other targets with header-units.
config.assign(TreatHUAsHeaderFile::YES);
config.assign(BigHeaderUnit::NO);
I lack time so did not investigate. And neither did I want to make any changes in the source-code except the header-includes. A total of 608 PCM and a total of 2608 .o files are compiled.
I think with C++20 modules, the spped-up could be >3x.
@ChuanqiXu mentioned in his blog
Some have argued that C++20 Modules are nothing more than standardized Precompiled Headers (PCH) or standardized Clang Header Modules. This is incorrect. PCH or Clang Header Modules reduce compile times by avoiding repetitive pre-processing and parsing across different TUs.
C++20 Modules go a step further by also preventing repetitive optimization and compilation of the same declarations in the compiler backend. For many projects, backend optimization and code generation are the main sources of long compilation times.
HMake is the only build-system that can do
- #include to C++20 header-unit transition without source-code changes needed(as demonstrated).
- 2-phase compilation of C++20 modules
- #include to C++20 modules transition without the immediate source-code changes needed in the consumers, thus avoiding the macro-mess (I would say impossible otherwise).
The hmake.cpp does not need any change to use C++20 modules. You can actually convert a header-file to C++20 module and it will work. However, you will have to replace all references of # include with import.
The most important metric in the benchmark I shared is I think “Percent of CPU this job got” as this reflects on the ability of the HMake to fully utilize all cores. You can see that it is approx Ninja level in both standard and hu configuration. I assure that HMake will beat Ninja in this as some optimizations are still pending including the one of directly invoking the clang compiler instead of the driver first. HMake can do these as it has much more context.
Based on this I think even if someone writes all the dependency info by hand, it would only be 1-2% faster than the IPC based approach. Also, it should be 7-8% faster than the scanning (every scanning is a new process). Also, HMake is at-least 4-5x faster than Ninja in zero-target build speeds. It consumes lesser memory as-well
HMake will have integrated support for distributed building like Bazel and host or remote caching like sccache/ccache. And it would work with C++20 header-units IPC. There is no reason for it to not to.
Please see Example for how HMake can support IPC based lit testing. I have simplified the IPC API and made it more generic. IMO, it would be trivial (1-2 week work).
I also have this proposal for C++20 modules and hu integration. This is very scalable, performant, extensible and easy to support for the 3 of HMake, IDE and the Language Server.
I compiled Clang relatively quickly with my software while I was struggling with Boost because LLVM follows good conventions. e.g. .h file does not depend on external non-configuration macros. Only .def and .inc do. Both of which are included as header-files while only .h are compiled as header-units. Similarly, LLVM uses same compile-command for all of its files’ compilation. I compiled all source including that of code-gen targets. Only the code-gen part itself is missing. I also compared the commands executed between Ninja ninja clang -nv and hbuild hbuild -n to check for similarity.
To support IPC compilation, HMake needs more information compared to the standard approach. e.g. all the include-names used in the files and to which files these map to. This information is cached in a key-map (header-name mapped to file-path(Node*)) and then used at build-time. adding same header-name or module-name twice is an error. File is also stored in a set and adding it again is an error. File must be uniquely owned by one of a target or its dependencies. This is a shield against symbol de-duplication problems.
Because now a file can not be a header-file in one dependency target with one include-name and a header-unit in another dependency target with another include-name. Only place where symbol de-duplication will be needed is when a header-file is being included by 2 dependency header-units. To prevent against this promote more and more .inc and .def files to header-units instead.
While HMake can send a header-unit for #include, it has another feature of replacing multiple small header-files with one big hu. This hu is called big-hu and these header-files are called composing headers. Suppose std.hpp has iostream.hpp. Now when the user specifies iostream.hpp, HMake at config-time ensures that only one target has iostream and also only that target has the iostream file. Now, while compiling the std.hpp, HMake will send the iostream file but for any other iostream request, HMake will send the std.pcm instead.
HMake can do this with modules as well. i.e. send std instead of iostream to all the dependents. During transition, once all references of iostream are replaced by std, either the iostream can be removed from composing headers or it can be treated as private at build-time. So, at config-time HMake still checks for uniqueness of iostream include-name and iostream file among the dependents (so 2 different modules not have composing iostream) but at build-time #include iostream will be a not found error.
The first step to #include to module transition is HandleHeaderIncludeOrImport function having the ability to parse a module plus a header-file (consisting of macros squeezed of that module) instead of a header-file for the include-name. This is needed only under ipc flag.
One more wart is that in this approach #include_next is not supported. So, HMake compiles the standard header-units using traditional way. It compiles 2 because macros in link.h and zlib.h collide with macros in LLVM as standard hu is compiled as one big-hu. This aspect is little random and needs more definition.
To complete the above features, I also seek financial sponsorship. I have been working on this project for a long time. You can see the scale of this.
For LLVM, with concerted effort, all of the above can be completed in next 4 months resulting in >3x speed-up in compilation and testing. While previously, I have been off in my estimates, I am in much better position now. I think HMake should be announced as replacement for the CMake and Bazel in LLVM for wide-scale interest and detailed review. We can always roll back the transition.
Please share if you were able to reproduce the results.
I would love to do a video chat presentation of my software which I think could be the quickest way to get familiar with my software.