[RFC] HMake for LLVM

Hi. I would like to request the review for this patch Added partial support for compiling C++20 modules and header-units without scanning. by HassanSajjad-302 · Pull Request #147682 · llvm/llvm-project · GitHub

And like to propose my software HMake. HMake supports C++20 header-units and modules using a pioneering IPC based compilation. Compiling a large code-base like LLVM with header-units can help improve the compiler itself besides the 10x faster compilation.

Soon, HMake will support ISOCPP SG15 List: Re: #include to modules transition and ISOCPP SG15 List: [Modules] [P3057] Two finer-grained compilation models for named modules

This reduces recompilations a lot and also makes the module adoption at scale possible (very difficult if not impossible). LLVM will be able to switch to modules very very quickly with these. Like hu, this will help improve the quality of implementation as well.

I am available for a desktop sharing presentation on my software and pull-request. This might be a quicker way to familiarize.

:cross_mark: This RFC was rejected in this message

As the maintainer, I have big concern for maintaining this, especially I can’t understand it in detail in limited time. I do suggest you to try to make it like a plugin and promote it to the community. If it shows there are a lot of people who is interested in it, we can have more passion for it.

If it does not get any traction, you can just remove the IPC2978 folder and remove the code. I have added 1-2 new functions in clang code. Other modifications are mostly just checking whether the N2978::managerCompiler is set and do things in that or in else. And one of checking whether the /noScanIPC flag is set to set the N2978::managerCompiler. Merging solves the chicken and egg problem. As otherwise someone using HMake will be in doubt about the implementation. Hence less users and hence less traction. Once this gets merged, I plan to reach out to boost and other mono-repos.

In my library, there is a frontend and backend API. Most of the stuff is backend. You can just trust it as the ClangTest.cpp that I added works. The frontend API is the public API of IPCManagerCompiler class + makeManagerCompiler() function + managerCompiler symbol.

These form a total of 9 symbols. Out of these, only 5 are currently being used. Other 2 are needed for shared memory files. While the other 2 are optional. I request for the implementation as I found shared-memory pcm files difficult to add.

N2978::managerCompiler (This pointer is set in cc1_main function if the -noScanIPC flag is provided using N2978::makeManagerCompiler()).

N297::makeManagerCompiler().

IPCManagerCompiler::findResponse(). Whenever the compiler parses an include or import statement it uses this function. Also used when the compiler deserializes a module. This will automatically fetch the request if it does not exist in cache.IPCManagerCompiler::lastMessage. This variable is used to write the compiler output instead of outputting it to console which is not allowed.

IPCManagerCompiler::sendLastMessage. When compilation finishes or an error occurs, this is used to send the last message. Used in ParseAST() once object-file is generated.

I request you to please do run the boost-example GitHub - HassanSajjad-302/HMake: C++ build system that uses C++ for build configuration. .

Reading documentation will be effective in getting the details of how HMake can do header-units while no others can but for now, you can instead do the following steps.

Step - By - Step

  1. You need Windows

  2. You need latest version of VS2022.

  3. You need to compile my patch https://github.com/llvm/llvm-project/pull/147682 or you can instead use the pre-compiled binary https://drive.google.com/file/d/1agPjaVW65Ae10yUeuuqEGlh7WiigEiqg/view

  4. Edit the last line of CppCompilerFeatures::initialize() to point to this fork

  5. After downloading my software, you need to build it with CMake in Release mode with the latest visual-studio 2022.

  6. Add the directory of the produced hhelper hbuild and htools in the path variable. I use clion and I add the cmake-build-release.

  7. Run htools with admin permissions. It creates C:\Program Files (x86)\HMake\toolsCache.json file.

  8. Download and un-zip from https://www.boost.org/releases/1.88.0/

  9. cd to the boost-1.88.0

  10. copy the contents of HMake/Projects/Boost/* to here ( ArrayDeclarations.hpp and hmake.cpp)

  11. mkdir Build

  12. cd Build

  13. hhelper

  14. hhelper

  15. cd huBig-d

  16. ptime hbuild

  17. cd ../conventional-d

  18. ptime hbuild

On my system conventional-d time was 2.3x higher than huBig-d time.

I have attached an output of this. I have also attached 2 images that show the 10x faster rebuild speeds and the colored output. To test rebuild, I touched a random boost/libs/hash2/Tests/ripemd.cpp. HMake is 2x faster than Ninja.

output.txt (245.5 KB)

LLVM due to its huge size will benefit >10x in both clean-build and re-build. Due to proper hu spport, there is no need for this [RFC] Extensions to export macros/(preprocessor states) for C++20 modules

With HMake, there will be no need to build again-and-again for clangd support, if the LLVM moves to modules New Compilation‐Database Proposal · HassanSajjad-302/HMake Wiki · GitHub . This also ensures full IDE support. This was a reply to C++20 Modules Support in Clangd . Thanks to the article, I believe is easy to add as-well.

I plan to add support for a module-map file, in which user defines which files are header-files and which are header-units in a directory. With this, the LLVM hmake.cpp would be simple and less than 1000 c++ lines and it can be completed in less than 2 months from now. Without the intermediate step of header-units, the direct switch to modules is extremely complicated. The boost effort is struggling e.g.

Please comment. I would like to mention the CMake integration maintainers @petrhosek and @Ericson2314 for comment. Please do run the Boost-Example, please. Ultra-Fast build is an experience of a kind. Thanks.

Please see Added partial support for compiling C++20 modules and header-units without scanning. by HassanSajjad-302 · Pull Request #147682 · llvm/llvm-project · GitHub.

Just wanted to share an update. Current status · Issue #2 · HassanSajjad-302/HMake · GitHub

We discussed this RFC in the Clang area team meeting, and we think it needs a more substantial design proposal if you want to move forward with upstreaming your PR.

I think there is actually a lot of shared project interest in having an IPC system:

  • Lit tests suite process creation time is driving interest in a daemonized, buildozer-like solution for running LLVM tool invocations
  • clangd uses grpc for IPC/RPC. gRPC is a portable, off-the-shelf solution. The PR uses Windows IPC APIs directly, and we aim to be a portable project. Have you considered gRPC, and weighed the tradeoffs? clangd is sort of a PCH-generation daemon, presumably similar in structure to what you’ve proposed.
  • Distributed ThinLTO is sort of doing the opposite, it is taking jobs that normally would run as threads and farming them out as separate process invocations that can be run elsewhere. Is there some potential for sharing design work here?

My main concern with any IPC system is that we want to keep the core of the compiler functional (as in “functional programming”) and deterministic. This is one of the major flaws of the MSVC type server PDB system (i.e. the /Zi flag as opposed to /Z7), which uses IPC to assign type IDs online, making its compiler output inherently non-deterministic.

I could easily see us merging some kind of scheduler / executor in LLVM that schedules compilation work if it is very careful about state sharing, but if the proposal is to reuse cached AST nodes in memory, I don’t think the project should take on that maintenance burden, speaking as a former intern who proposed a very similar project 16 years ago.

HMake can have great support for lit testing using IPC. I only have high-level idea of lit. To support IPC based testing, HMake needs the list of tests and the list of all the steps of these tests. For this, HMake can do IPC with the lit tool itself to get this data in the specified data format. If this is a slow process, then the work can be distributed into multiple processes. As soon as HMake receives this data, HMake can start executing the tests. For these it will launch processes like opt and llc on demand. These binaries will not exit and instead write to stdout followed by a 32-byte delimiter. Build-system will read this info and then execute the next test step(if any). HMake will also mark the opt or llc process as unused. The next step of the test could be FileCheck. HMake will launch FileCheck on demand and on its exit, it will also mark this process as unused. Now, for the next test, we can use these cached process by writing on stdin. This way if we have max-concurrent-process for HMake as 32 (by-default the max hardware threads), we will never have more than 32 launches of opt, llc or FileCheck. It could be less infact if these are exiting very quickly.

I am working on supporting the IPC approach using stdout followed by a 32-byte delimiter and stdin instead of separate pipe and socket on Windows and Linux respectively. This did not cross my mind before.

If you have anything in mind, please share. HMake might support it. I have little context here.

I don’t think the way HMake deploys IPC based compilation affects the functional nature or determinism. HMake clang compilation will be deterministic, provable by a similar build-target like stage3 in the LLVM ninja build-system.

If I understood this comment correctly, HMake does not do the shared memory sharing of AST nodes. It does that of PCM files to reduce their reads across the processes. Though it could possibly do AST memory sharing as well if there is to be further speed-up by removing the serialization cost.

I am on verge of successfully compiling Clang with C++20 header-units using IPC in HMake and will share an update soon.

Also, if during IPC, a test fails by crashing (diagnosed as eof received). HMake can report that. It can also report the previous tests this process executed, to ensure that it is not the state of previous tests that leaked and resulted in crash.

Most of my comments were more focused on the code changes to Clang that you proposed. Maybe I need to read the PR more deeply, but I saw a hand-coded IPC system, as you put it, using Windows APIs writing to stdout, and I think what I’d want to see is an abstraction that serves other known use cases before we merge support for anything like this into Clang.

I compiled Clang with C++20 header-units and achieved 25% speed-up.
https://claude.ai/share/c459bad8-06b5-45c0-946c-1d05603f8777

It only works on Linux and there are a few warts currently.

To reproduce,

  1. you need to build my patch. Name the build-dir as my-fork. It is hard reference in hmake.cpp.
  2. Modify the last line in ToolsCache::detectAndInitialize to point to the custom clang
  3. Build my project
  4. add release-build-dir in path
  5. Run CMake target htools. It creates /home/toolsCache.json.
  6. Run CMake target Dir-Mapping in llvm-project. It adjusts the include-names in the llvm and clang source-code. A single header-unit or header-file can be included by more than one include-names in current target and its dependents. It will output lines like Replaced #include "CXXRecordDeclDefinitionBits.def" to #include "clang/AST/CXXRecordDeclDefinitionBits.def" in file /home/hassan/Projects/llvm-project/clang/include/clang/AST/DeclCXX.h. Build will fail without this.
  7. Create LLVMBuild in llvm-project.
  8. cd to LLVMBuild
  9. Run hhelper
  10. Run hhelper

Now you are ready to build by running hbuild. But there are a few considerations.
Running hbuild in LLVMBuild will start building both configurations. Running it in
standard or hu will build only the respective one.
First run ulimit -n 4096. As multiple simultaneous processes are launched while HMake waits for a
dependency header-unit or module. You might also run into OOM Killer, especially if you have less than
2GB/thread. In that case, you can run hbuild -hu in the hu configuration. This will take approx 4s
and will build all the hu processes only. Build algorithm will be enhanced to not have indefinite compilation launches if these all are waiting on a single one. I assure you this would be supported without any performance penalty.

Another wart is that you have to build support/LLVMSupportC first
before building hu configuration.
LLVMSupport target has some C and assembly files.
Building these files with header-units IPC does not work yet.
So, I created a phony target LLVMSupportC just with C and assembly files.
In standard configuration, this target is build
and LLVMSupport has a dependency on it.
While in hu configuration LLVMSupport has dependency
on the prebuilt static library of standard/LLVMSupportC.

You might think that these are
lots of warts for such small improvement.
But all of these are can easily be permanently addressed.

While I previously estimated >10x speed-up, I still expect >2.5x for complete LLVM build with hu. Only, 7-8 targets were compiled as header-units. Compilation failed with others. Compilation failed with big-hu as well. Most of this is fixable by small changes. Like header-units can not have non-inline variable definitions. Or, if using big header-units, there could be macro collision. Or it is failing because of a compiler bug. You can modify the following lines or place these somewhere else to try other targets with header-units.

    config.assign(TreatHUAsHeaderFile::YES);
    config.assign(BigHeaderUnit::NO);

I lack time so did not investigate. And neither did I want to make any changes in the source-code except the header-includes. A total of 608 PCM and a total of 2608 .o files are compiled.

I think with C++20 modules, the spped-up could be >3x.

@ChuanqiXu mentioned in his blog

Some have argued that C++20 Modules are nothing more than standardized Precompiled Headers (PCH) or standardized Clang Header Modules. This is incorrect. PCH or Clang Header Modules reduce compile times by avoiding repetitive pre-processing and parsing across different TUs.

C++20 Modules go a step further by also preventing repetitive optimization and compilation of the same declarations in the compiler backend. For many projects, backend optimization and code generation are the main sources of long compilation times.

HMake is the only build-system that can do

  1. #include to C++20 header-unit transition without source-code changes needed(as demonstrated).
  2. 2-phase compilation of C++20 modules
  3. #include to C++20 modules transition without the immediate source-code changes needed in the consumers, thus avoiding the macro-mess (I would say impossible otherwise).

The hmake.cpp does not need any change to use C++20 modules. You can actually convert a header-file to C++20 module and it will work. However, you will have to replace all references of # include with import.

The most important metric in the benchmark I shared is I think “Percent of CPU this job got” as this reflects on the ability of the HMake to fully utilize all cores. You can see that it is approx Ninja level in both standard and hu configuration. I assure that HMake will beat Ninja in this as some optimizations are still pending including the one of directly invoking the clang compiler instead of the driver first. HMake can do these as it has much more context.

Based on this I think even if someone writes all the dependency info by hand, it would only be 1-2% faster than the IPC based approach. Also, it should be 7-8% faster than the scanning (every scanning is a new process). Also, HMake is at-least 4-5x faster than Ninja in zero-target build speeds. It consumes lesser memory as-well

HMake will have integrated support for distributed building like Bazel and host or remote caching like sccache/ccache. And it would work with C++20 header-units IPC. There is no reason for it to not to.

Please see Example for how HMake can support IPC based lit testing. I have simplified the IPC API and made it more generic. IMO, it would be trivial (1-2 week work).

I also have this proposal for C++20 modules and hu integration. This is very scalable, performant, extensible and easy to support for the 3 of HMake, IDE and the Language Server.

I compiled Clang relatively quickly with my software while I was struggling with Boost because LLVM follows good conventions. e.g. .h file does not depend on external non-configuration macros. Only .def and .inc do. Both of which are included as header-files while only .h are compiled as header-units. Similarly, LLVM uses same compile-command for all of its files’ compilation. I compiled all source including that of code-gen targets. Only the code-gen part itself is missing. I also compared the commands executed between Ninja ninja clang -nv and hbuild hbuild -n to check for similarity.

To support IPC compilation, HMake needs more information compared to the standard approach. e.g. all the include-names used in the files and to which files these map to. This information is cached in a key-map (header-name mapped to file-path(Node*)) and then used at build-time. adding same header-name or module-name twice is an error. File is also stored in a set and adding it again is an error. File must be uniquely owned by one of a target or its dependencies. This is a shield against symbol de-duplication problems.

Because now a file can not be a header-file in one dependency target with one include-name and a header-unit in another dependency target with another include-name. Only place where symbol de-duplication will be needed is when a header-file is being included by 2 dependency header-units. To prevent against this promote more and more .inc and .def files to header-units instead.

While HMake can send a header-unit for #include, it has another feature of replacing multiple small header-files with one big hu. This hu is called big-hu and these header-files are called composing headers. Suppose std.hpp has iostream.hpp. Now when the user specifies iostream.hpp, HMake at config-time ensures that only one target has iostream and also only that target has the iostream file. Now, while compiling the std.hpp, HMake will send the iostream file but for any other iostream request, HMake will send the std.pcm instead.

HMake can do this with modules as well. i.e. send std instead of iostream to all the dependents. During transition, once all references of iostream are replaced by std, either the iostream can be removed from composing headers or it can be treated as private at build-time. So, at config-time HMake still checks for uniqueness of iostream include-name and iostream file among the dependents (so 2 different modules not have composing iostream) but at build-time #include iostream will be a not found error.

The first step to #include to module transition is HandleHeaderIncludeOrImport function having the ability to parse a module plus a header-file (consisting of macros squeezed of that module) instead of a header-file for the include-name. This is needed only under ipc flag.

One more wart is that in this approach #include_next is not supported. So, HMake compiles the standard header-units using traditional way. It compiles 2 because macros in link.h and zlib.h collide with macros in LLVM as standard hu is compiled as one big-hu. This aspect is little random and needs more definition.

To complete the above features, I also seek financial sponsorship. I have been working on this project for a long time. You can see the scale of this.

For LLVM, with concerted effort, all of the above can be completed in next 4 months resulting in >3x speed-up in compilation and testing. While previously, I have been off in my estimates, I am in much better position now. I think HMake should be announced as replacement for the CMake and Bazel in LLVM for wide-scale interest and detailed review. We can always roll back the transition.

Please share if you were able to reproduce the results.

I would love to do a video chat presentation of my software which I think could be the quickest way to get familiar with my software.

I have simplified a lot my pull-request and my build-system both. Previously HMake was multi-threaded. But I think it was getting in the way of adoption and benefit was not worth the price. It was ideal under ideal conditions where everything else was a library and is loaded and used with thread-synchronization. Now, we loose this speed but we get the advantage of process security and vastly enhanced simplicity. Also a child process can do multi-threading. the build-system will reserve some threads for it so there is no over-saturation.

Pull-request is simplified as now HMake uses stdin and stdout for IPC instead of named-pipes on Windows and sockets on Linux.

A lot of what you have read in above message is for protection against symbol de-duplication bugs. With HMake, you can’t even mistakenly, consume a file as 2 of Header-File, Header-Unit, Module. This means performance is never compromised. This forms the basis for scalable widespread module adoption in projects. While the API might be a bit difficult / underdeveloped today. With time and with some module-map files, it will be as easy as the include-directories used today.

This is one more feature that only HMake can achieve.

Hello everyone. Please share if you were able to reproduce. what are your thoughts?

Please share if some error happened. I look forward to your thoughts.

To reproduce,

  1. you need to build my patch. Name the build-dir as my-fork. It is hard reference in hmake.cpp. Configure it with cmake ../ -GNinja -DLLVM_ENABLE_PROJECTS=clang -DLLVM_TARGETS_TO_BUILD=X86 -DCMAKE_CXX_STANDARD=20 -DCMAKE_MAKE_PROGRAM=/usr/bin/ninja -DCMAKE_DISABLE_PRECOMPILE_HEADERS=ON. Configure it with last option if you want to benchmark without precompiled headers.
  2. Modify the last line in ToolsCache::detectAndInitialize to point to the custom clang
  3. Build my project
  4. add release-build-dir in path
  5. Run CMake target htools. It creates /home/toolsCache.json.
  6. Run CMake target Dir-Mapping in llvm-project. It adjusts the include-names in the llvm and clang source-code. A single header-unit or header-file can not be included by more than one include-names in current target and its dependents. It will output lines like Replaced #include "CXXRecordDeclDefinitionBits.def" to #include "clang/AST/CXXRecordDeclDefinitionBits.def" in file /home/hassan/Projects/llvm-project/clang/include/clang/AST/DeclCXX.h. Build will fail without this.
  7. Copy hmake.cpp from HMake/Projects/LLVMProject/hmake.cpp to llvm-project/hmake.cpp (I missed this in previous mesage).
  8. Create LLVMBuild in llvm-project.
  9. cd to LLVMBuild
  10. Run hhelper
  11. Run hhelper

@rnk @ChuanqiXu

In that latest commit, I have added 3 new features resolving one wart and fixing one work-around and simplifying hmake.cpp a lot (reduced lines by >100).

HMake now no longer launches unlimited number of process while waiting for a dependency header-unit to compile. This was resulting in OOM Killer activation. maxSimultaneousProcessDesired is set to std::thread::max_hardware_concurrency() * 32. Once total number of launched processes reaches to this level, the build-system switches to single-process mode. Now, build-system will launch only one process a time and not in bulk. This is to ensure that build-system does not stall in situation where even after reaching the maxSimultaneousProcessDesired, there is no single active process and all of them are waiting for a header-unit / module which has not yet been scheduled. In practice, this should never happen as HMake schedules the dependency CppTarget header-units before the dependent’s.

The result is exact same performance and no OOM killer activation and no need to run hbuild -hu, thus fixing this wart. This option was added as a hack around this and will be removed.

HMake now reports the completed tasks and tasks left while compilation. This is not completely accurate as new dependency relationship discoveries increases the count.

Renamed N2978 to P2978 as suggested in the GitHub issue.

As previously mentioned, in HMake IPC, you could only import a header-unit from the target itself or the dependencies and it is checked at compile-time that no 2 targets provide the same include-name or the same file. To make this work, lots of phony targets were added in hmake.cpp and it was still not working as multiple targets are including header-files from their multiple dependent targets ( the dependency relationship being same as in the CMakeLists of these targets). To workaround, I was using a temporary hack of allowing header-units even from dependent targets which was causing some rebuild bugs as well.

Now, I have made this a feature with proper checks at configure-time. if UseConfigurationScope::YES is set for the Configuration, then C++20 modules can import a header-unit from any of the CppTarget in that respective Configuration. At configure-time, the name collision is also checked with the whole Configuration scope. The downside is that 2 unrelated CppTarget can not provide the same module-name or include-name.

There was just one collision of include-name TableGenBackends.h in both llvm/utils/TableGen and clang/utils/TableGen. I disabled the clangTableGen exe as it is unused for now anyway.

Another downside is, as mentioned above, HMake schedules dependency CppTarget header-units first before the dependent’s to reduce number of waiting processes. Now, this optimization is partially circumvented.

So, I removed the additional phony targets and additional dependency relationships besides those in the CMakeLists.txt. Projects/LLVM/hmake.cpp is so much more clean now. This approach is extensively tested / supported. And achieves the exact performance for LLVM. Nonetheless, UseConfigurationScope::NO remains the ideal approach.

Following are the results for latest LLVM https://claude.ai/share/e49bb9c0-1660-44e9-b690-f8f95dce027b. While hu is 25% slower than latest (much better now) PCH, please remember that very very few CppTarget are being compiled as hu for now.

What should I do next? I would love to design LIT testing IPC protocol and do Code Generation steps removing CMake build-directory dependency but I have very little knowledge of these and would appreciate collaboration.

I request you to find a bug in my software. I am confident you won’t.

We discussed this RFC again during today’s Clang Area Team meeting and we believe there is not consensus to proceed with this proposal, largely due to a lack of community engagement. We think there are interesting ideas worth exploring in this space, so we’d welcome a similar RFC with more focus on the design of the system in the future.

With the latest commit, HMake now has significantly improved documentation. Please give it another look.

New feature: HMake can now warn when the same header file is included by two different dependencies (header units or modules), enabling guaranteed zero symbol deduplication. This matters both for compilation performance and for correctness — deduplication bugs are notoriously subtle. The back-end for this feature is complete; the user-facing API is still in progress.

On the architecture: HMake has a clean separation between its core layer and the C++ layer built on top of it. The C++ layer is more complex due to the breadth of features it supports, but the user-facing API is intentionally simple and can be simplified further. Importantly, LIT testing support only requires the core layer.

On the IPC protocol: In this commit I have worked on the IPC protocol. My understanding of the LIT internals is still limited, so please treat this as a first draft and share any feedback.

On reproducibility: A comment by @dblaikie raised the question of reproducibility on failure. HMake addresses this by writing a replay file containing all inputs sent to the tool up to the point of failure. From the protocol design file:

/// Failure handling:
/// If a test fails during IPC execution, HMake re-runs it in a fresh
/// process (outside IPC) using llvm-lit. If it fails again, it is a genuine test failure.
/// If it passes in isolation, it is a state-leakage failure caused by a
/// prior test. In that case, HMake dumps a replay file containing all inputs
/// the tool received up to that point, so you can reproduce and debug the
/// exact sequence that caused the leak. struct LitTest

Debugging state-leakage failures is straightforward because the IPC protocol requires tools to support an array of inputs. From the same file:

/// Contains exactly one Response during normal IPC operation.

/// If a test fails only under IPC (passes when re-run in isolation), it
/// indicates a state-leakage bug in the tool. In that case, HMake kills the
/// daemon process (as it is now in a bad state) and writes a binary replay
/// file containing all Response messages that were sent to that process
/// in order, serialized in the same format as originally used in IPC.
/// The tool can then be pointed at this replay file directly to reproduce the
/// failure deterministically.

/// To support this, the tool must always parse this field as a loop regardless
/// of whether it is running under IPC or from a replay file — the only
/// difference is the number of entries (1 in IPC, N in the replay file).
/// vector responses;

After killing a corrupt process, if that tool type is still needed and the pool is below capacity, HMake will launch a replacement. The number of live processes per tool is capped at std::thread::hardware_concurrency(), and processes are launched on demand.

On process reduction: The current IPC design assumes each llc or opt invocation is followed by exactly one FileCheck. If that holds, then as @nikic notes there are approximately 45k such invocations — meaning this approach eliminates around 90k process launches.

There is also an opportunity to eliminate a redundant file read. Currently, a .ll file is read once by lit and again by llc. With this approach, lit reads the file once, composes the data in the required format, and sends it to HMake over IPC — removing the second read entirely.

Both the lit-side changes (producing data in the required format) and the tool-side changes (IPC support) can be opt-in, allowing a gradual transition that keeps the existing approach working alongside the new one.

On determinism: The process scheduling is inherently non-deterministic, so test ordering across daemon processes will vary between runs. However, the failure handling described above — re-running failing tests in isolation and writing replay files — effectively addresses any issues this introduces.

On scope: If the C++ build system aspects of HMake are not of interest but this IPC infrastructure is, HMake could be slimmed down and contributed under llvm/utils, where it would be invocable with a single command much like lit today.

If this can be completed by Saturday, I can share reproducible benchmark results by Monday. In the meantime, I plan to remove HMake’s dependency on the CMake build directory for the Clang build, which will eliminate another rough edge in the setup.

I would also like to request a meeting with interested LLVM team members to present HMake in detail. I estimate 40–50 minutes would be sufficient to cover the architecture, the IPC protocol, and the C++20 modules work. Please let me know if there is interest and what time works.

edit: formatting.

The concern of maintainability still exists. This is still a big PR for a big new feature. We have to be careful.

And from my point of view, we need to see significant benefit from the new model to be willing to take the additional burden.

Here, the benefit is not only about comparing header units to normal builds without any serialization. Since we already have PCH and clang header modules. And again, I believe header units in most cases have exactly the same performance improvements as clang (explicit) header modules. So what I want to see, or what makes me more interesting is, with the new model (the IPC approach you proposed), how many improvement we can get comparing to current clang explicit header modules.

I think you can try to build LLVM (or any project big enough) with clang explicit header modules, and build it again with your model, and show us the result. I feel this is a more plausible path.

For clang explicit header modules, I remember @mpark mentioned me that Meta’s Buck’s implementation is open sourced, but they never said about it. (Meta uses header units as far as I know, but actually it is clang explicit header modules from the implementation’s point of view). And I am not sure if bazel open sourced their implementation for clang explicit modules. I am not sure who is the person I can ping here, properly, maybe @boomanaiden154-1 ?

1 Like

I’m not too familiar with the details, but from what I can tell, there is support for clang header modules, at least looking at the features supported by bazel directly (seems to be implemented in core bazel rather than rules_cc like I would have thought).

There’s also some support for C++20 modules.

There is currently no project that is using explicit clang header modules / header-units, the way HMake makes it possible. HMake allows header-units for every single header-file. To do this in any other build-system, you would need to maintain mod_map file for [every single file * every single configuration] which is impossibly difficult. You can see the MS Office header-units blogs. They use only for immutable core libraries. Conventional model limitations impact whether it be clang header-modules or header-units. IPC model has been discussed often in the past, however, HMake is the first ever implementation of it. Just search for word server in these links.

In my latest commit, I built more libraries with c++20 header-units. A total of 2513 header-units were compiled, an increase from previously >500. I achieved 18% faster build compared to the pre-compiled headers, link. Please do reproduce.

However, compiling the remaining (approx 1000) either failed or cause OOM or were too slow. Header-units of target cppStaticAnalyzerCore takes more than minute each to compile. You can reproduce by editing this line from YES to NO. You can confirm from htop that it is clang process that it taking 100% cpu and takes more than a minute to finish each. Otherwise, the rest of header-units compile very very quickly.
In latest commit, I have added feature to be able to reproduce with a script. After build completion, HMake will create a script. Now, to reproduce on some other system, you would just need to create the root build-dir and run the script. This script will have commands to automatically create the sub build-dirs, compile dependencies(potentially 100s) and then build the faulting module/hu providing it all the dependencies on the command-line upto the point of crash.
However, I did not use this and filed bug report since the Clang command-line has limitations for header-file to header-unit translation.

From your message, I observe that you have favorable opinion of Clang header-modules support in buck/bazel. But this is incomplete as @boomanaiden154-1 mentioned and Support C++20 modules · Issue #4005 · bazelbuild/bazel · GitHub. As I mentioned, bazel/buck supports C++20 modules / Clang header-modules using mod-map files which was rejected by CMake due to its maintenance cost. CMake went with runtime dependency generation step that itself could not support C++20 header-units.

HMake benefits include (not limited to), I repeat:

  1. #include to module translation.
  2. #include to hu translation.
  3. 2-phase module compilation.
  4. Guaranteed De-duplication.
  5. Zero build-configuration changes needed to move to modules (almost literally). Zero immediate changes needed in the consumers. You just convert a file to module and it works.

Because minimal build-system changes are needed to move to modules, minimal build-system changes will be needed to rollback a file in-case of compiler-bug, just like I roll-backed from >2600 to >2500, once I observed very slow hu compilations. The importance of this feedback loop can not be understated.
For even stronger demonstration (>4x build speed-up, compiling rest of 1000 files as header-units, enabling DeDuplicationWarnings::YES, clang quality of implementation improvements), I seek your help but your help is contingent on stronger demonstration. Please see chicken and egg cycle.

In latest commit, I added an example of IPC. This proposal is scheduled for discussion in upcoming infra meeting. By then, I hope to prepare a significant speed-up demonstration by compiling majority of tests using IPC.

Merging the PR would be a great first step, and I am happy to work through any specific maintainability concerns you have.

This is not true based on my knowledge. I believe both Google and Meta use them.

HMake allows header-units for every single header-file.

I am not talking about HMake or CMake or any build systems. Again, the only thing I cared here is, the advantage of IPC with the current explicit clang header modules VS the current explicit clang header modules. I just want to see the number to decide if we want to take the new burden. Code are debt. We need to see significant improvements to decide if we want the new debt.