[RFC] Entrypoint cleanup

Our entrypoints are currently a bit of a random mess based on what people remembered to add things to when they implemented a function. There were 17 separate platform-specific entrypoints.txt files (e.g., linux/x86_64, baremetal/arm, gpu/spirv) with a lot of nearly-identical lists between them.

I spent some time (obDisclosure: me and a bot) unifying things into portfolios to see how it would look. Comparing llvm:main...kaladron:entrypoints · llvm/llvm-project · GitHub is the result. Highlights:

  • Granular “Opt-In” Portfolios: Introduce a domain-specific, purely additive portfolio system under libc/config/portfolios/.
    • Sub-portfolios represent specific headers or groups (e.g., cstd/string.txt, posix/unistd.txt, linux/sys_ioctl.txt).

    • Meta-portfolios aggregate these sub-portfolios into broad standards (e.g., cstd.txt, posix.txt, linux.txt, baremetal.txt).

  • Platform Minimalism: Each platform’s entrypoints.txt is reduced to a simple, high-level list of the portfolios it opts into.

So far the portfolios themselved contain a lot of the exceptions that used to exist. My goal would be to first commit with all the exceptions in the porfolios, and then clean them up in subsequent PRs as I confirm which ones were omissions and which ones were intentional. I’d like the portfolios to generally be without many conditionals but clean-up is needed to get there!

I think this will also make it easier to remove some of the conditionals for GPUs that exist, and also make things like FreeBSD and UEFI easier to support without making the cmake files more gnarly than they have to be. I also imagine us being more explicit about what goes into overlay mode versus full build by selecting which portfolios go in - but that is intentionally deferred to later. This only brings the capability to have that conversation.

Please note that the above change is a proof of concept, not a PR. It’s intentionally only minimally cleaned up and verified to show the direction and get feedback rather than submitting.

Tks,
Jeff Bailey

Also as a quick note:

“Showing 81 changed files with 2,569 additions and 11,327 deletions.”

Gives a hint as to the amount of complexity reduction that I’m hoping comes from this =)

2 Likes

Overall, I like the looks of this. Thank you for working on it.

As someone relatively new to the project, I have two questions/suggestions:

  • When adding a new entry point, would there be a default expectation to enable it for all targets using a given “portfolio” (obviously, unless the implementation is architecture-specific or you have other reasons to believe it will not work on a given target)? I think the choice of defaults here is important, as I suspect that a lot of the current divergencies were down to people not wanting to break targets that they cannot test. At least, that has been my worry whenever I was doing this. If contributors are overly cautious in enabling entry points, then the result may not look that clean..
  • What would you say if, instead the portfolio file excluding the entry point based on the target, we moved those exclusions to the target files instead. I don’t know if the exclusion would be done via list(REMOVE_ITEM) or something more elaborate, but I feel that this would be a more principled approach (a default implementation and an override; instead of the portfolio needing to know what all of its users are). And it avoids the situation, where if we have many long-lived exceptions, the in-portfolio approach degenerates into something even messier than what we have now. If the exceptions are listed in the target file, then you limit the complexity to the “odd” targets (and maybe motivate them to do something about it).
1 Like

When adding a new entry point, would there be a default expectation to enable it for all targets using a given “portfolio” (obviously, unless the implementation is architecture-specific or you have other reasons to believe it will not work on a given target)?

I think so, yes, but I’d like feedback on the idea. I can see it being hard for Windows, FreeBSD and other platforms that are less common development environments so this is definitely a place where we need to make sure we understand what’s expected of us as contributors and whether or not there are port owners or something who can help. Or perhaps we can fast-follow with a disable if something fails in post-submit?

What would you say if, instead the portfolio file excluding the entry point based on the target, we moved those exclusions to the target files instead.

Yes. These are currently groups together in the portfolios because it was easier to see them this way. I’d like to commit them like this and then I will move them over to the target files. As we figure these out per-platform. Some of the exclusions can just be removed and some might indicate that the portfolios aren’t the right ones. I’d love feedback from baremetal and GPU folks on these.

Tks,
Jeff Bailey

FWIW, that’s fine by me.

:+1:

This was discussed in today’s monthly meeting: Monthly LLVM libc meeting - #69 by michaelrj-google

I like the idea of creating larger groups of functions which can be combined. That would definitely simplify out build. We do have an existing way to exclude individual functions, using excludes.txt: llvm-project/libc/config/linux/x86_64/exclude.txt at main · llvm/llvm-project · GitHub
(more specifically adding entrypoints to TARGET_LLVMLIBC_REMOVED_ENTRYPOINTS since all the *.txt files in /config are just interpreted as cmake files).

I am all in on reducing duplication. This seems like a good idea overall.

However I agree with this point from the meeting notes:

Using a system like the proposal, it would be easy to have “everything in this group” but possibly more difficult to do “everything in this group except this one function”

I’m not keen on having conditionals in these portfolio files. We will have reduced the duplication, but the fine-grained inclusions/exclusions would still look messy. Messier than it is now.

Perhaps the idea of using exclusion lists is a better solution than the conditionals.

Right - I never got to an answer here that I thought could satisfy everyone, so I didn’t push it further.

In general, I don’t like negative/filter lists because it makes it so much harder to reason about. In principle, what I want is for the desktop versions of the libc to all get all the things without duplicating entrypoints, and to make it easy to opt in to sets of things. But also, preserve the ability to not start from one of those lists and just do your own things. So, like optional macros. But ones that make it easy to say “Hey, we’ve brought up PA-RISC[0], and we can just link in C99, C23, POSIX, Linux and we’ll be along for the ride”

If this isn’t much easier for the people who opt-in with basically zero cost to those who don’t, then I think we’ll have done it wrong.

Tks,
Jeff Bailey

[0] This is not the weirdest hobby I’ve ever had.