[RFC] Compress Intrinsic Name Table

I checked internally at NVIDIA and the size of the driver component that packages LLVM and other shader compiler related code is not a significant contributor to the total driver size, and as such its code size is not a significant concern. If compiling out intrinsics was a less intrusive change, then yes, we would be happy to have it, but if it’s a more involved effort, then the consensus seems to be that for that use case at least a 1MB saving (out of 64MB) is not that important given other driver components are bigger in size.

It also seems that we (at NVIDIA) had in the past attempted to do this by changing the code (remove target intrinsic TD files from Intrinsics.td and #ifdef’ing out common code that uses these target intrinsic IDs) for code size savings, but it created significant maintenance burden for minimal code size gains, so we decided to stop doing this and eat that cost.

So, the takeaway from this discussion seems to be:

  1. Multiple folks have suggested and there has been atleast 1 attempt in the past to compile out target specific intrinsics (excluding this one).
  2. Shipping LLVM with just one or a few targets enabled is a common production use case
  3. Compiling out target intrinsics has logistical issues in terms of test coverage & maintenance and adding new tests.
  4. The ! for code/data size saving is use-case specific. For desktop GPU device drivers, it seems it’s not that much. For embedded use cases this might be significant, given that folks at Apple have attempted this in the past.
  5. There are also questions about LLVM’s API/LangRef spec: as @jrtc27 pointed, do we consider target intrinsics a part of LLVM Core IR and hence always available and usable vs being tied to a target.

Given that (3) and (5) are significant barriers, the effort to compile out intrinsics seems less worthwhile. I’d think it would be good to clarify/get consensus on (5) to be clear if we are not pursuing this due to it being illegal to do (i.e., say that targets intrinsics should be always usable, and what that implies, for ex, can one target start using other target’s intrinsics? may be that already happens today?) or just due to the high logistical cost.

As we keep adding more target to LLVM over time and/or add new intrinsic related features (like pretty printing) or new intrinsics, this may come up again in future so hoping this takeaway can help summarize the status as of now.