[RFC] Introducing memory(fresh) to LLVM IR

Related Threads:

  • [llvm-dev] [RFC] How to manifest information in LLVM-IR, or, revisiting llvm.assume
  • Discourse: @llvm.assume blocks optimization (#71609)

1. Summary

This RFC proposes a new memory location category, fresh, for the LLVM IR memory(...) function attribute. The fresh location models a call-private, non-aliasing side-effect sink: effects written to fresh are guaranteed not to interact with any other memory location, including other fresh accesses from different calls.

Only none (the default) and write are meaningful for fresh; read and readwrite are semantically invalid and will be rejected.

The primary motivation is to unblock optimizations that are currently pessimized by @llvm.assume and its outlined counterpart "assume_fn" , as extensively discussed in Discourse #71609.

2 Prior Art: assume_mem

Johannes Doerfert proposed memory(assume_mem: write) as an intermediate step in Discourse #71609. This isolates assume side effects from regular memory, malloc, and global state. However, assume_mem is still a shared abstract location: multiple @llvm.assume calls all write to the same assume_mem, so the compiler must still reason about ordering and interference between them.

We need a stronger guarantee.

3. Design: memory(fresh)

3.1 Semantics

fresh introduces a new IRMemLocation that behaves unlike any existing location category (argmem , inaccessiblemem , errnomem , target_mem , etc.).

Core axiom:

Every call site has its own private fresh location. Two accesses to fresh from different call sites never alias, regardless of the access type.

This means:

  • fresh: write from Call A does not alias fresh: write from Call B.
  • fresh: write does not alias any other memory location (including argmem, inaccessiblemem, Other, etc.).
  • There is no “read” from a fresh location that yields data written by another call.

3.2 Legal and Illegal Forms

Because fresh is a private sink that is never observable by other operations, reading from it is meaningless :

Form Status Rationale
memory(fresh: none) Legal (default) No fresh access.
memory(fresh: write) Legal Side effects sink into a private hole.
memory(fresh: read) Iilegal No prior data exists in a fresh location; reading is undefined/self-referential.
memory(fresh: readwrite) Iilegal The read component is meaningless.

Parser/verifier will reject fresh: read and fresh: readwrite .

3.3 Alias Analysis Rules

The implementation introduces specialized handling within BasicAliasAnalysis (and potentially AAResults) to exploit the independence guarantee of fresh memory:

When analyzing two memory locations, if both LocA and LocB are marked as fresh, the alias analysis must report AliasResult::NoAlias unconditionally. This holds true even when both locations originate from different invocations of the same function, as each fresh allocation denotes a distinct, non-overlapping region of memory.

This behavior represents a fundamental departure from assume_mem. Under assume_mem, multiple calls may still reference a shared abstract memory location, leaving room for potential interference between them. In contrast, fresh provides a hard independence guarantee: any two fresh locations are strictly disjoint, enabling the alias analyzer to definitively rule out aliasing without further context or analysis.

Can you please explain in more detail what specific problem you are trying to solve here? When we talk about assumes interfering with optimizations, it’s usually either because they introduce extra uses, or because they keep control flow live. I can’t say I’ve ever seen the memory effects of the assume itself meaningfully contribute to optimization issues.

It would be easy enough to teach BasicAA that two assumes are reorderable without a special memory location kind (see the experimental guard handling for where this would happen), I just struggle to think of a case where this would actually be useful.

LLVM currently lacks a way to express “writes that are semantically meaningful for optimization but provably unobservable at runtime”, which forces llvm.assume to rely on intrinsic-specific handling and blocks a sound design for assume_fn.

Marking llvm.assume as having a “memory write” is a hack; we do it because we don’t want to audit a bunch of places to figure out how they should handle “this doesn’t have any side-effects, but please don’t erase it”. It’s not actually touching memory in any practical sense; modeling it as a memory attribute is just confusing. (If the assumption fails, the behavior is undefined… but lots of instructions have undefined behavior without writing to memory. Like dividing by zero.)

Sorry, I still don’t understand what specific problem you are trying to solve here. How is this related to assume_fn?

Can you maybe give some IR examples that illustrate what optimizations this is going to enable?