[RFC] A Vulkan-style memory model for AMDGPU (and beyond)

In most communication patterns, not all memory accesses performed by a thread need to be exposed to other threads. Even when they do need to be exposed, not all threads may need to observe these memory accesses. We propose a memory model incubated under the AMDGPU target, which is weaker than the LLVM memory model. The AMDGPU memory model allows the user to control how the side-effects of memory accesses are propagated across threads. The implementation can then choose a more efficient mechanism to complete them, such as the cache policy bits in an AMDGPU device.

Although we started with a target-specific draft, we believe that the model itself is useful for other LLVM Targets. To this end, we have deliberately couched our model in terms that are compatible with the Vulkan memory model. We are inviting comments on whether this serves as a subset of the Vulkan memory model expressed in LLVM terms.

The memory model introduces the following concepts that will be familiar to any Vulkan developer:

  • Availability and visibility
  • Scope instances
  • Location order

The following features are notable:

  1. The new model is weaker than the LLVM memory model; the former allows more outcomes than the latter. At the same time, a simple mapping can be used to implement this model in terms of the LLVM memory model. Thus, there exists a safe-by-default implementation that produces executions that are valid in both models.
  2. The new model is built on top of the happens-before order defined by the LLVM memory model. But happens-before by itself is not sufficient to describe its observable effects. Instead, the new model uses availability and visibility to describe how the side-effects of these operations propagate to other threads.
  3. The new model does not change the structure of happens-before, but changes the rules that determine how operations may observe the side-effects of other operations that happen before them.
  4. Metadata is used to weaken existing synchronizing instructions like atomic operations and fences. Additional AMDGPU intrinsics are added to describe instructions that cannot be expressed as metadata.

Implementation roadmap:

  1. Introduce the proposed intrinsics for store-available and load-visible operations (#172090, after renaming to match the new spec).
  2. Implement the proposed metadata “!amdgcn-av-none”.
  3. Expand the model to include asynchronous memory operations on AMDGPU architectures.
  4. Expand the model to be parameterized by address spaces.
4 Likes