Hello all,
I would like to ask for feedback on a possible design direction in One-Shot Bufferize.
Today, bufferization already exposes useful extension points for several important cases, such as function-boundary type conversion, unknown-type fallback, and default memory-space selection. This seems to work well for the current use cases. At the same time, I am wondering whether the customization surface is gradually becoming organized more by where a decision is made than by what kind of decision it is.
Relevant code:
-
* Bufferize.cpp option wiring (llvm/llvm-project/blob/main/mlir/lib/Dialect/Bufferization/Transforms/Bufferize.cpp#L70-L120)
The concern is not that the current approach is wrong. Rather, downstream users integrating custom tensor-like types often need to express one coherent policy for questions such as:
-
how target buffer types should be chosen,
-
how memory space should be resolved,
-
how layout mismatches should be merged,
-
and what should happen when there is still a mismatch (e.g. cast, copy, or failure).
Today, these decisions appear to be distributed across several call sites.
For example:
-
allocateTensorForShapedValuehas its own memory-space resolution path, with a local fallback chain involvinggetBufferType(...)anddefaultMemorySpaceFn(...):
* BufferizableOpInterface.cpp -
SCF merge logic handles mismatches locally. In
scf.ifdiffering memory spaces lead to failure, while differing layouts are promoted to a fully dynamic layout map:
* SCF if merge (llvm/llvm-project/blob/main/mlir/lib/Dialect/SCF/Transforms/BufferizableOpInterfaceImpl.cpp#L313-L323) -
Similar logic appears again in loop
iter_argmerging:
* SCF loop merge(llvm/llvm-project/blob/main/mlir/lib/Dialect/SCF/Transforms/BufferizableOpInterfaceImpl.cpp#L540-L586) -
Some downstream transformations also impose their own layout constraints. For example,
BufferResultsToOutParamscurrently accepts only static identity layout or fully dynamic layout maps:
* BufferResultsToOutParams.cpp
There are also existing tests that already exercise several of these conflict families.
alloc_tesor_copy_from_non_default_space_no_castshows an awkward interaction between memory-space conversion and function-boundary expectations: the test currently relies on a bufferization.to_tensor round-trip and even contains a TODO noting that this should likely become illegal once function boundaries are bufferized.
return_extract_sliceshows that layout-conversion policy is not merely cosmetic: different layout modes change whether the result can be returned directly or requires an additional alloc/copy/generalization step, with an additional TODO around cleanup of inserted buffers.
materialize_in_destinationshows a similar cross-memory-space resolution problem in a different usepoint, where materialization into a destination tensor with another encoding/memory space is handled through a separate local path.
Because of this, I am wondering whether it would make sense to gradually move toward a more unified policy surface for buffer type resolution and conflict handling, while keeping the current behavior unchanged by default.
What seems important here is that the bufferization process itself is often implemented in op-specific interface models. At the same time, many of the options that control type conversion, layout, and memory-space decisions are global in intent.
For example, BufferizationOptions exposes global knobs such as function-boundary type conversion, unknown-type fallback, and default memory-space selection, and the pass wires them globally. However, when an op needs specific mismatch behavior, the resolution is often implemented locally in that op’s bufferization path instead of going through a more general policy surface. Function boundaries, for instance, already have dedicated hooks for argument/result type conversion, while caller/callee type mismatch is handled locally in CallOp via castOrReallocMemRefValue.(code) Similar local mismatch logic appears in SCF merge points and in strided-layout paths such as tensor.extract_slice / tensor.insert_slice.(inside of inferRankReducedResultType for insert slice, direct creation of meref in extract slice)
One possible direction could be a small set of reusable policy hooks, for example:
-
TypeConversionPolicy: compute the target buffer-like type for a given context,
-
MemorySpacePolicy: resolve memory space from tensor/encoding/copy context,
-
LayoutMergePolicy: decide how to merge competing layout choices,
-
MismatchActionPolicy: decide whether a remaining mismatch should become a cast, copy, or failure.
I am not attached to this exact decomposition. It may be better as a single callback over a richer request object, or as a small policy interface/object rather than several callbacks.
What I would mainly like feedback on is the direction:
-
Does this seem like a real gap in the current customization model?
-
Is it better to continue adding usepoint-specific hooks as needed, or would a more semantic policy surface be preferable long-term?
-
If a unified policy makes sense, should it be modeled as one structured callback or as a small set of orthogonal hooks?
-
Which call sites would be the best first candidates to route through such a policy without changing default behavior?
If this direction sounds reasonable, I can prepare an incremental PR plan that preserves current defaults and only centralizes selected decision points first.
Thanks!