RFC: Generalizing buffer type resolution in One-Shot Bufferize

Hello all,

I would like to ask for feedback on a possible design direction in One-Shot Bufferize.

Today, bufferization already exposes useful extension points for several important cases, such as function-boundary type conversion, unknown-type fallback, and default memory-space selection. This seems to work well for the current use cases. At the same time, I am wondering whether the customization surface is gradually becoming organized more by where a decision is made than by what kind of decision it is.
Relevant code:

  • * BufferizationOptions:319

  • * Bufferize.cpp option wiring (llvm/llvm-project/blob/main/mlir/lib/Dialect/Bufferization/Transforms/Bufferize.cpp#L70-L120)

The concern is not that the current approach is wrong. Rather, downstream users integrating custom tensor-like types often need to express one coherent policy for questions such as:

  • how target buffer types should be chosen,

  • how memory space should be resolved,

  • how layout mismatches should be merged,

  • and what should happen when there is still a mismatch (e.g. cast, copy, or failure).

Today, these decisions appear to be distributed across several call sites.

For example:

  • allocateTensorForShapedValue has its own memory-space resolution path, with a local fallback chain involving getBufferType(...) and defaultMemorySpaceFn(...):
    * BufferizableOpInterface.cpp

  • SCF merge logic handles mismatches locally. In scf.ifdiffering memory spaces lead to failure, while differing layouts are promoted to a fully dynamic layout map:
    * SCF if merge (llvm/llvm-project/blob/main/mlir/lib/Dialect/SCF/Transforms/BufferizableOpInterfaceImpl.cpp#L313-L323)

  • Similar logic appears again in loop iter_arg merging:
    * SCF loop merge(llvm/llvm-project/blob/main/mlir/lib/Dialect/SCF/Transforms/BufferizableOpInterfaceImpl.cpp#L540-L586)

  • Some downstream transformations also impose their own layout constraints. For example, BufferResultsToOutParams currently accepts only static identity layout or fully dynamic layout maps:
    * BufferResultsToOutParams.cpp

There are also existing tests that already exercise several of these conflict families.
alloc_tesor_copy_from_non_default_space_no_castshows an awkward interaction between memory-space conversion and function-boundary expectations: the test currently relies on a bufferization.to_tensor round-trip and even contains a TODO noting that this should likely become illegal once function boundaries are bufferized.
return_extract_sliceshows that layout-conversion policy is not merely cosmetic: different layout modes change whether the result can be returned directly or requires an additional alloc/copy/generalization step, with an additional TODO around cleanup of inserted buffers.
materialize_in_destinationshows a similar cross-memory-space resolution problem in a different usepoint, where materialization into a destination tensor with another encoding/memory space is handled through a separate local path.

Because of this, I am wondering whether it would make sense to gradually move toward a more unified policy surface for buffer type resolution and conflict handling, while keeping the current behavior unchanged by default.

What seems important here is that the bufferization process itself is often implemented in op-specific interface models. At the same time, many of the options that control type conversion, layout, and memory-space decisions are global in intent.

For example, BufferizationOptions exposes global knobs such as function-boundary type conversion, unknown-type fallback, and default memory-space selection, and the pass wires them globally. However, when an op needs specific mismatch behavior, the resolution is often implemented locally in that op’s bufferization path instead of going through a more general policy surface. Function boundaries, for instance, already have dedicated hooks for argument/result type conversion, while caller/callee type mismatch is handled locally in CallOp via castOrReallocMemRefValue.(code) Similar local mismatch logic appears in SCF merge points and in strided-layout paths such as tensor.extract_slice / tensor.insert_slice.(inside of inferRankReducedResultType for insert slice, direct creation of meref in extract slice)

One possible direction could be a small set of reusable policy hooks, for example:

  • TypeConversionPolicy: compute the target buffer-like type for a given context,

  • MemorySpacePolicy: resolve memory space from tensor/encoding/copy context,

  • LayoutMergePolicy: decide how to merge competing layout choices,

  • MismatchActionPolicy: decide whether a remaining mismatch should become a cast, copy, or failure.

I am not attached to this exact decomposition. It may be better as a single callback over a richer request object, or as a small policy interface/object rather than several callbacks.

What I would mainly like feedback on is the direction:

  1. Does this seem like a real gap in the current customization model?

  2. Is it better to continue adding usepoint-specific hooks as needed, or would a more semantic policy surface be preferable long-term?

  3. If a unified policy makes sense, should it be modeled as one structured callback or as a small set of orthogonal hooks?

  4. Which call sites would be the best first candidates to route through such a policy without changing default behavior?

If this direction sounds reasonable, I can prepare an incremental PR plan that preserves current defaults and only centralizes selected decision points first.

Thanks!

@matthias-springer If you have time, I’d really appreciate your perspective on this RFC

A few thoughts on this:

  • A type conversion policy may have to be somewhat op-specific. E.g., when bufferizing tensor.extract_slice, we generate memref.subview, which works only on memref types and not other custom buffer types. Similarly, some ops may support only memref types with identity layout and using a different layout map may result in invalid IR.
  • Memory space policy: Note that the tensor encoding is not the equivalent of a memref memory space. You can encode arbitrary data in there. E.g., the sparse tensor dialect uses it to encode the storage format.
  • Memory space policy / layout merge policy: A “copy policy” would be easy to implement: we can insert allocations + memcpy ops (using the lambdas in BufferizationOptions) to reconcile memory space / layout mismatches such as the ones that you are seeing during scf.if bufferization. We could generalize these lambdas to allow for failure / cast generation. Maybe something along the lines of ReconcileBufferTypeMismatchFn, which defaults to alloc+copy.

Overall, I think nobody really thought much about how to deal with mismatching memory spaces or layouts so far. I think this is worth investigating.

Thanks, this is very helpful feedback.

I agree with your points:

I will prepare a small POC to make the discussion more concrete.

So far, I see several issues on the SCF level and on memref subview semantics. Here is a concrete example of what we want. (snippet of my prototyping for an Intel NPU compiler )

Before bufferization:

%res = scf.for … iter_args(%iter = %init) → tensor<1x1x2000x2000xf16, {order=#NCHW}> {
    %tile = tensor.extract_slice %src[…] […] […]
                : tensor<1x1x2000x2000xf32>
                to tensor<1x1x?x2000xf32, {bounds=…, order=#NCHW}>
    %conv = NPU.Op(%tile)
                : tensor<1x1x?x2000xf32, {bounds=…, order=#NCHW}>
                → tensor<1x1x?x2000xf16, {bounds=…, order=#NCHW}>
    %next = tensor.insert_slice %conv into %iter[…] […] […]
                : tensor<1x1x?x2000xf16, {bounds=…, order=#NCHW}>
                 into tensor<1x1x2000x2000xf16, {order=#NCHW}>
    scf.yield %next : tensor<1x1x2000x2000xf16, {order=#NCHW}>
}

After bufferization:

%res = scf.for … iter_args(%iter = %init_memref) → memref<1x1x2000x2000xf16, {order=#NCHW}> {
    %src_view = memref.subview %src_memref[…] […] […]
                        : memref<1x1x2000x2000xf32>
                        to memref<1x1x?x2000xf32, {order=#NCHW, runtimeShapeInfo=…, strides=…}>
    %tmp = memref.alloc(%dyn)
                        : memref<1x1x?x2000xf16, {order=#NCHW, runtimeShapeInfo=…}>
    %dma = NPU.LowerOp inputs(%src_view) outputs(%tmp)
    %dst_view = memref.subview %iter[…] […] […]
                        : memref<1x1x2000x2000xf16, {order=#NCHW}>
                        to memref<1x1x?x2000xf16, {order=#NCHW, runtimeShapeInfo=…, strides=…}>
    memref.copy %dma, %dst_view
    scf.yield %iter : memref<1x1x2000x2000xf16, {order=#NCHW}>
}

As you see, we are at least keeping order attribute, which is our compiler-specific parameter.

We have mainly 2 unclear things so far in upstream:

  1. Making SCF infrastructure support conversion from custom tensor-encoding to memref layout. (similarly to conversion from tensor-like to buffer-like, but for parts of builtin types)
    1. Type mismatch problem seems resolvable with ReconcileBufferTypeMismatchFn or analog.
    2. Conversion mechanism (tensor-encoding → memref layout) itself is not clear - in theory, we can take a similar approach to functionArgTypeConverterFn hook
  2. Make memref.subview work on “abstract” memref layout that exposes StidedLayoutAttr-like API (for example, the memref layout interface already has getStridesAndOffset method). That way, we can use our own layout object while still conforming to the subview’s internal contract.

If this direction sounds reasonable, we will reflect it directly in the second POC and API proposal.

The following sequence of PRs has resolved this issue:
1.

2.

3.

4.

5.

With these changes, we can now resolve conflicts between buffer types while correctly accounting for custom tensor encodings.

Many thanks to @matthias-springer for the help and support!