# \[RFC\] Continuing with bufferization::{TensorLike, BufferLike} - op semantics update in bufferization

**URL:** <https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983>\
**Category:** MLIR\
**Created:** [April 22, 2025, 12:49pm UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983 "2025-04-22T12:49:22Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![andrey-golubev](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/andrey-golubev/32/27277_2.png) [@andrey-golubev](https://discourse.llvm.org/u/andrey-golubev)\
**Post date:** [April 22, 2025, 12:49pm UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/1 "2025-04-22T12:49:22Z")

</div>

As a follow-up to the previous [discussion](https://discourse.llvm.org/t/rfc-changing-base-types-for-tensors-and-memrefs-from-c-base-classes-to-type-interfaces/85509/17) around extending one-shot bufferization support to user-specified types, I would like to continue along this path as I am in the process of using the newly added `TensorLike` and `BufferLike` type interfaces inside of the implementation.

In particular, my proposal is this: Bufferization’s operations (`to_tensor`, `to_memref`, `clone` but also `alloc_tensor` / `dealloc_tensor`?) need to be extended to work with `TensorLike` / `BufferLike`.

I can justify this for `to_tensor` and `to_memref` - from my perspective, these are the “unrealized\_conversion\_casts” with more semantics.  
I can not really justify it for other ops however: tensor / memref cloning kind of depends on what is a tensor / memref - so user-specific; allocation / deallocation is also something that is user-specific?

In general, the pitfall to me is this:

- if bufferization ops remain “as is”, users have to provide a complete set of similar ops themselves → this suggests introducing _op_ interfaces and bufferization starts to rely on interfaces instead of actual operations
  - downside: every new op has to have an associated interface
  - this also ultimately renders the new type interfaces useless (TensorLike / BufferLike)

- if bufferization ops change, it is not really clear what is “tensor/memref allocation”, “tensor/memref copy”, etc.
  - perhaps this is fine? - i.e. downstream projects could lower from bufferization-specific ops to “proper” ops separately, after one-shot bufferization is done

I do not think that “more customization points” (i.e. new callbacks in [bufferization options](https://github.com/llvm/llvm-project/blob/e428afdfcf56ccadbbcff16e8fe52e51622baed7/mlir/include/mlir/Dialect/Bufferization/IR/BufferizableOpInterface.h#L252)) would help here as the underlying issue is still the same: “roll your own ops” vs “support custom types in existing bufferization ops”.

This RFC is to discuss whether changing operations (effectively, their semantics) is viable and hopefully become aware of potential issues.

As an aside, more customization points are probably needed in any case (e.g. for [`allocateTensorForShapedValue`](https://github.com/llvm/llvm-project/blob/e428afdfcf56ccadbbcff16e8fe52e51622baed7/mlir/include/mlir/Dialect/Bufferization/IR/BufferizableOpInterface.h#L585)).

---

<div class="post-metadata">

**Author:** ![ftynse](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/ftynse/32/18644_2.png) [@ftynse](https://discourse.llvm.org/u/ftynse)\
**Post date:** [April 23, 2025, 7:01am UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/2 "2025-04-23T07:01:24Z")

</div>

Tensor copy can be just `to_tensor(to_memref(x))`, so IMO there is no difference between allowing these two ops support custom types and not allowing copies. All of them are similarly “aware” of the underlying type structure.

---

<div class="post-metadata">

**Author:** ![matthias-springer](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/matthias-springer/32/19066_2.png) [@matthias-springer](https://discourse.llvm.org/u/matthias-springer)\
**Post date:** [April 23, 2025, 8:56am UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/3 "2025-04-23T08:56:39Z")

</div>

> [@andrey-golubev](#):
>
> In particular, my proposal is this: Bufferization’s operations (`to_tensor`, `to_memref`, `clone` but also `alloc_tensor` / `dealloc_tensor`?) need to be extended to work with `TensorLike` / `BufferLike`.

`alloc_tensor` is inserted by the bufferization framework. See `TensorCopyInsertion.cpp`. That’s an internal pass that brings the IR into a form where no further analysis is needed and all tensors can be directly replaced with memrefs.

So I think you have to extend `alloc_tensor` (and for consistency also `dealloc_tensor`) in the same way. The implementation of `AllocTensorOp::bufferize` calls lambdas from `BufferizationOptions` to insert the buffer alloc/copy ops. These lambdas must create the correct buffer ops. That’s where the user-specific code is.

Is there a problem with this approach?

> if bufferization ops change, it is not really clear what is “tensor/memref allocation”, “tensor/memref copy”, etc.

The naming of `alloc_tensor` may not be ideal. It returns a tensor that is guaranteed to bufferize to a new buffer. I.e., the future buffer of the `alloc_tensor` result is guaranteed to be distinct from the future buffer of the `copy` operand (or any other “previous” buffer). In essence, this op gives you a tensor that will turn into buffer that doesn’t alias with anything (up to that point).

> [@andrey-golubev](#):
>
> I do not think that “more customization points” (i.e. new callbacks in [bufferization options](https://github.com/llvm/llvm-project/blob/e428afdfcf56ccadbbcff16e8fe52e51622baed7/mlir/include/mlir/Dialect/Bufferization/IR/BufferizableOpInterface.h#L252)) would help here as the underlying issue is still the same: “roll your own ops” vs “support custom types in existing bufferization ops”.

Which additional callbacks would be needed? I think we already have everything that we need.

> [@andrey-golubev](#):
>
> As an aside, more customization points are probably needed in any case (e.g. for [`allocateTensorForShapedValue`](https://github.com/llvm/llvm-project/blob/e428afdfcf56ccadbbcff16e8fe52e51622baed7/mlir/include/mlir/Dialect/Bufferization/IR/BufferizableOpInterface.h#L585)).

I haven’t looked into this function in detail, but you probably assume that `shapedValue` has a type that implements either `TensorLike` or `BufferLike`. That should make it possible to implement this function in a generic way. (By querying type interface methods to convert from tensor-like → buffer-like types.)

---

<div class="post-metadata">

**Author:** ![andrey-golubev](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/andrey-golubev/32/27277_2.png) [@andrey-golubev](https://discourse.llvm.org/u/andrey-golubev)\
**Post date:** [April 23, 2025, 9:55am UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/4 "2025-04-23T09:55:30Z")

</div>

> [@matthias-springer](#):
>
> So I think you have to extend `alloc_tensor` (and for consistency also `dealloc_tensor`) in the same way. The implementation of `AllocTensorOp::bufferize` calls lambdas from `BufferizationOptions` to insert the buffer alloc/copy ops. These lambdas must create the correct buffer ops. That’s where the user-specific code is.
> 
> Is there a problem with this approach?

Not really. I am just not sure upstream “allocation” makes sense for user types. In any case, thanks for the hints, I think it does indeed look “expected” to have `alloc_tensor` / `dealloc_tensor` extended.

What about `dealloc` and `clone` (both operate on memrefs)? They seem pretty memref-related to me (i also see that `options.getMemCpy` ends up producing `memref::CopyOp` by default so i’m not even sure why `clone` is special).

> [@matthias-springer](#):
>
> Which additional callbacks would be needed? I think we already have everything that we need.

Judging by the initial work in [[mlir][bufferization] Use TensorLike, BufferLike type interfaces by andrey-golubev · Pull Request #136736 · llvm/llvm-project · GitHub](https://github.com/llvm/llvm-project/pull/136736), I would “swap” `bufferization::getMemRefType` and `options.unknownTypeConverterFn`. Right now, the logic seems to be:

- “cannot bufferize via `BufferizableOpInterface::getBufferType`” → call `getMemRefType` → (internally) calls `options.unknownTypeConverterFn` when “ranked tensor without layout”.

I’d rather propose to do it this way:

- “cannot bufferize via `BufferizableOpInterface::getBufferType`” → call `options.unknownTypeConverterFn` → (internally) calls `getMemRefType`

This way we hide the TensorType / BaseMemRefType details behind options. But I’ll need to dwell on this as I’m not sure if it’s good in general (there are multiple places where the `getMemRefType` is being used).

> [@matthias-springer](#):
>
> I haven’t looked into this function in detail, but you probably assume that `shapedValue` has a type that implements either `TensorLike` or `BufferLike`.

So far, I can ignore the problem in `one-shot-bufferize` (at least) by assuming `shapedValue` is a builtin tensor/memref. Generally though, you’re correct.

> [@matthias-springer](#):
>
> That should make it possible to implement this function in a generic way. (By querying type interface methods to convert from tensor-like → buffer-like types.)

Do you propose to introduce `TensorLikeType::getBufferType` and perhaps a handful of other interface methods? Doing it this way might end up reducing the need for certain option-level hooks (and maybe eliminate op-level interfaces methods e.g. `BufferizableOpInterface::getBufferType`) and also actually allow one to use `one-shot-bufferize` pass directly. For instance, the current issue is that one needs to set options manually in order to use: function boundary bufferization, `allocateTensorForShapedValue` and likely other primitives that rely on the “not BufferizableOpInterface then call `getMemRefType`” kind of model.

Edit: for “tensor-like → buffer-like types” conversion, I guess we also need to agree on whether this should be _external_ or _internal_ w.r.t. the types. External: smth.getBufferType(TensorLike) → BufferLike; internal: TensorLike.getBufferType() → BufferLike.

I am still thinking about TensorLike / BufferLike to be eventually hoisted to builtins. If that ever to be considered, having “internal conversion” model is a hard blocker.

---

<div class="post-metadata">

**Author:** ![matthias-springer](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/matthias-springer/32/19066_2.png) [@matthias-springer](https://discourse.llvm.org/u/matthias-springer)\
**Post date:** [April 24, 2025, 7:46am UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/5 "2025-04-24T07:46:24Z")

</div>

> [@andrey-golubev](#):
>
> What about `dealloc` and `clone` (both operate on memrefs)? They seem pretty memref-related to me (i also see that `options.getMemCpy` ends up producing `memref::CopyOp` by default so i’m not even sure why `clone` is special).

`bufferization.clone` is used by the buffer deallocation pass. I don’t think we use it anywhere else. I’m also not sure why we need the op at all. We could probably remove it and replace it with `memref.alloc + memref.copy`.

> [@andrey-golubev](#):
>
> “cannot bufferize via `BufferizableOpInterface::getBufferType`” → call `options.unknownTypeConverterFn` → (internally) calls `getMemRefType`

It’s been too long since I looked at this… `unknownTypeConverterFn` does not take a layout, so I’m not sure if it can we wired like that. I would look into the places that are calling `unknownTypeConverterFn`. What if you put a failed assertion in the lambda? What tests are failing? I think this lambda is kind of a “fallback” when we don’t know which type to use. But when is that actually the case? I don’t remember…

`getMemRefType` produces a memref type from a tensor type. Now that we have multiple buffer types, this sounds like a candidate for a type interface method to me.

> [@andrey-golubev](#):
>
> Do you propose to introduce `TensorLikeType::getBufferType`

yes

> [@andrey-golubev](#):
>
> and maybe eliminate op-level interfaces methods e.g. `BufferizableOpInterface::getBufferType`

I think you’re still going to need that. This function is predicting the future buffer type for a tensor result. E.g., for `tensor.extract_slice`, the bufferized result must have a certain layout map. We need to predict this type when bufferizing a loop op: if I remember correctly, the loop op is bufferized before the loop body.

Inside of `BufferizableOpInterface::getBufferType`, the builtin MemRefType can be hard-coded. This is an interface method on a specific operation. E.g., `memref.subview` (bufferized operation of `tensor.extract_slice`) supports only MemRefType and no custom buffer types, so it’s OK to hardcode MemRefType.

> [@andrey-golubev](#):
>
> Edit: for “tensor-like → buffer-like types” conversion, I guess we also need to agree on whether this should be _external_ or _internal_ w.r.t. the types. External: smth.getBufferType(TensorLike) → BufferLike; internal: TensorLike.getBufferType() → BufferLike.

What’s the difference? The first one is a member function, the second one is a static function? What is `smth`?

---

<div class="post-metadata">

**Author:** ![andrey-golubev](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/andrey-golubev/32/27277_2.png) [@andrey-golubev](https://discourse.llvm.org/u/andrey-golubev)\
**Post date:** [April 24, 2025, 9:50am UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/6 "2025-04-24T09:50:08Z")

</div>

> [@matthias-springer](#):
>
> > [@andrey-golubev](#):
> >
> > and maybe eliminate op-level interfaces methods e.g. `BufferizableOpInterface::getBufferType`
> 
> I think you’re still going to need that. This function is predicting the future buffer type for a tensor result. E.g., for `tensor.extract_slice`, the bufferized result must have a certain layout map. We need to predict this type when bufferizing a loop op: if I remember correctly, the loop op is bufferized before the loop body.

I see. So the `BufferizableOpInterface::getBufferType` is a building block of `BufferizableOpInterface::bufferize` kind of _and simultaneously_ the method to preliminary query the output buffer without actually bufferizing. I guess the renewed logic is then something like:

- Try `BufferizableOpInterface::getBufferType` if op supports BufferizableOpInterface
- Go to “new getBufferType API” otherwise

(the cases could likely be enclosed into the free-standing `getBufferType` function or something along these lines)

> [@matthias-springer](#):
>
> Inside of `BufferizableOpInterface::getBufferType`, the builtin MemRefType can be hard-coded.

Do you mean at API level or in implementation? I guess the latter does make sense indeed. I’d still prefer `BufferLikeType BufferizableOpInterface::getBufferType(TensorLikeType, ...)` at the signature level though.

> [@matthias-springer](#):
>
> What’s the difference? The first one is a member function, the second one is a static function? What is `smth`?

Yes. Member function - clear enough (again, main concern is that we couple “types” and “operations” on these types). Non-member function:

- **bufferizationOptions**.getBufferType(TensorLike) → BufferLike (`smth` - options object)
- **converter**.getBufferType(TensorLike) → BufferLike (`smth` - some converter object that is a [DialectInterface](https://mlir.llvm.org/docs/Interfaces/#dialect-interfaces) that one just creates for user-specified tensor/memref - they likely live in own dialect anyway)
  - this has to live inside the context forever though (so slightly higher memory usage)

From the standpoint of custom tensors / memrefs support, options is probably most cumbersome because they require one to basically reimplement `one-shot-bufferization` pass itself (not the underlying `mlir::runOneShotModuleBufferize()` function with the actual logic though) - because that’s the only way to “seed” options object? For instance, this is what we do in our downstream and it works fine, I am not against this in general.

---

<div class="post-metadata">

**Author:** ![andrey-golubev](https://sea1.discourse-cdn.com/flex021/user_avatar/discourse.llvm.org/andrey-golubev/32/27277_2.png) [@andrey-golubev](https://discourse.llvm.org/u/andrey-golubev)\
**Post date:** [April 24, 2025, 1:20pm UTC](https://discourse.llvm.org/t/rfc-continuing-with-bufferization-tensorlike-bufferlike-op-semantics-update-in-bufferization/85983/7 "2025-04-24T13:20:20Z")

</div>

> [@andrey-golubev](#):
>
> > [@matthias-springer](#):
> >
> > What’s the difference? The first one is a member function, the second one is a static function? What is `smth`?
> 
> Yes. Member function - clear enough (again, main concern is that we couple “types” and “operations” on these types). Non-member function:

One more thing I just remembered, our downstream also extends builtin tensor (via [encoding](https://github.com/llvm/llvm-project/blob/d7f3c3129344b133859d89d962fcdd5058702f72/mlir/include/mlir/IR/BuiltinTypes.h#L261)) and memref (via [layout interface](https://github.com/llvm/llvm-project/blob/d7f3c3129344b133859d89d962fcdd5058702f72/mlir/include/mlir/IR/BuiltinTypes.h#L204)).  
In general, this would require a customization point to overwrite builtin.tensor → builtin.memref conversion as well [1] - we want to be able to convert `encoding` to `layout` in a particular fashion ourselves. Which is why having `TensorLikeType::getBufferType` might be problematic - the TensorLike interface is not “promised” for builtins but hard-attached instead (see [details](https://github.com/llvm/llvm-project/pull/134220#issuecomment-2781093571)) - so supplying a custom implementation is hard-ish.

[1]: I think so far we manage to avoid solving this issue on our end by a combination of: `options.copyBeforeWrite = false; options.testAnalysisOnly = true;` (these two luckily “turn off” any tensor copies and allocations during one-shot bufferization). Any `Op::bufferize` implementation also uses our own “get buffer type” method that also handles builtins.
