# \[RFC\] Add NV-GPU dialect (HW specific extension of GPU dialect for Nvidia GPUs)

**URL:** https://discourse.llvm.org/t/rfc-add-nv-gpu-dialect-hw-specific-extension-of-gpu-dialect-for-nvidia-gpus/61466
**Category:** MLIR
**Created:** [April 4, 2022, 8:34pm UTC](https://discourse.llvm.org/t/rfc-add-nv-gpu-dialect-hw-specific-extension-of-gpu-dialect-for-nvidia-gpus/61466 "2022-04-04T20:34:51Z")
**Posts on this page:** 1
**Showing post:** 10

<div class="post-metadata">

### Author: ![bondhugula](https://avatars.discourse-cdn.com/v4/letter/b/cc9497/32.png) [@bondhugula](https://discourse.llvm.org/u/bondhugula)
#### Post date: [April 7, 2022, 6:31am UTC](https://discourse.llvm.org/t/rfc-add-nv-gpu-dialect-hw-specific-extension-of-gpu-dialect-for-nvidia-gpus/61466/10 "2022-04-07T06:31:15Z")

</div>

> [@ThomasRaoux](#):
>
> In some cases there is a large abstraction gap, making the lowering non trivial and preventing us from doing higher level transformation based on hardware details.

This part (“preventing us from doing higher-level transformation”) isn’t entirely clear to me. An example here would help. In the past, we’ve added wmma-level ops to the `gpu` dialect (although these were specific to NVIDIA GPUs) for the lack of an `nvgpu` dialect – these ops as you know use memrefs and GPU dialect-specific types, and it has been so far considered okay to add certain hardware-specific ops the `gpu` dialect itself: the key is that these ops still worked on neutral (MLIR builtin) types although their “actions” were GPU-specific (nvidia or AMD). Examples for general reference:

```auto
%C = gpu.subgroup_mma_load_matrix %22[%c0, %c0] {leadDimension = 16 : index} : memref<16x16xf32> -> !gpu.mma_matrix<16x16xf32, "COp">
...
%R = gpu.subgroup_mma_compute %A, %B, %C : !gpu.mma_matrix<16x16xf16, "AOp">, !gpu.mma_matrix<16x16xf16, "BOp"> -> !gpu.mma_matrix<16x16xf32, "COp">

```

> [@ThomasRaoux](#):
>
> This proposal is for adding a new `Nvidia/PTX` Dialect that

PTX in the name would appear to be out of place for a dialect like this. It looks like you want to have a specialized GPU dialect for NVIDIA GPUs: `nvgpu` instead?

> [@ThomasRaoux](#):
>
> Since there doesn’t seem to be any objections so far, I sent a patch introducing the new dialect for review:  
> [https://reviews.llvm.org/D123266](https://reviews.llvm.org/D123266)

It’ll be good to have more discussion here before we create it. I don’t think `nvptx` is the right name here (being the name of the final LLVM backend) – a big jump in abstraction through GPU → nvvm → LLVM → nvPTX.

---

_[View the full topic](https://discourse.llvm.org/t/rfc-add-nv-gpu-dialect-hw-specific-extension-of-gpu-dialect-for-nvidia-gpus/61466)._
