Upstreaming basic support for accelerators

Meta and NVIDIA have been working together on a couple of branches that provide basic support for debugging AMD and NVIDIA GPUs:

We have reached a state in which LLDB can be used for debugging kernels very decently but with some restrictions. For NVIDIA, the restrictions are: one kernel at a time, <20-30K threads (otherwise it gets too slow; fixable by adding denser gdb-remote packets, etc.), no step-over (you can only use breakpoints and do resume for now). At least for debugging small to medium kernels, this is already very useful. There’s still a lot of work ahead, but we think that now is a good moment to start upstreaming our changes, start getting contributions and ideas, foster collaboration within our own companies (yep, not everyone internally can use our fork directly), and at the same time have a lower rebase burden internally.

The good news is that so far we have made very few changes to the core of LLDB and 95% of our code is in the form of plugins. So I have a good feeling about the upstreamability of this. I will summarize the most important changes to get some early feedback and avoid wasting cycles during code review.

  • lldb-server plugins: We added this feature that allows a CPU target to initialize an lldb-server plugin based on some event. For example, in the NVIDIA case, if the CPU target hits a certain breakpoint (cuda_init) then the corresponding plugin will get initialized. We are using this plugin to create a secondary accelerator target that debugs only the GPU. Plugins can be disabled and enable by the user as well, so that you pay the initialization cost of the plugin if you really want to use it.
  • accelerator actions: The cpu and accelerator targets can execute actions on each other usin gaccelerator actions. They come in a few flavors: e.g. synchronize the targets upon a stop, set a breakpoint, fetch some symbols upon a breakpoint is hit, etc. They are meant to simplify the complex interactions between targets.
  • address space support: accelerators like gpus often have multiple address spaces, which means that whenever data needs to be read, an address space identifier needs to be set. Each accelerator plugin provides an address space spec that is used by lldb to do the correct data fetching and dwarf parsing. We also had some changes to dwarf expressions to make use of these spaces but they are very localized.
  • mock accelerator plugin: NVIDIA and AMD support need to link lldb-server with some specific libraries, which makes generic testing a problem. Because of this, we created a mock accelerator plugin that can test most codepaths without relying on any external library.

With this, we have been able to enable debugging of accelerators by having a target for the cpu and one target for the accelerator. We could theoretically support multiple accelerators at once for the same target, but we don’t have a good use case for that yet. We also didn’t want to mix the cpu and the accelerator in a single target because that opens a can of worms: the platforms are different; they both might get events simultaneously from signals or the driver making synchronization a tad difficult if they are the same target; the way you navigate through each set of threads is different. For example, it’s common to traverse over CPU threads by ids, but for GPUs, you traverse by coordinates, such as blockIdx (1, 2, 123) threadIdx (0, 0, 100). It’s relatively easy to add custom behavior to the user interaction via Platform plugins when the targets are independent. Not only that, we also want to support accelerator-only debugging as well, in which case the CPU target is dettached upon the initialization of the accelerator.

The long term vision we have for this is that we want to add the necessary foundation so that anyway can enable their own accelerator in lldb. And besides that, we want to eventually add search and filtering capabilities to LLDB. In fact, when doing GPU debugging, you are dealing with potentially more than a million threads, so the big question is how can the debugger help you search quickly and identify the threads that satisfy certain conditions you care about, or how can you narrow down the set of threads that are interesting to you as a way to reduce noise. We have a few ideas, like adding a query language that can optimize searches by executing some actions on the server and some others in the client. But we don’t want to overthink this as we’d like first to have some widespread user feedback to figure out solutions to these problems.

As a final note, we’ve decided to use the name accelerator instead of GPU because we are experimenting with non-gpu architectures as well, and making this generic is a good idea anyway. If you can suggest a better name with a convincing reason, that would be very nice as well!

Any feedback or ideas are welcome. I plan to start upstreaming this with the Meta folks soon.

13 Likes

Thanks for starting this discussion, Walter. I appreciate the context. Here are my initial thoughts:

  • The plugin approach sounds promising and will certainly help with upstreaming and reviewing.
  • Separate targets for the CPU and the accelerator are the right call. I really can’t imagine that working out any other way.
  • I’m very excited about address space support. I’d love to hear more. At this point, it sounds limited to the accelerators, but we have other use cases for it in LLDB (like Wasm), so hopefully, we can find a way to thread this through all of LLDB if it isn’t already.
  • I’d also love to hear more about the protocol extensions you have or are considering.

I do have a few questions, mostly requests for more details:

  1. Do you have a plan for how you’ll be upstreaming this? The context here is helpful, but for someone that’s not been involved in this, it’s not immediately clear what the different steps and deliverables are.
  2. Can you share a bit more about the testing strategy? You mention the mock accelerators. Does that mean that most testing is done by debugging an lldb-server that runs one of those mock plugins? Do you plan on setting up any bots with real hardware?

Thanks, Water, for kicking this off.

Jonas, regarding #2, I can briefly share the testing strategies we’ve used in this work, including items specific to the AMD plugin.

First, we absolutely do not want to regress any existing CPU LLDB tests. At the moment, I believe the dual-target design is causing a number of test failures that we need to resolve before upstreaming.

For new tests, we’ve added coverage at multiple levels:

  1. Accelerator plugin design

    We added new mock plugin tests to validate the new accelerator plugin architecture.

  2. AMD plugin (owned by our company)

    We’ve built several layers of testing:

Interestingly, cuda-gdb and rocgdb (nvidia and amd debuggers correspondingly), do the debugging using a single target by synchronizing the run/stop states of the GPU and CPU, i.e. you can’t have the CPU running while you are stopped at a GPU breakpoint. However, this design breaks, at least in the NVIDIA case, because some deadlocks can happen while debugging due to the forced state synchronization between the processors. This is rare and I have only seen them when people so CPU-GPU spin locks, but it’s still worth mentioning.
Another interesting example, which is actually more relevant, is supporting dual python and GPU debugging. We had an experiment in which we wanted to use both the python debugger and gdb to debug a kernel launched by point. We observed that on GPU stops, gdb was stopping the CPU, which caused the python debugger to halt, because the python debugger uses the interpreter under the hood, which was halted by gdb. Therefore the user couldn’d do python debugging anymore! I realized at that moment that a better architecture is to keep the gpu and cpu as individual targets to allow for great flexibility when designing more complex debugging solutions.

What we did was to add one gdb-remote packet to read memory given an address space identifier. The identifiers along with their names are provided by the target during initialization. We also added support for this in the dwarf parser. CUDA programs user the DW_TAG_ADDRESS_CLASS to denote the address space of variables. We might move to an ADDRESS_SPACE tag created in DWARF6, but that will take a while. In any case, the dwarf changes were relatively small, and we didn’t add support for caching on the client side, but we should.
This is the original PR that added support for address spaces Add support for memory spaces. · clayborg/llvm-project@0a6ca1b · GitHub if you want to take a look. It’s mostly boilerplate though.

At least for the short term, I mostly envision creating batch versions of several thread-based packets we have now. For example, if there’s a packet that reads some registers from a thread, we’ll most likely need a version that does the same but for multiple threads at once. These are the easy ones and I don’t think they will require much thought.
The ones that are tricky are related to stepping. In the gpu case, when you step an individual thread, you are also stepping all the other threads in its warp/wave, which are 32 or 64 threads. This means that, in contrary to lldb’s current expecations, stepping a thread has side effects on other threads and the vCont packet should be augmented to return to the client the information of the threads that have been affected. Currently, the vCont packet assumes that, when stepping a single thread, only that threads moves.

The other extension I mentioned in my original post is related to accelerator actions. Greg added an extension to the stop-reply packet so that it can return a set of accelerator actions to the client, which can trigger some actions. There’s no gpu action packet per se, it’s just metadata added to the stop-reply packet.

Beyond those kinds of packets and the extensions for memory addresses, I can’t see right now other packets we might need. I’m actually happy that we got nvidia debugging in such a decent state with almost no packets added.

Most likely the progression will be:

  • Add support for lldb-remote plugins tested with an empty mock accelerator plugin.
  • Add support for basic accelerator actions to initialize the mock plugin only under certain conditions.
  • Add support for a proper accelerator target created by the mock plugin.
  • Add support for some level of synchronization between the CPU and accelerator plugins
  • Add support for breakpoints, memory read, fetching dylibs, etc. in the mock plugin
    etc etc

All of this is already tested in our dev branch and doesn’t require any special libs linked with lldb.
After this work is done, I will upstream NVIDIA’s plugin. The Meta folks will decide when to upstream the Amd plugin.

I think this progression is quite good because it can be done in small-to-medium chunks avoiding gigantic patches.

Yes, testing right now requires launching a server even for the mock plugin.

I w.r.t. bots, I don’t think NVIDIA will provide one at least this year. They might do if lldb is at some point on track to take over gdb for NVIDIA though.

Hello,

Isn’t this a consequence of GDB’s all-stop mode, which is default behavior, rather than target configuration? In non-stop mode, the user would be able to keep other (CPU and GPU) threads running while inspecting a stopped GPU thread. Even in all-stop mode, it should be possible to resume a single thread/wave asynchronously (with the continue & command using scheduler-locking), then switch to a CPU thread and resume it, too, so that the deadlock is resolved (I’m assuming the accelerator supports individual resume of threads/waves; it’s already possible for CPU threads.)

Handling the GPU part as a separate inferior means that there will be moments where the device would be idle, i.e. no threads are running. Skimming through the git repo, I noticed that you defined a shadow/fallback thread, so that there would always exist a thread from LLDB’s perspective. Do you plan to keep this fallback thread or get rid of it in the future?

Thank you.

First of all, thank you very much for skimming through the code and leaving feedback. This makes me very happy :slight_smile:

You are right to some extent. For sure the gdb gpu architects could have used two targets. Nothing prevented that, but they opted for a single target, which perhaps is more aligned with the heterogenous programming philosophy that the compiler tries to adhere to. It turns out that in practice that philosophy has some drawbacks for debugging, as explained above. W.r.t. to all-stop vs non-stop, all-stop is the default on all platforms and non-stop is only available on certain targets, so the restrictions on the lock-step situation are based on the fact that the only configuration that works for all platforms is all-stop. Still, there are some interesting situations that arise with non-stop. When you step a gpu thread, you are actually stepping a group of threads, which would require some new abstractions on top of non-stop mode; this is all doable, but it turns out that there’s no easy path with the non-stop situation.

Yes, this is something inconvenient. What complicates further this situation is that once you initialize the gpu, you never shut it down, which means that if your host will launch multiple kernels sequencially, there will indeed be multiple periods of time in which there will be no gpu threads, and after the last kernel finished, the gpu will have no running threads any more and lldb doesn’t know that there will no additional kernels. We have two paths there: 1) Teach lldb how to handle a target that might not have threads. Right now it’s expecting forcefully at least one thread. 2) Use a fake/placeholder thread that represents an idle gpu. I opted for the second option in my current implementation because it’s the easiest and it hasn’t created any bad behavior yet. I’m sure that if someone tries to use lldb via scripts they will need to add a special handler for this situation, but it’s just that 1) is way more work and it’s very intrusive to lldb. I’m open to do 1) if we eventually see that it’s really necessary though.

Thanks for the great discussion in this thread about upstreaming basic accelerator support!
Building on this discussion, we at Intel have been developing LLDB support specifically for Intel accelerators.
This work aligns with the broader vision of bringing accelerator debugging capabilities into the LLVM ecosystem.
As our contribution, we’ll soon be sharing a branch with the IntelGPU plugin implementation in this thread (still WIP, so no firm timeline yet).
Feel free to check it out and let us know what you think!
Looking forward to collaborating with everyone and jumping into code reviews and discussions.

Music to my ears. I’ll sync up with the Meta and NVIDIA folks about this and try to set up a connection with you guys :slight_smile:

FYI we recently added GDB Remote Protocol Extensions - 🐛 LLDB. Though that packet is just for breakpoints, the format of it is a wrapper for a batch of packets we already support.

So if what you want is along the lines of “do X on these N things” where X is something we can already do one at a time, consider that format.

I am late to this discussion and luckily I don’t have anything major to bring up. Looking at the first PR now.

Can you show a small demo of this? I am most interested in how users will select a target (assuming they have to) and how backtrace will work. Assuming this is an offload type scenario where you “call” into the accelerator, I don’t know if that’s the case for you though.

You have one target that’s the CPU and one for the GPU. A debugger could either:

  • Present them as one multi arch target, somehow.
  • Present two targets where events on one target effect the other. (you are proposing this style?)

I think I understand how this would work, and based on that, I think I agree with the choice to have separate targets (apart from anything else, it saves making large parts of lldb multi-arch aware).

But yeah, some short demo output would be really helpful.

I created a small demo for debugging AMD HIP Kernel, hopefully it helps to answer the questions.

Checkout this gist : gpu_demo.hip · GitHub

Thanks for the demo, that makes things clear.

1 Like

@radocaj We started upstreaming meta’s implementation. at this point 60-70% framework is merged upstream with the following PRs.

[lldb-server] Add accelerator plugin infrastructure for debugging hardware accelerators like gpus

[lldb-server] Add breakpoint support to accelerator plugin protocol
[lldb] Add accelerator plugin connection support
[lldb] Handle accelerator plugin breakpoint actions on the client

In next 2-3 weeks we will complete the upstreaming of our implementation.

Letting you know if it makes to reuse any of this work.
We are also open for collaboration and discussion.

cc: @jeffreytan81 , @clayborg

@satyajanga Thank you for the update and for keeping us in the loop on the upstreaming progress. This is really exciting news!

IntelGT Dev Status: We are currently in the final phase of polishing our basic support for Intel GPU (IntelGT), which is based on the accelerator architecture discussed here. This work should not take long, and our next step will be to share a dev branch to collect your feedback.

Once the branch is shared, it would be a great opportunity to sync up and discuss how we can best leverage what you’ve built.
We believe this will be a fantastic way to collaborate and ensure our efforts are well aligned with the upstream direction.

Looking forward to hearing your thoughts!

1 Like

We are excited to share that we have reached a presentable, functional solution that we would like to share with the LLDB community!

Intel GPU Plugin (IntelGT) PR: :backhand_index_pointing_right: Intel GPU Plugin by radocaj-int · Pull Request #117 · clayborg/llvm-project · GitHub

Our goal with this PR is to provide comprehensive LLDB support for IntelGT through an implementation of the lldb-server plugin.

“Intel GT” is the umbrella name for Intel’s GPU IP — the same graphics and compute engine that ships in Intel client CPUs (as an integrated GPU, iGPU) and in the Data Center GPU product lines (as a discrete GPU). The “GT” originally stood for “Graphics Technology” and today refers to the Xe family of GPU architectures.

The PR currently supports kernel debugging on the following GPUs:

Alchemist
Ponte Vecchio
Battlemage

All IntelGT share the same programmable execution engine and the same debug interface, meaning that a single LLDB plugin aims to cover the entire IntelGT family.

IntelGT Programming Model
Intel GT leverages Level Zero as its low-level compute API, with higher-level programming support provided through SYCL/DPC++ and OpenCL. Kernels are first compiled to SPIR-V and then finalized into the Intel EU ISA by the graphics compiler at load time.

IntelGT Debug Interface
Debugging of Intel GT is facilitated through the Level Zero debug extension, which provides the necessary hooks for tools like LLDB to interact with the GPU execution environment.

Features
The proposed implementation for the Intel GPU plugin is based on the lldb-server plugin model (@clayborg’s plugin model) including:

  • IntelGT lldb-server Plugin: The CPU target initializes an lldb-server plugin in response to an event. For Intel GT, hitting the CPU-side breakpoint on zeModuleCreate (the Level Zero entry point invoked at the first GPU kernel launch) initializes the IntelGT plugin. The plugin then creates a secondary accelerator target that debugs only the GPU, alongside the ongoing CPU process.

  • Address Space Support

  • Live Debugging: Users can get real work done as most run-control functionality is supported. However, please note that known bugs exist — use it at your own risk.

  • Offline Debug Support: Currently not supported.

  • Documentation: Currently a work in progress (WIP).

Requirements:
Linux x86_64 with i915 or Xe DRM driver with debug enabled
Intel GPU — currently supported architectures: Alchemist and Battlemage
Intel oneAPI Level Zero loader + driver
libiga64 (from intel-graphics-compiler) for disassembly

(more install info could be found at Installing LTS 2523.x — Intel® software for general purpose GPU capabilities documentation )

Example Outputs and Screenshots
Following request from @DavidSpickett , providing an example outputs and screenshot:

:backhand_index_pointing_right: LLDB IntelGT demo · GitHub

Again, we see this PR as the perfect starting point for syncing up and building upon each other’s work. We acknowledge that it is larger than a typical PR — this was necessary in order to catch up with the ongoing effort in this thread. We apologize for that and appreciate your patience. Counting on your feedback. Together, we can ensure our efforts remain aligned. Thank you all.

1 Like

Thanks!

On one hand, I expected to see something wild but actually it’s pretty familiar aside from what looks like a very large amount of threads. Now I’m curious how it will look in IDEs via. DAP, but this should “just work” :slight_smile:

Anyway, I’m not a GPU expert so I’ll just say that this work is exciting and I’m happy to see so many architectures being part of it.

I see plenty of opportunity to split it up later, so I will likely comment on that part of it but leave the real review to the experts in this thread.

I think you’re the first person to mention this actually. I’m curious how GPU core files would work, but I think it’s safe to say it doesn’t need to be part of this current effort.