One year has passed since we started implementing the infrastructure in Lighthouse, and we’d like to give the community an update on the state it is now, and what are the next steps.
Co-written by: @asiemien @teekarna
Initial Goals
When we started the project, we set out to cover a number of initial goals.
The first goal was to test and validate assumptions about using MLIR and create a common infrastructure to utilize it. Import programs into MLIR, create various pipelines and transform schedules, then execute them on actual hardware. We have achieved that goal a few months back by integrating KernelBench into a set of pipelines and benchmarked them locally on CPU and GPU.
The second (and current) goal is to find common patterns in those pipelines and start combining the strategies in a user-facing manner, so that people can build their own compilers without building them from scratch and utilizing well known transforms and pipelines for common patterns. We are pursuing that goal by integrating Lighthouse into two existing projects – AI-Bench and TPP-MLIR – as a demonstration for other projects.
Current State
The KernelBench tool can run and validate output for most of the first 200 kernels (~80%) on CPU, with remaining failures being tracked with issues in the Lighthouse repo. Bugs are mostly in torch-mlir and are being slowly fixed as we go.
However, just having a tool that runs specific kernels doesn’t show the main expectation above, so we have now integrated the same infrastructure into a third-party benchmark harness (AI-Bench). The performance between Lighthouse and AI-Bench is almost identical, which means any improvements we do in Lighthouse will translate directly into its users.
On CPUs, the overall performance isn’t on par with existing MLIR implementations yet. We compare it with our own efforts (e.g., TPP-MLIR), and we’re getting around 60%~80% of the reference performance on average. This is mainly because Lighthouse doesn’t yet identify all the special high-level patterns that these other tools do. The lower level optimizations are identical (upstream), when the patterns match.
On GPUs, we have implemented lowering pipelines and kernel execution infrastructure for Intel GPUs. We have a kernel auto-tuning framework in development for GEMM-like kernels. We can currently lower and execute most GEMM and elementwise kernels in KernelBench, as well as some layer norms. The performance of tuned kernels is on par with reference implementations (e.g., oneDNN).
One key long-term goal with Lighthouse is to move all that knowledge upstream. Transforms, special patterns, and matchers should go into MLIR, whereas the glue code, auto-tuning, and pipeline building logic should stay in Lighthouse. In principle, none of the latter should be hosted in (public facing) downstream projects.
We have a pre-commit CI loop that runs all examples and tests on Linux x86 and arm CPU architectures. We also have a post-commit CI that runs on XeGPU targets. We’d like to expand that matrix to more CPU and GPU architectures, as well as more operating systems.
How to Use
Lighthouse hasn’t been designed as a full fledge compiler, but as a module that wraps the MLIR compiler framework and exposes to its users.
The library itself provides:
- Module imports: e.g. using Torch-MLIR to convert PyTorch models into MLIR modules.
- Execution environments: Using MLIR’s JIT functionality, which can be driven either directly (via a compiler driver) or as a Torch Dynamo plugin (via
@torch.compileannotation). - Dialect creation: wrapped from IRDL, allows users to create dialects and operations directly in Python (and usable within Lighthouse).
- Transforms and schedules: Using the infrastructure above, allows us to create transforms operations and schedules based on those and upstream transforms, to provide common patterns for code transformation.
- Pipeline schedules: Parametrized YAML files that can reuse upstream passes & transforms as well as Lighthouse ones, that users can compose to create entire pipelines on their compilers.
- Auto-tuners: Experimental modules that provide some auto-tuning functionality.
All of that is used by the internal tool kernel-bench, and tested here. This gives users a good idea how to use the modules, schedules and import their models into MLIR.
AI-Bench Usage
AI-bench is a framework for executing KernelBench PyTorch kernels and benchmarking their performance across multiple backends and devices. For MLIR execution, it relies on Lighthouse’s MLIR backend through a lean wrapper integrated with torch.compile. In this flow, TorchDynamo captures the PyTorch model and dispatches the resulting graph to the Lighthouse backend, which imports it into Linalg-on-Tensors and lowers it for execution on CPU and Intel GPU targets.
Most of the integration is provided by the Lighthouse’s compiler and runner stack, with AI-bench adding only the glue needed to configure the target-specific lowering and runtime environment. It then connects the MLIR path to the common correctness and performance infrastructure, allowing MLIR kernels to use the same validation and benchmarking flow as the other backend options supported by AI-bench. In particular, the MLIR route operates directly on the original KernelBench model, without requiring a separate MLIR-specific model implementation.
TPP-MLIR Usage
In TPP-MLIR, our main concern is importing PyTorch models into MLIR, so we can consume in our compiler. For that, we have added an import tool that wraps Lighthouse Torch-MLIR direct import, mimicking the kernel-bench functionality, to get Linalg-on-Tensor graphs into our compiler.
This has shown some repeat code across the three projects, and perhaps opportunities to simplify the import process. But only because we have followed the KernelBench model building strategy (for args and parameters), so it’s not necessarily a generic observation.
XeGPU Auto Tuning
We have added a preliminary GEMM autotunning tool for Intel GPUs. The tool uses a cost model to generate prominent workgroup, subgroup and reduction tile configurations and valid thread-cooperative prefetch strategies for A and B. The best configurations can be fine-tuned in subsequent iterations (e.g., optimizing prefetch depth). Typically, GEMM tuning takes a couple of minutes and achieves 80% to 100% of the SOTA performance.
Next Steps
The major items being worked on right now are:
- Improvement on the reuse and extensibility of pipelines: From YAML parametrization to transform schedule reuse, to hardware heuristics and cost modelling. The goal is to have mostly upstream transforms, parametrized by the (given or auto-detected) target architecture via some simple cost modelling logic in Lighthouse.
- Improvement on pattern complexity: Extending decision making to more complex graphs, allowing auto-tuned parameters to span across multiple layers (similar sub-graphs with different shapes), and compose that with the pipeline above.
We do not have the bandwidth to add new targets other than Intel ones. So, this is a call for action: help us test and extend this to run on other CPU and GPU architectures. (@banach-space @ftynse @Matthias-springer)
We’d also like to encourage usage, including outside of our comfort zone.
Acknowledgement
We’d like to thank all contributors to the project for the efforts so far. We hope more can join the effort to expand usage, test coverage and new technologies.
@rengolin @asiemien @teekarna @fschlimb @charitha22 @kurapov-peter @jopperm @rolfmorel @nbpatel @dchigarev @MattPD @FedericoBruzzone