Nova: An End-to-End MLIR Compiler for Deep Learning
06:00 · August 4, 2026 · arXiv cs.AI RSS

The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and granular control over hardware and memory required to maximize physical hardware utilization natively. To bridge this gap, we designed Nova, an automated end-to-end JIT compiler whose defining purpose is to achieve absolute control over this hardware mapping: fusing operations across operation boundaries, optimizing complex memory hierarchies, and tuning execution down to the register level. By capturing eager executions and unifying forward and backward passes into a single value-semantic dialect, Nova unlocks aggressive whole-graph optimizations. It then utilizes an Analytic Configurator to deterministically derive optimal execution schedules based on arithmetic intensity, dropping search time to zero. Backed by a structural hashing runtime, Nova synthesizes fine-grained kernels directly from the computation's structure. In our evaluations on an RTX 3060, Nova matches or modestly exceeds cuBLAS and XLA on TF32 matmuls on most shapes, maintaining a stringent < 5e-4 relative error. At the model level, Nova achieves up to 10.6% greater throughput than PyTorch and 4.4% greater than XLA on a 42-million parameter model, without compromising on numerical fidelity. Crucially, by reducing the memory footprint by up to 29% relative to PyTorch, Nova successfully trains a 144-million parameter model at 17,900 tokens/s where PyTorch encounters Out-Of-Memory (OOM) failures on the same 12 GB consumer GPU.
Summary
Nova is an end-to-end JIT compiler built on MLIR that translates eager deep-learning code into optimized machine instructions for training. It addresses the limitations of tensor frameworks whose separate kernel launches and lack of whole-graph visibility prevent cross-operator fusion and force intermediate results back to global memory. Instead, Nova captures both forward activations and backward gradients during a training step, then lowers the computation through a progressive sequence of dialects that preserve global optimizations down to register-level scheduling on the target GPU.
The compiler introduces the nova dialect, a value-semantic intermediate representation that unifies the forward and backward passes into a single execution block. An Analytic Configurator reads device parameters such as arithmetic intensity and memory hierarchy limits to derive tile sizes, warp mappings, and MMA intrinsics deterministically, eliminating autotuning search. A structural-hashing runtime then emits fine-grained kernels directly from the computation graph, while a strict caching layer ensures compilation occurs only once per distinct structure.
On an RTX 3060, the resulting kernels match or modestly exceed cuBLAS and XLA throughput on TF32 matrix multiplications for most shapes while keeping relative error below 5e-4. At the model level, Nova delivers up to 10.6 percent higher throughput than PyTorch and 4.4 percent higher than XLA on a 42-million-parameter network. Its memory footprint is reduced by as much as 29 percent relative to PyTorch, enabling a 144-million-parameter model to train at 17,900 tokens per second on a 12 GB consumer GPU where the baseline framework encounters out-of-memory failures.
Why it matters
Nova's approach to maximizing GPU efficiency and reducing memory overhead is highly relevant for Dutch AI researchers and SMEs aiming to train models cost-effectively and sustainably. Its deep technical insights into MLIR and hardware-aware optimizations provide actionable knowledge for advancing AI infrastructure and Green AI initiatives in the Netherlands.









