Skip to content

Julia JIT support for device APIs #32

Description

@ggkountouras

As header-only libraries based on CUTLASS, these only work in C++.

It would be nice to have e.g. the new DGEMM via IMMA (https://github.com/NVIDIA/CUDALibrarySamples/tree/master/MathDx/cuBLASDx/16_dgemm_emulation) inside Julia kernels.

Unfortunately, Warp/numba can only JIT compile a subset of python types, and not the (more general) Julia types that we require (e.g. differential equation solvers).

A potential way forward would be to compile libmathdx to PTX, link it in via LLVM, and get full performance via LTO. However, these often contain NVVM IR, which LLVM cannot handle.

Related discussion: https://discourse.julialang.org/t/using-cublasdx-in-julia/125527

Activity

  1. ZzEeKkAa commented on Jun 24, 2025

    @ZzEeKkAa
    Contributor

    Hi @ggkountouras! I'm not very familiar with Julia, but this is how we bring mathdx libraries to python using libmathdx:

    1. compile separately mathdx types and functions as an LTO(s) through libmathdx
    2. compile numba-cuda kernel with function references to LTO from previous step.
    3. Link kernel and and libmathdx produced LTO using libnvvm
    4. Dispatch the kernel

    Please let me know if the same approach would work at Julia.

    BTW, we have DGEMM via IMMA implementation already available at nvmath: https://github.com/NVIDIA/nvmath-python/blob/main/examples/device/cublasdx_fp64_emulation.py

  2. ggkountouras commented on Jul 14, 2025

    @ggkountouras
    Author

    It would work, with changes to the Julia compiler to support libNVVM.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions