Skip to content

Support 1 * 128 and 128 * 128 block-wise quant? #38

Description

@zfan2356

In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quantize approach. I wonder do we have any plan for support this approach?

https://docs.nvidia.com/cuda/cublas/index.html#cublasltmatmulmatrixscale-t

Activity

  1. self-assigned this
    on Aug 6, 2025
  2. szkarpinski commented on Aug 6, 2025

    @szkarpinski

    Hi @zfan2356 , thank you for raising this issue. The lower-level nvmath.bindings for those scaling types should be included in the next release of nvmath-python. I'll also start working on adding support to our higher-level host APIs in nvmath.linalg.advanced.matmul.

  3. zfan2356 commented on Aug 11, 2025

    @zfan2356
    Author

    @szkarpinski Thanks for your reply! I also recommend supporting the pointer array mode for Group GEMM. Both of these features are commonly used in modern LLM pretrain. nvmath-python is a user-friendly Python library, and I hope it can support these two features as soon as possible. Thank you very much!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions