Skip to content

Fix legacy RC QP indexing for NVSHMEM 3.5.19+ - #7

Open
heyselbi wants to merge 3 commits into
mainfrom
heyselbi/fix-nvshmem-rc-layout
Open

heyselbi wants to merge 3 commits into
mainfrom
heyselbi/fix-nvshmem-rc-layout

Conversation

@heyselbi

@heyselbi heyselbi commented Sep 30, 2026 •

Copy link
Copy Markdown

Summary

  • Port deepseek-ai/DeepEP#696 / the fix vendored by vllm-project/vllm#58159 into midstream DeepEP
  • Fix legacy IBGDA ibgda_get_rc indexing for NVSHMEM 3.5.19+ (QP-major, PE-interleaved RC QP layout); keep PE-major indexing for older NVSHMEM
  • Add compile-time version gates plus a runtime host/device layout compatibility check before NVSHMEM init to avoid CUDA illegal memory access on first RDMA dispatch
  • Bump midstream version to 2.0.1+rhaiv.5 (setup.py + pyproject.toml) so the next tag matches the wheel version (ADR-170)

Context

DeepEP v1 legacy low-latency / high-throughput all2all IBGDA kernels used the pre-3.5.19 RC QP layout. Building against NVSHMEM 3.5.19+ (including PyTorch 2.15 / NVSHMEM 3.7) can fault on first RDMA dispatch. The original fix in deepseek-ai#564 was lost in the DeepEPv2 upgrade; deepseek-ai#696 restores it.

Once this lands, cut tag v2.0.1+rhaiv.5. vLLM can then drop its temporary build-time patch once it pins that revision.

Test plan

  • Confirm patch matches upstream DeepEP#696 / vLLM vendored deepep_nvshmem_rc_qp.patch
  • Build against NVSHMEM ≥3.5.19 and exercise multi-node dispatch/combine (no IMA / hang)
  • Confirm wheel version resolves to 2.0.1+rhaiv.5
  • Optionally verify older NVSHMEM (<3.5.19) still uses PE-major indexing path

heyselbi and others added 2 commits September 30, 2026 16:17
NVSHMEM 3.5.19+ stores IBGDA RC QPs in QP-major order; use that layout
at compile time and reject mismatched host/device NVSHMEM versions at
init. Port of deepseek-ai#696 (also vendored in vllm-project/vllm#58159).

Co-authored-by: caoxiaoyuyuyuyuyu <170225252+caoxiaoyuyuyuyuyu@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Bump midstream local suffix so the next tag after the NVSHMEM RC QP
layout fix matches the wheel version (ADR-170).

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@heyselbi

heyselbi commented Oct 1, 2026

Copy link
Copy Markdown
Author

Note for anyone rebuilding this against a stripped production RHAIIS image (e.g. registry.redhat.io/rhaii/vllm-cuda-rhel9):

Building DeepEP via uv pip install --no-build-isolation against the torch bundled in those images can hit half/bf16 operator errors in nvshmem.cu / NVSHMEM's reduce.cuh. That comes from PyTorch's cpp_extension.py injecting -D__CUDA_NO_HALF_OPERATORS__ (and related flags), which conflicts with NVSHMEM's unconditional half/bf16 reduce macros.

AIPCC's fondue/builder path does not need this — leave nvcc_flags as-is for the real midstream build. For a local rebuild only, appending these to setup.py's nvcc_flags unblocks it:

'-U__CUDA_NO_HALF_OPERATORS__',
'-U__CUDA_NO_HALF_CONVERSIONS__',
'-U__CUDA_NO_BFLOAT16_CONVERSIONS__',
'-U__CUDA_NO_BFLOAT16_OPERATORS__',
'-U__CUDA_NO_HALF2_OPERATORS__',

Those production images are also runtime-only (no -devel headers for NVSHMEM/NCCL/cuBLAS/cuSPARSE/cuSOLVER/NVRTC, and no mlx5dv.h). Matching nvidia-*-cu13 wheels from PyPI plus exact-version devel RPMs from NVIDIA's public CUDA repo (header extract only) are enough to get a local rebuild working; again, not needed for AIPCC.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants