Conversation
NVSHMEM 3.5.19+ stores IBGDA RC QPs in QP-major order; use that layout at compile time and reject mismatched host/device NVSHMEM versions at init. Port of deepseek-ai#696 (also vendored in vllm-project/vllm#58159). Co-authored-by: caoxiaoyuyuyuyuyu <170225252+caoxiaoyuyuyuyuyu@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Bump midstream local suffix so the next tag after the NVSHMEM RC QP layout fix matches the wheel version (ADR-170). Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
Note for anyone rebuilding this against a stripped production RHAIIS image (e.g. Building DeepEP via AIPCC's fondue/builder path does not need this — leave '-U__CUDA_NO_HALF_OPERATORS__',
'-U__CUDA_NO_HALF_CONVERSIONS__',
'-U__CUDA_NO_BFLOAT16_CONVERSIONS__',
'-U__CUDA_NO_BFLOAT16_OPERATORS__',
'-U__CUDA_NO_HALF2_OPERATORS__',Those production images are also runtime-only (no |
Summary
ibgda_get_rcindexing for NVSHMEM 3.5.19+ (QP-major, PE-interleaved RC QP layout); keep PE-major indexing for older NVSHMEM2.0.1+rhaiv.5(setup.py+pyproject.toml) so the next tag matches the wheel version (ADR-170)Context
DeepEP v1 legacy low-latency / high-throughput all2all IBGDA kernels used the pre-3.5.19 RC QP layout. Building against NVSHMEM 3.5.19+ (including PyTorch 2.15 / NVSHMEM 3.7) can fault on first RDMA dispatch. The original fix in deepseek-ai#564 was lost in the DeepEPv2 upgrade; deepseek-ai#696 restores it.
Once this lands, cut tag
v2.0.1+rhaiv.5. vLLM can then drop its temporary build-time patch once it pins that revision.Test plan
deepep_nvshmem_rc_qp.patch2.0.1+rhaiv.5