Conversation
* nits * hack unrolled warp copy * Revert "nits" This reverts commit 3e1b28d.
Each thread is responsible for one target rank
* Fix mbarrier * Remove redundant store
* Suppress kineto output * Add pressure test mode * Add `x_pure_rand_e4m3` test * Add more results into hash value
* Remove redundant TMA flushes * Less barrier initialization overhead * Simplify `elect_one_sync` * Use `elect_one_sync` instead of lanes * Minor fix * Polish testing prints * Refactor for internode kernels * Better performance
* Fix hidden_size % 128 != 0 * Add `align_down()` function * Use the full warp to wait TMA store * Support arbitrary hidden sizes in fp8 cast * lint
Co-authored-by: Yifei Zhang <219273404+yifeizhang-c@users.noreply.github.com>
…rence (deepseek-ai#370) * support EP elastic shrink * Add test script for EP shrink failover * add coll buffer and allgather api * remove allgather api, merge barrier and clean kernels, rename elastic to shrink * fix bug * refine code, add shrink test script * refine code * refine code * fix bug * add/remove blank lines --------- Co-authored-by: Jiaqi Gao <jiaqi.g@alibaba-inc.com> Co-authored-by: Yichi Xu <xuyichi.xyc@alibaba-inc.com>
* Check `QP_DEPTH` for low latency kernels. * Revert the default `QP_DEPTH`
* Support EP24 for internode kernels. * Skip check for round_scale test
* Add format scripts. * Formated * Fix ruff check.
* Add format workflow * Remove debug code
* Fix OOB * Add more assertions * More checks on channels
Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com> Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>
* Compatible with older versions of python and torch * Format
…-ai#217) * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * more * add flag * add test * fix * more * apply
Signed-off-by: wangfakang <fakangwang@gmail.com>
Signed-off-by: Salman Muin Kayser Chishti <13schishti@gmail.com>
🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Contributor <contributor@example.com> Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* DeepEP/csrc: Force CUDA RDC when NVSHMEM is used. When DeepEP is against a source build of NVSHMEM starting with NVSHMEM 3.3.20, multiple instances of nvshmemi_ibgda_device_state_d are detected. One from NVSHMEM, another from the kernels. Signed-off-by: Seth Howell <sethh@nvidia.com> * DeepEP/csrc: Update RC selection for NVSHMEM 3.5. The locations of QPs within the device state struct changed with the addition of QP-specific APIs. This change updates DeepEP to work with the new layout. The ibgda_get_rc function is also templated so that main remains backwards compatible with previous versions of NVSHMEM. Tested with both internode and low_latency kernels on 4 H100 nodes. Signed-off-by: Seth Howell <sethh@nvidia.com> * Check Format Signed-off-by: Seth Howell <sethh@nvidia.com> --------- Signed-off-by: Seth Howell <sethh@nvidia.com>
…supports (deepseek-ai#605) * Merge with private repo * Fix compilation * Add readme * Update docs * Fix submodule * Fix cu128 compilation * enhance: add env:EP_NIC_NAME to config nic name (deepseek-ai#610) * enhance: add env:EP_NIC_NAME to config nic name * enhance: add env:EP_NIC_NAME to config nic name * enhance: add env:EP_NIC_NAME to config nic name * enhance: add env:EP_NIC_NAME to config nic name * enhance: add env:EP_NIC_NAME to config nic name --------- Co-authored-by: fujianhao.fjh <fujianhao.fjh@alipay.com> --------- Co-authored-by: Shangyan Zhou <sy.zhou@deepseek.com> Co-authored-by: AlphaBaby <fujianhao1997@qq.com> Co-authored-by: fujianhao.fjh <fujianhao.fjh@alipay.com>
What's New Bug Fixes Fix a rare numerical error in multi-node hybrid combine Features Add CPU buffer support for Engram Remove alignment requirement for all-gather Improvements Reuse torch NCCL comm by default to reduce GPU memory usage Improve combine/combine-epilogue performance for certain cases
* fix missing num_worst_tokens arg * fix
This commit changes the library matching, so that it avoids accidentally matching "/dev/shm/nccl-XXXXXX" shared memory mappings that might be present in /proc/self/map.
…deepseek-ai#627) When NVSHMEM and NCCL are installed via the official NVIDIA pip wheels (`nvidia-nvshmem-cu12`, `nvidia-nccl-cu12`), the wheel only ships the SONAME-suffixed library (e.g. `libnvshmem_host.so.3`, `libnccl.so.2`) and omits the unversioned `libnvshmem_host.so` / `libnccl.so` symlink that a system tarball install would create. The build then fails at link time with: /usr/bin/ld: cannot find -l:libnvshmem_host.so: No such file or directory because `-l:NAME` is exact-name matching and refuses to consider `libnvshmem_host.so.3` when only `libnvshmem_host.so` is requested. `get_nvshmem_host_lib_name()` already existed for this exact reason, but was never wired into `extra_link_args` — the link string still hard-coded `-l:libnvshmem_host.so`. NCCL had the same hard-coded `-l:libnccl.so` and the same wheel-vs-tarball mismatch, just no helper to back it up. This change: * generalises the helper into `_find_versioned_so(base_dir, prefix)`, preferring the unversioned symlink (tarball install) and falling back to the SONAME file (pip wheel install); * adds `get_nccl_lib_name` mirroring the NVSHMEM helper; * uses both helpers when assembling `extra_link_args` so the `-l:NAME` flag carries the real on-disk filename. Validated by building a wheel inside a container that has only the pip-wheel installs of `nvidia-nvshmem-cu12==3.6.5` and `nvidia-nccl-cu13==2.30.4`. After the patch the link succeeds and `ldd deep_ep/_C*.so` resolves to the wheel-shipped SONAME files. No behaviour change for tarball installs that already provide the unversioned symlink — the helper picks it up first.
…deepseek-ai#642) * Add fence.proxy.async.shared::cta between mbarrier wait and TMA load. * Move fence before "mbarrier_arrive"
Co-authored-by: ZhenghangRen <zhenghangr@nvidia.com>
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com> Co-authored-by: Rui Tian <tianr22@deepseek.com> Co-authored-by: Shengyu Liu <shengyuliu@deepseek.com> Co-authored-by: Yi Qian <qianyi@deepseek.com>
…epseek-ai#688) Store the runtime-sized device communicator behind jit::NoRefPtr so JIT launches pass its backing storage directly without host-side dereference. Co-authored-by: Katie Gioioso <kgioioso@nvidia.com>
…ns NVLink and RDMA (deepseek-ai#715) * fix: Add fence.acq_rel.sys before the GIN barrier for mixed-fabric scaleup * fix: fence mixed-fabric writes in GIN barrier --------- Co-authored-by: Shangyan Zhou <sy.zhou@deepseek.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
deepseek-ai/DeepEPmain branch into the neuralmagic fork[project]/[build-system]config and adopting upstream's lint configurationTest plan
🤖 Generated with Claude Code