Skip to content

Sync with upstream deepseek-ai/DeepEP - #1

Open
pazqo wants to merge 52 commits into
neuralmagic:mainfrom
pazqo:sync-upstream-main
Open

pazqo wants to merge 52 commits into
neuralmagic:mainfrom
pazqo:sync-upstream-main

Conversation

@pazqo

@pazqo pazqo commented Aug 26, 2026

Copy link
Copy Markdown

Summary

  • Merges latest changes from deepseek-ai/DeepEP main branch into the neuralmagic fork
  • Resolves pyproject.toml conflict by keeping neuralmagic's [project]/[build-system] config and adopting upstream's lint configuration

Test plan

  • Verify build still works with the existing pyproject.toml settings
  • Confirm no regressions from upstream changes

🤖 Generated with Claude Code

fzyzcjy and others added 30 commits September 1, 2025 17:06
* nits

* hack unrolled warp copy

* Revert "nits"

This reverts commit 3e1b28d.
Each thread is responsible for one target rank
* Fix mbarrier

* Remove redundant store
* Suppress kineto output

* Add pressure test mode

* Add `x_pure_rand_e4m3` test

* Add more results into hash value
* Remove redundant TMA flushes

* Less barrier initialization overhead

* Simplify `elect_one_sync`

* Use `elect_one_sync` instead of lanes

* Minor fix

* Polish testing prints

* Refactor for internode kernels

* Better performance
* Fix hidden_size % 128 != 0

* Add `align_down()` function

* Use the full warp to wait TMA store

* Support arbitrary hidden sizes in fp8 cast

* lint
Co-authored-by: Yifei Zhang <219273404+yifeizhang-c@users.noreply.github.com>
…rence (deepseek-ai#370)

* support EP elastic shrink

* Add test script for EP shrink failover

* add coll buffer and allgather api

* remove allgather api, merge barrier and clean kernels, rename elastic to shrink

* fix bug

* refine code, add shrink test script

* refine code

* refine code

* fix bug

* add/remove blank lines

---------

Co-authored-by: Jiaqi Gao <jiaqi.g@alibaba-inc.com>
Co-authored-by: Yichi Xu <xuyichi.xyc@alibaba-inc.com>
* Check `QP_DEPTH` for low latency kernels.

* Revert the default `QP_DEPTH`
* Support EP24 for internode kernels.

* Skip check for round_scale test
* Add format scripts.

* Formated

* Fix ruff check.
* Add format workflow

* Remove debug code
* Fix OOB

* Add more assertions

* More checks on channels
Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com>
Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com>
* Compatible with older versions of python and torch

* Format
…-ai#217)

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* more

* add flag

* add test

* fix

* more

* apply
yifeizhang-c and others added 22 commits November 5, 2025 15:33
Signed-off-by: wangfakang <fakangwang@gmail.com>
Signed-off-by: Salman Muin Kayser Chishti <13schishti@gmail.com>
🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Contributor <contributor@example.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* DeepEP/csrc: Force CUDA RDC when NVSHMEM is used.

When DeepEP is against a source build of NVSHMEM starting
with NVSHMEM 3.3.20, multiple instances of nvshmemi_ibgda_device_state_d
are detected. One from NVSHMEM, another from the kernels.

Signed-off-by: Seth Howell <sethh@nvidia.com>

* DeepEP/csrc: Update RC selection for NVSHMEM 3.5.

The locations of QPs within the device state struct changed
with the addition of QP-specific APIs.

This change updates DeepEP to work with the new layout.

The ibgda_get_rc function is also templated so that main remains
backwards compatible with previous versions of NVSHMEM.

Tested with both internode and low_latency kernels on 4 H100 nodes.

Signed-off-by: Seth Howell <sethh@nvidia.com>

* Check Format

Signed-off-by: Seth Howell <sethh@nvidia.com>

---------

Signed-off-by: Seth Howell <sethh@nvidia.com>
…supports (deepseek-ai#605)

* Merge with private repo

* Fix compilation

* Add readme

* Update docs

* Fix submodule

* Fix cu128 compilation

* enhance: add env:EP_NIC_NAME to config nic name (deepseek-ai#610)

* enhance: add env:EP_NIC_NAME to config nic name

* enhance: add env:EP_NIC_NAME to config nic name

* enhance: add env:EP_NIC_NAME to config nic name

* enhance: add env:EP_NIC_NAME to config nic name

* enhance: add env:EP_NIC_NAME to config nic name

---------

Co-authored-by: fujianhao.fjh <fujianhao.fjh@alipay.com>

---------

Co-authored-by: Shangyan Zhou <sy.zhou@deepseek.com>
Co-authored-by: AlphaBaby <fujianhao1997@qq.com>
Co-authored-by: fujianhao.fjh <fujianhao.fjh@alipay.com>
What's New
Bug Fixes

Fix a rare numerical error in multi-node hybrid combine
Features

Add CPU buffer support for Engram
Remove alignment requirement for all-gather
Improvements

Reuse torch NCCL comm by default to reduce GPU memory usage
Improve combine/combine-epilogue performance for certain cases
* fix missing num_worst_tokens arg

* fix
This commit changes the library matching, so that it avoids accidentally
matching "/dev/shm/nccl-XXXXXX" shared memory mappings that might be present in
/proc/self/map.
…deepseek-ai#627)

When NVSHMEM and NCCL are installed via the official NVIDIA pip wheels
(`nvidia-nvshmem-cu12`, `nvidia-nccl-cu12`), the wheel only ships the
SONAME-suffixed library (e.g. `libnvshmem_host.so.3`, `libnccl.so.2`)
and omits the unversioned `libnvshmem_host.so` / `libnccl.so` symlink
that a system tarball install would create.

The build then fails at link time with:
  /usr/bin/ld: cannot find -l:libnvshmem_host.so: No such file or
directory
because `-l:NAME` is exact-name matching and refuses to consider
`libnvshmem_host.so.3` when only `libnvshmem_host.so` is requested.

`get_nvshmem_host_lib_name()` already existed for this exact reason,
but was never wired into `extra_link_args` — the link string still
hard-coded `-l:libnvshmem_host.so`. NCCL had the same hard-coded
`-l:libnccl.so` and the same wheel-vs-tarball mismatch, just no helper
to back it up.

This change:
  * generalises the helper into `_find_versioned_so(base_dir, prefix)`,
    preferring the unversioned symlink (tarball install) and falling
    back to the SONAME file (pip wheel install);
  * adds `get_nccl_lib_name` mirroring the NVSHMEM helper;
  * uses both helpers when assembling `extra_link_args` so the
    `-l:NAME` flag carries the real on-disk filename.

Validated by building a wheel inside a container that has only the
pip-wheel installs of `nvidia-nvshmem-cu12==3.6.5` and
`nvidia-nccl-cu13==2.30.4`. After the patch the link succeeds and
`ldd deep_ep/_C*.so` resolves to the wheel-shipped SONAME files.

No behaviour change for tarball installs that already provide the
unversioned symlink — the helper picks it up first.
…deepseek-ai#642)

* Add fence.proxy.async.shared::cta between mbarrier wait and TMA load.

* Move fence before "mbarrier_arrive"
Co-authored-by: ZhenghangRen <zhenghangr@nvidia.com>
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
Co-authored-by: Rui Tian <tianr22@deepseek.com>
Co-authored-by: Shengyu Liu <shengyuliu@deepseek.com>
Co-authored-by: Yi Qian <qianyi@deepseek.com>
…epseek-ai#688)

Store the runtime-sized device communicator behind jit::NoRefPtr so JIT launches pass its backing storage directly without host-side dereference.

Co-authored-by: Katie Gioioso <kgioioso@nvidia.com>
…ns NVLink and RDMA (deepseek-ai#715)

* fix: Add fence.acq_rel.sys before the GIN barrier for mixed-fabric scaleup

* fix: fence mixed-fabric writes in GIN barrier

---------

Co-authored-by: Shangyan Zhou <sy.zhou@deepseek.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.