Skip to content

Fix/win ninja tl2 q6k - #627

Open
manuelmen2a-blip wants to merge 12 commits into
microsoft:mainfrom
manuelmen2a-blip:fix/win-ninja-tl2-q6k
Open

manuelmen2a-blip wants to merge 12 commits into
microsoft:mainfrom
manuelmen2a-blip:fix/win-ninja-tl2-q6k

Conversation

@manuelmen2a-blip

@manuelmen2a-blip manuelmen2a-blip commented Sep 13, 2026 •

Copy link
Copy Markdown

Remove dead/non-building code + add portable Windows/Ninja CI

Removed (dead code)

  • gpu/bitnet_kernels/gemm_native.cu (69 lines): no references anywhere in
    the tree (not in any CMakeLists.txt, compile.sh/setup.py; grep for
    gemm_native/native_gemm/ggml_bitnet_gemm under gpu/ → 0 matches).
    Never compiled, never called; its kernel also produced numerically wrong
    results in local testing.
  • src/ggml-bitnet-lut.cpp — TL2 x86 ggml_bitnet_mul_mat (46 lines):
    calls ggml_bitnet_mul_mat_task_init/ggml_bitnet_mul_mat_task_compute,
    which are declared in include/ggml-bitnet.h but defined nowhere in the repo,
    so any build with BITNET_X86_TL2 fails to link. The working path is I2_S,
    not TL2.

Changed (build portability)

  • CMakeLists.txt: ggml-base arch is now a cache variable
    BITNET_GGML_ARCH (default native) instead of a hardcoded
    -march=native. -march=native is not portable across CI runners (one
    runner hit undeclared AVX-512 intrinsics in ggml.c); CI builds with
    -DBITNET_GGML_ARCH=x86-64-v3 (the AVX2 microarchitecture level).
  • .github/workflows/ci.yml: new GitHub Actions workflow
    (windows-latest, clang + Ninja) that configures, builds, and smoke-tests
    llama-cli --version (which loads ggml-base.dll, so a link failure
    surfaces). Configure flags: -DBITNET_GGML_ARCH=x86-64-v3 -DLLAMA_BUILD_COMMON=ON -DLLAMA_BUILD_EXAMPLES=ON -DLLAMA_BUILD_TOOLS=ON.

Kept (unchanged)

  • The Win/Ninja I2_S link fix: src/ggml-bitnet-mad.cpp and
    src/ggml-bitnet-shim.c compiled directly into ggml-base via
    target_sources (fixes the quantize_i2_s / dequantize_row_i2_s
    undefined-symbol link error caused by GGML_SOURCES_BITNET never being
    consumed).
  • bitlinear_int8xint2_batched in gpu/bitnet_kernels/bitnet_kernels.cu
    (referenced by gpu/model.py:107 for the batched prefill path).

Verification

  • CI green on the fork (manuelmen2a-blip/BitNet, run
    36258669883): Configure → Build → Smoke test all pass (7m47s).
  • llama-cli.exe --version exits 0 and loads ggml-base.dll.
  • The set of compiled translation units is unchanged versus the previous
    working I2_S build (only unreferenced code was removed).

Test plan

  • cmake --build build configures, compiles and links on Windows/Ninja
  • ggml-bitnet-mad.cpp compiled with an AVX2-level ISA (x86-64-v3)
  • BITNET_X86_TL2 path no longer carries a file that cannot link
  • llama-cli --version smoke test passes
  • Cross-repo PR CI needs maintainer approval to run on
    microsoft/BitNet (GitHub runs base-branch workflows only for fork PRs)

🤖 Generated with Claude Code

@manuelmen2a-blip

Copy link
Copy Markdown
Author

fix: win/ninja I2_S build + TL2 x86 mul_mat + GPU int2 prefill + GEMM M>1

@manuelmen2a-blip

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

bitnet-pruebas added 6 commits September 26, 2026 10:17
gemm_native.cu had no references anywhere in the tree and its kernel produced numerically wrong results. The TL2 x86 mul_mat added in bc3b20c calls ggml_bitnet_mul_mat_task_init/compute, which are only declared in include/ggml-bitnet.h and defined nowhere in the repo, so any build with BITNET_X86_TL2 fails to link. The working path (Win/Ninja I2_S + GPU int2 prefill fallback) is unchanged.
Verify the Windows/Ninja clang build configures, compiles and links, then smoke-test the produced binary (llama-cli --version loads ggml-base.dll, so a link failure surfaces here). Triggers on PRs and pushes to main.
LLAMA_BUILD_EXAMPLES defaults to LLAMA_STANDALONE (OFF in a fork), so llama-cli.exe was never built and the smoke test failed on a missing binary. Enable LLAMA_BUILD_EXAMPLES/LLAMA_BUILD_TOOLS.
-march=native is not portable across CI runners (one runner hit undeclared AVX-512 intrinsics in ggml.c). Add a BITNET_GGML_ARCH cache var (default native) and have CI build with -DBITNET_GGML_ARCH=avx2.
-march=avx2 is invalid (-march expects a CPU name, not an ISA). Use x86-64-v3, the AVX2 microarchitecture level, which is a valid -march target on clang 12+.
Examples and tools are gated on LLAMA_BUILD_COMMON (llama.cpp/CMakeLists.txt:212,217), which defaulted to OFF, so llama-cli.exe was never produced. Enable it alongside EXAMPLES/TOOLS.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant