Fix/win ninja tl2 q6k - #627
Open
manuelmen2a-blip wants to merge 12 commits into
Open
manuelmen2a-blip wants to merge 12 commits into
manuelmen2a-blip wants to merge 12 commits into
Conversation
added 6 commits
September 12, 2026 15:38
…shim dequantize_row_i2_s
Author
|
fix: win/ninja I2_S build + TL2 x86 mul_mat + GPU int2 prefill + GEMM M>1 |
Author
|
@microsoft-github-policy-service agree |
added 6 commits
September 26, 2026 10:17
gemm_native.cu had no references anywhere in the tree and its kernel produced numerically wrong results. The TL2 x86 mul_mat added in bc3b20c calls ggml_bitnet_mul_mat_task_init/compute, which are only declared in include/ggml-bitnet.h and defined nowhere in the repo, so any build with BITNET_X86_TL2 fails to link. The working path (Win/Ninja I2_S + GPU int2 prefill fallback) is unchanged.
Verify the Windows/Ninja clang build configures, compiles and links, then smoke-test the produced binary (llama-cli --version loads ggml-base.dll, so a link failure surfaces here). Triggers on PRs and pushes to main.
LLAMA_BUILD_EXAMPLES defaults to LLAMA_STANDALONE (OFF in a fork), so llama-cli.exe was never built and the smoke test failed on a missing binary. Enable LLAMA_BUILD_EXAMPLES/LLAMA_BUILD_TOOLS.
-march=native is not portable across CI runners (one runner hit undeclared AVX-512 intrinsics in ggml.c). Add a BITNET_GGML_ARCH cache var (default native) and have CI build with -DBITNET_GGML_ARCH=avx2.
-march=avx2 is invalid (-march expects a CPU name, not an ISA). Use x86-64-v3, the AVX2 microarchitecture level, which is a valid -march target on clang 12+.
Examples and tools are gated on LLAMA_BUILD_COMMON (llama.cpp/CMakeLists.txt:212,217), which defaulted to OFF, so llama-cli.exe was never produced. Enable it alongside EXAMPLES/TOOLS.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Remove dead/non-building code + add portable Windows/Ninja CI
Removed (dead code)
gpu/bitnet_kernels/gemm_native.cu(69 lines): no references anywhere inthe tree (not in any
CMakeLists.txt,compile.sh/setup.py;grepforgemm_native/native_gemm/ggml_bitnet_gemmundergpu/→ 0 matches).Never compiled, never called; its kernel also produced numerically wrong
results in local testing.
src/ggml-bitnet-lut.cpp— TL2 x86ggml_bitnet_mul_mat(46 lines):calls
ggml_bitnet_mul_mat_task_init/ggml_bitnet_mul_mat_task_compute,which are declared in
include/ggml-bitnet.hbut defined nowhere in the repo,so any build with
BITNET_X86_TL2fails to link. The working path is I2_S,not TL2.
Changed (build portability)
CMakeLists.txt:ggml-basearch is now a cache variableBITNET_GGML_ARCH(defaultnative) instead of a hardcoded-march=native.-march=nativeis not portable across CI runners (onerunner hit undeclared AVX-512 intrinsics in
ggml.c); CI builds with-DBITNET_GGML_ARCH=x86-64-v3(the AVX2 microarchitecture level)..github/workflows/ci.yml: new GitHub Actions workflow(
windows-latest, clang + Ninja) that configures, builds, and smoke-testsllama-cli --version(which loadsggml-base.dll, so a link failuresurfaces). Configure flags:
-DBITNET_GGML_ARCH=x86-64-v3 -DLLAMA_BUILD_COMMON=ON -DLLAMA_BUILD_EXAMPLES=ON -DLLAMA_BUILD_TOOLS=ON.Kept (unchanged)
src/ggml-bitnet-mad.cppandsrc/ggml-bitnet-shim.ccompiled directly intoggml-baseviatarget_sources(fixes thequantize_i2_s/dequantize_row_i2_sundefined-symbol link error caused by
GGML_SOURCES_BITNETnever beingconsumed).
bitlinear_int8xint2_batchedingpu/bitnet_kernels/bitnet_kernels.cu(referenced by
gpu/model.py:107for the batched prefill path).Verification
manuelmen2a-blip/BitNet, run36258669883): Configure → Build → Smoke test all pass (7m47s).llama-cli.exe --versionexits 0 and loadsggml-base.dll.working I2_S build (only unreferenced code was removed).
Test plan
cmake --build buildconfigures, compiles and links on Windows/Ninjaggml-bitnet-mad.cppcompiled with an AVX2-level ISA (x86-64-v3)BITNET_X86_TL2path no longer carries a file that cannot linkllama-cli --versionsmoke test passesmicrosoft/BitNet(GitHub runs base-branch workflows only for fork PRs)🤖 Generated with Claude Code