Skip to content

Add fixed MatMul scheduling controls (step 2 of #1002) - #1047

Open
Mikyx-1 wants to merge 2 commits into
google:devfrom
Mikyx-1:feat/matmul-scheduling
Open

Mikyx-1 wants to merge 2 commits into
google:devfrom
Mikyx-1:feat/matmul-scheduling

Conversation

@Mikyx-1

@Mikyx-1 Mikyx-1 commented Oct 3, 2026

Copy link
Copy Markdown

Implements step 2 of the six-PR series proposed in #1002, following step 1 in #1041.

Timing-dependent MatMul autotuning can change floating-point results between otherwise identical evaluation runs. This adds fixed scheduling controls so baseline-versus-target comparisons can isolate changes from subsequent quantization work.

  • Adds --matmul_schedule auto|fixed|fixed_min_k, keeping auto as the default. fixed selects the first legal candidate; fixed_min_k selects the candidate with the fewest K partitions, preserving candidate order for ties.
  • Follows the review feedback on #1003: select a single candidate in matmul.cc, without adding scheduling mode state to MMAutoTune. Single-candidate tuning completes after its first measurement.
  • Applies the policy to activation conversion and optional oneDNN BRGeMM, and makes it immutable within a MatMulEnv.
  • Reserves distinct tuning keys for BF16, row-A8, and block-A8 activation representations.
  • Adds CLI, candidate-selection, bucket-order, and repeated-output tests, plus usage documentation. Fixes a null-pointer zero-length copy exposed by sanitizers and supplies Highway test helpers when its own test suite is disabled.

Fixed schedules may be slower. Reproducibility requires the same binary, hardware, topology/thread settings, weights, and inputs; --deterministic continues to control sampling separately.

Dependency and review scope

#1041 is still open, so this branch is based on its head and targets dev. The full PR diff currently includes step 1. The changes specific to step 2 are isolated in commit 8716968.

Validation

  • Release builds passed for MatMul/argument/model-comparison tests, MMLU, benchmark, Gemma CLI/API server, and DeepSeek runner; the hello-world example passed a syntax check.
  • 31 CTest checks passed, including the 18-test Python comparison suite.
  • 7 focused MatMul tests passed with AddressSanitizer, UndefinedBehaviorSanitizer, and debug assertions on AVX2.
  • Gemma 3 270M SFP: two separate runs over all 83 MMLU questions for each of fixed and fixed_min_k. Both modes produced exactly zero full-vocabulary KL divergence for every question, identical label logits, no prediction/answer changes, and 20/83 accuracy in each run.
  • Optional oneDNN BRGeMM compilation and candidate-selection checks passed; its runtime/AMX path was not exercised.
  • Bazel was unavailable, so the updated Bazel dependency was not tested.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant