gh-158283: Force dispatch tail duplication with Clang 19 and Apple clang 17 - #158286
Open
matthiasgoergens wants to merge 1 commit into
Open
matthiasgoergens wants to merge 1 commit into
matthiasgoergens wants to merge 1 commit into
Conversation
matthiasgoergens
requested review from
AA-Turner,
corona10,
emmatyping,
erlend-aasland and
itamaro
as code owners
September 27, 2026 14:35
matthiasgoergens
pushed a commit
to matthiasgoergens/cpython
that referenced
this pull request
Sep 27, 2026
…ple clang 17 LLVM 19 limits tail duplication of blocks ending in an indirect branch (llvm/llvm-project#78582), so the computed-goto interpreter is compiled with a single shared dispatch jump instead of one per instruction. Apple clang from Xcode 16.3-16.4 has the same bug; Xcode 26.0-26.3 merges most of the dispatch jumps. LLVM 20.1.1 fixed it (llvm/llvm-project#114990). Detect the affected compilers and pass -mllvm -tail-dup-pred-size=1000 when compiling ceval.c, and to the linker's LTO backend under --with-lto. With Clang 19 this made pyperformance 8.4-8.7% faster (PGO+LTO, and thin LTO without PGO).
matthiasgoergens
force-pushed
the
pr/clang19-dispatch
branch
from
September 27, 2026 22:46
384d72b to
65772c6
Compare
matthiasgoergens
pushed a commit
to matthiasgoergens/cpython
that referenced
this pull request
Sep 27, 2026
matthiasgoergens
added a commit
to matthiasgoergens/cpython
that referenced
this pull request
Sep 28, 2026
GitHub's ubuntu-24.04 runners come with different CPUs (EPYC 7763, 9V74, 9V45, Xeon Platinum 8573C, Xeon 6973P-C were all seen), and the regex_effbot slowdown of pythongh-158286 is largest on EPYC 7763. blockbench now records the CPU model in every row, and `analyze --cpu REGEX` restricts to matching blocks. backfill_cpu.py recovers the CPU of older runs from job logs while GitHub still has them; all-cpu.jsonl for exp3, exp9b and exp11b were made with it. build_arms.py accepts "makefile_sed" to edit the generated Makefile (and fails if an expression changes nothing). exp12-sre uses it to compile Modules/_sre/sre.c with -fno-lto, to test whether keeping the regex engine's dispatch out of the link-time option avoids the regression. make_branch.sh now pushes to $PERF_CI_REMOTE without --force: in a clone where origin is python/cpython, the old `git push -f origin` would have targeted upstream. Also drop a stray .pyc and ignore __pycache__.
matthiasgoergens
added a commit
to matthiasgoergens/cpython
that referenced
this pull request
Sep 28, 2026
…ostly layout Clang 19 main / PR / PR + asm goto _sre, each linked with lld --shuffle-sections under 3 seeds, 11 jobs (33 blocks). Pooled over seeds, regex_effbot is +4.9% with the PR on EPYC 7763 and the regex geomean +0.2% (n.s.). Per seed on EPYC 7763 the PR's regex_effbot effect ranges +1.2 to +8.9%, merged _sre -4.3 to +5.7%, and main's own layouts differ by up to 4.7% on the newer runners: layout moves regex_effbot about as much as the change does. The earlier 7-14% came from single fixed layouts.
Contributor
Author
|
The CI failures are flakes. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Clang 19 and Apple clang from Xcode 16.3–26.3 merge most or all of the computed-goto dispatch jumps into one shared indirect jump (gh-158283). This adds a configure check for those compilers that passes
-mllvm -tail-dup-pred-size=1000when compilingceval.c, which restores one dispatch jump per opcode. The option controls the limit that llvm/llvm-project#78582 introduced.With
--with-ltothe option also has to reach the linker's LTO backend, where it applies to the whole LTO link rather than justceval.c. The clang driver does not forward-mllvmthere: with plain-mllvmon the link line it only warns "argument unused during compilation", and a thin-LTO build keeps 1 dispatch jump with Clang 19 (GNU ld or lld) or 94 with Apple clang 1700.6 (ld64). So it is passed as-Wl,-plugin-opt=for lld and GNU ld/gold with LLVMgold, and as-Wl,-mllvm,for ld64. Because it applies to the whole LTO link, it also un-merges the computed-goto dispatch of the regex engine (Modules/_sre/sre_lib.h):sre_ucs1_match,sre_ucs2_matchandsre_ucs4_matchgo from 7, 14 and 18 indirect jumps to 33 each with Clang 19 (thin LTO, x86-64), and from 10–12 indirect branches to 29–30 with Apple clang 1700.6 (LTO, arm64). Without any flag, GCC 16 has 18–19 and Clang 22 has 21–25.This is where the 6–9% regex_effbot slowdown reported in the issue comes in, and its size depends strongly on code layout. With each build linked in three different function orders (
-ffunction-sections -Wl,--shuffle-sections=.text*=<seed>, lld, thin LTO), regex_effbot is between 1% and 9% slower with this PR on the AMD EPYC 7763 runners (4.9% averaged over the layouts), regex_v8 and regex_dna get 3.0% and 1.1% faster, and the geometric mean of the three regex benchmarks is +0.2% (95% CI −0.1% to +0.5%). On an i9-13900K, an M4 Max and an M3 MacBook Air, all three get faster (geometric mean −7.1%, −3.7% and −5.7%). Keeping_sremerged, by making its shared dispatch block the target of an emptyasm goto, removes the regex_effbot slowdown on the EPYC 7763 runners but is slower than this PR on the i9, M4 Max and M3, and under Stabilizer (LLVM 21, fresh code layout per process) it is also slower on the EPYC 7763 (+5.4% on the three regex benchmarks). So the PR leaves_srealone. Data: shuffled layouts, Stabilizer, i9, M4 Max and M3.The check sits next to the Clang 22
-finline-max-stacksizeworkaround (gh-148284) and is written the same way. It is limited to the affected versions because Clang 18 and older reject the option.Dispatch jumps in
_PyEval_EvalFrameDefaultin the final binary, before → after:--with-lto=thin--enable-optimizations --with-lto--with-lto--with-ltoOn GitHub's macOS runners the check matches Xcode 16.3, 16.4, 26.0.1, 26.1.1, 26.2 and 26.3, and not Xcode 15.0.1–15.4, 16.0–16.2 or 26.4.1–26.6. On Linux it matches Clang 19 but not Clang 18, Clang 21 or GCC.
pyperformance, this PR vs main: 8.4% faster with Clang 19 PGO+LTO and 8.7% with Clang 19 thin LTO (FreeBSD's configuration) on Linux x86-64; 11.4% faster with Xcode 16.4 and 1.4% with Xcode 26.3 on macOS arm64. Confidence intervals and method are in the issue.
The version test excludes GCC, Clang 18 and older, and Clang 20 and newer; MSVC builds do not use configure.
This PR was prepared with the help of an AI assistant (Claude Code) and reviewed by me. Raw data and scripts are on the
clang19-dispatch-databranch of my fork.