Skip to content

gh-158283: Force dispatch tail duplication with Clang 19 and Apple clang 17 - #158286

Open
matthiasgoergens wants to merge 1 commit into
python:mainfrom
matthiasgoergens:pr/clang19-dispatch
Open

matthiasgoergens wants to merge 1 commit into
python:mainfrom
matthiasgoergens:pr/clang19-dispatch

Conversation

@matthiasgoergens

@matthiasgoergens matthiasgoergens commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Clang 19 and Apple clang from Xcode 16.3–26.3 merge most or all of the computed-goto dispatch jumps into one shared indirect jump (gh-158283). This adds a configure check for those compilers that passes -mllvm -tail-dup-pred-size=1000 when compiling ceval.c, which restores one dispatch jump per opcode. The option controls the limit that llvm/llvm-project#78582 introduced.

With --with-lto the option also has to reach the linker's LTO backend, where it applies to the whole LTO link rather than just ceval.c. The clang driver does not forward -mllvm there: with plain -mllvm on the link line it only warns "argument unused during compilation", and a thin-LTO build keeps 1 dispatch jump with Clang 19 (GNU ld or lld) or 94 with Apple clang 1700.6 (ld64). So it is passed as -Wl,-plugin-opt= for lld and GNU ld/gold with LLVMgold, and as -Wl,-mllvm, for ld64. Because it applies to the whole LTO link, it also un-merges the computed-goto dispatch of the regex engine (Modules/_sre/sre_lib.h): sre_ucs1_match, sre_ucs2_match and sre_ucs4_match go from 7, 14 and 18 indirect jumps to 33 each with Clang 19 (thin LTO, x86-64), and from 10–12 indirect branches to 29–30 with Apple clang 1700.6 (LTO, arm64). Without any flag, GCC 16 has 18–19 and Clang 22 has 21–25.

This is where the 6–9% regex_effbot slowdown reported in the issue comes in, and its size depends strongly on code layout. With each build linked in three different function orders (-ffunction-sections -Wl,--shuffle-sections=.text*=<seed>, lld, thin LTO), regex_effbot is between 1% and 9% slower with this PR on the AMD EPYC 7763 runners (4.9% averaged over the layouts), regex_v8 and regex_dna get 3.0% and 1.1% faster, and the geometric mean of the three regex benchmarks is +0.2% (95% CI −0.1% to +0.5%). On an i9-13900K, an M4 Max and an M3 MacBook Air, all three get faster (geometric mean −7.1%, −3.7% and −5.7%). Keeping _sre merged, by making its shared dispatch block the target of an empty asm goto, removes the regex_effbot slowdown on the EPYC 7763 runners but is slower than this PR on the i9, M4 Max and M3, and under Stabilizer (LLVM 21, fresh code layout per process) it is also slower on the EPYC 7763 (+5.4% on the three regex benchmarks). So the PR leaves _sre alone. Data: shuffled layouts, Stabilizer, i9, M4 Max and M3.

The check sits next to the Clang 22 -finline-max-stacksize workaround (gh-148284) and is written the same way. It is limited to the affected versions because Clang 18 and older reject the option.

Dispatch jumps in _PyEval_EvalFrameDefault in the final binary, before → after:

compiler build x86-64 arm64
Clang 19.1.7 --with-lto=thin 1 → 276
Clang 19.1.7 --enable-optimizations --with-lto 1–7 → 359–365
Xcode 16.4 (Apple clang 1700.0.13) --with-lto 1 → 276 1 → 300
Xcode 26.3 (Apple clang 1700.6) --with-lto 62 → 276 94 → 300

On GitHub's macOS runners the check matches Xcode 16.3, 16.4, 26.0.1, 26.1.1, 26.2 and 26.3, and not Xcode 15.0.1–15.4, 16.0–16.2 or 26.4.1–26.6. On Linux it matches Clang 19 but not Clang 18, Clang 21 or GCC.

pyperformance, this PR vs main: 8.4% faster with Clang 19 PGO+LTO and 8.7% with Clang 19 thin LTO (FreeBSD's configuration) on Linux x86-64; 11.4% faster with Xcode 16.4 and 1.4% with Xcode 26.3 on macOS arm64. Confidence intervals and method are in the issue.

The version test excludes GCC, Clang 18 and older, and Clang 20 and newer; MSVC builds do not use configure.

This PR was prepared with the help of an AI assistant (Claude Code) and reviewed by me. Raw data and scripts are on the clang19-dispatch-data branch of my fork.

…ple clang 17

LLVM 19 limits tail duplication of blocks ending in an indirect branch
(llvm/llvm-project#78582), so the computed-goto interpreter is compiled with
a single shared dispatch jump instead of one per instruction.  Apple clang
from Xcode 16.3-16.4 has the same bug; Xcode 26.0-26.3 merges most of the
dispatch jumps.  LLVM 20.1.1 fixed it (llvm/llvm-project#114990).

Detect the affected compilers and pass -mllvm -tail-dup-pred-size=1000 when
compiling ceval.c, and to the linker's LTO backend under --with-lto.  With
Clang 19 this made pyperformance 8.4-8.7% faster (PGO+LTO, and thin LTO
without PGO).
matthiasgoergens pushed a commit to matthiasgoergens/cpython that referenced this pull request Sep 27, 2026
matthiasgoergens added a commit to matthiasgoergens/cpython that referenced this pull request Sep 28, 2026
GitHub's ubuntu-24.04 runners come with different CPUs (EPYC 7763, 9V74,
9V45, Xeon Platinum 8573C, Xeon 6973P-C were all seen), and the regex_effbot
slowdown of pythongh-158286 is largest on EPYC 7763.  blockbench now records the
CPU model in every row, and `analyze --cpu REGEX` restricts to matching
blocks.  backfill_cpu.py recovers the CPU of older runs from job logs while
GitHub still has them; all-cpu.jsonl for exp3, exp9b and exp11b were made
with it.

build_arms.py accepts "makefile_sed" to edit the generated Makefile (and
fails if an expression changes nothing).  exp12-sre uses it to compile
Modules/_sre/sre.c with -fno-lto, to test whether keeping the regex
engine's dispatch out of the link-time option avoids the regression.

make_branch.sh now pushes to $PERF_CI_REMOTE without --force: in a clone
where origin is python/cpython, the old `git push -f origin` would have
targeted upstream.

Also drop a stray .pyc and ignore __pycache__.
matthiasgoergens added a commit to matthiasgoergens/cpython that referenced this pull request Sep 28, 2026
…ostly layout

Clang 19 main / PR / PR + asm goto _sre, each linked with lld
--shuffle-sections under 3 seeds, 11 jobs (33 blocks).  Pooled over
seeds, regex_effbot is +4.9% with the PR on EPYC 7763 and the regex
geomean +0.2% (n.s.).  Per seed on EPYC 7763 the PR's regex_effbot
effect ranges +1.2 to +8.9%, merged _sre -4.3 to +5.7%, and main's own
layouts differ by up to 4.7% on the newer runners: layout moves
regex_effbot about as much as the change does.  The earlier 7-14% came
from single fixed layouts.
@matthiasgoergens

Copy link
Copy Markdown
Contributor Author

The CI failures are flakes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant