When CPython is built with Clang 19, or with Apple clang from Xcode 16.3 or 16.4, the computed-goto interpreter ends up with a single shared indirect jump instead of one dispatch jump per opcode. Per-opcode branch prediction is the point of computed gotos; restoring it makes pyperformance 8–9% faster on Linux and 11% faster on macOS arm64 with Xcode 16.4. Put the other way round, current builds with these compilers are 9–13% slower.
LLVM 19 limited tail duplication of blocks that end in an indirect branch (llvm/llvm-project#78582). Clang lowers all computed gotos to one shared indirectbr block and relies on tail duplication to copy it back into each predecessor; the new limit stops that. LLVM 20.1.0 fixed this partially (llvm/llvm-project#116072) and 20.1.1 fully (llvm/llvm-project#114990). No 19.x release has the fix. Xcode 26.0–26.3 has the partial fix; it still merges most dispatch jumps, but the cost is only 1.4%.
GCC builds and the tail-calling interpreter (--with-tail-call-interp) are not affected.
Nelson Elhage found in March 2025 that this LLVM 19 regression accounted for most of the originally reported speedup of the 3.14 tail-calling interpreter. His issue about merged dispatch jumps, gh-129987, was closed after gh-132295 and gh-132530, which changed GCC's SLP vectorization; Clang 19 builds still have one dispatch jump.
Measurements
Indirect jumps in _PyEval_EvalFrameDefault, in Python/ceval.o built at -O3 (x86-64) and in the final PGO+LTO binary. PGO builds are not deterministic, so the PGO+LTO counts are ranges over the independent builds:
| compiler |
-O3 object |
PGO+LTO binary |
| GCC 12 / 13 / 14 |
257 |
234 |
| Clang 18 |
269 |
not measured |
| Clang 19.1.7 |
1 |
1–9 |
Clang 19.1.7 + -mllvm -tail-dup-pred-size=1000 |
269 |
356–360 |
| Clang 21 |
268 |
354–362 |
pyperformance on x86-64 Linux GitHub runners, main at 6af40a6. Each experiment used a randomized block design with the builds interleaved per benchmark and a same-binary control; the 95% confidence interval is a bootstrap over independent builds. Negative is faster.
| comparison |
build |
builds |
geomean |
95% CI |
Clang 19 + -mllvm -tail-dup-pred-size=1000 vs Clang 19 |
PGO+LTO |
6 |
−8.6% |
[−9.6, −7.4] |
| Clang 19 vs GCC 13 |
PGO+LTO |
6 |
+7.7% |
[+6.4, +8.7] |
| Clang 19 + flag vs GCC 13 |
PGO+LTO |
6 |
−1.6% |
[−1.8, −1.3] |
| gh-158286 vs main, Clang 19 |
PGO+LTO |
6 |
−8.4% |
[−9.5, −6.9] |
| gh-158286 vs main, Clang 19 |
thin LTO, no PGO (FreeBSD's configuration) |
8 |
−8.7% |
[−9.2, −8.3] |
The same-binary controls were within ±0.2%. The biggest single-benchmark gains are 18–24% (unpack_sequence, deepcopy_memo, nbody, scimark_sor); regex_effbot gets 6–9% slower. Under LTO the option also un-merges the regex engine's own computed-goto dispatch, and the regex_effbot slowdown depends strongly on code layout: with three randomised layouts per build it is 1–9% on the EPYC 7763 runners, and the three regex benchmarks together show no significant change (data; details in gh-158286).
macOS 15 arm64 GitHub runners, PGO+LTO, gh-158286 vs main:
| compiler |
dispatch jumps before → after |
builds |
geomean |
95% CI |
| Xcode 16.4 (Apple clang 1700.0.13) |
1 → 356–357 |
5 |
−11.4% |
[−13.2, −9.8] |
| Xcode 26.3 (Apple clang 1700.6) |
110–121 → 356–358 |
5 |
−1.4% |
[−2.2, −0.7] |
These runners are noisier; the same-binary control there was −0.4% [−1.8, +0.9].
Raw data, configurations and scripts: https://github.com/matthiasgoergens/cpython/tree/clang19-dispatch-data
Affected builds
Checked against the compiler string in shipped binaries where possible:
- FreeBSD 14 and 15 packages, python311 to python314 (base clang 19.1.7, thin LTO)
- OpenBSD 7.8 and 7.9, python 3.12 and 3.13 (base clang 19.1.7)
- OpenMandriva Lx 6.0, python 3.11 (clang 19.1.7)
- macOS builds made with Xcode 16.3 or 16.4, for example MacPorts python313 and python314 on macOS 15: one dispatch jump left
- macOS builds made with Xcode 26.0–26.3, for example Homebrew's macOS 15 bottles: 94 of ~300 (arm64) or 62 of ~276 (x86-64) dispatch jumps left in a
--with-lto build
python-build-standalone/uv (January–February 2025) and conda-forge's macOS builds also used Clang 19 for a while, but have since moved on. Xcode 26.4 and later still merge some dispatch jumps (200 in ceval.o, 328 with the option added by hand); gh-158286 leaves them alone.
Proposed fix
gh-158286 passes -mllvm -tail-dup-pred-size=1000 when compiling ceval.c with Clang 19 or Apple clang 1700.x (Xcode 16.3–26.3), and to the linker with LTO.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux, macOS
When CPython is built with Clang 19, or with Apple clang from Xcode 16.3 or 16.4, the computed-goto interpreter ends up with a single shared indirect jump instead of one dispatch jump per opcode. Per-opcode branch prediction is the point of computed gotos; restoring it makes pyperformance 8–9% faster on Linux and 11% faster on macOS arm64 with Xcode 16.4. Put the other way round, current builds with these compilers are 9–13% slower.
LLVM 19 limited tail duplication of blocks that end in an indirect branch (llvm/llvm-project#78582). Clang lowers all computed gotos to one shared
indirectbrblock and relies on tail duplication to copy it back into each predecessor; the new limit stops that. LLVM 20.1.0 fixed this partially (llvm/llvm-project#116072) and 20.1.1 fully (llvm/llvm-project#114990). No 19.x release has the fix. Xcode 26.0–26.3 has the partial fix; it still merges most dispatch jumps, but the cost is only 1.4%.GCC builds and the tail-calling interpreter (
--with-tail-call-interp) are not affected.Nelson Elhage found in March 2025 that this LLVM 19 regression accounted for most of the originally reported speedup of the 3.14 tail-calling interpreter. His issue about merged dispatch jumps, gh-129987, was closed after gh-132295 and gh-132530, which changed GCC's SLP vectorization; Clang 19 builds still have one dispatch jump.
Measurements
Indirect jumps in
_PyEval_EvalFrameDefault, inPython/ceval.obuilt at-O3(x86-64) and in the final PGO+LTO binary. PGO builds are not deterministic, so the PGO+LTO counts are ranges over the independent builds:-O3object-mllvm -tail-dup-pred-size=1000pyperformance on x86-64 Linux GitHub runners, main at 6af40a6. Each experiment used a randomized block design with the builds interleaved per benchmark and a same-binary control; the 95% confidence interval is a bootstrap over independent builds. Negative is faster.
-mllvm -tail-dup-pred-size=1000vs Clang 19The same-binary controls were within ±0.2%. The biggest single-benchmark gains are 18–24% (unpack_sequence, deepcopy_memo, nbody, scimark_sor); regex_effbot gets 6–9% slower. Under LTO the option also un-merges the regex engine's own computed-goto dispatch, and the regex_effbot slowdown depends strongly on code layout: with three randomised layouts per build it is 1–9% on the EPYC 7763 runners, and the three regex benchmarks together show no significant change (data; details in gh-158286).
macOS 15 arm64 GitHub runners, PGO+LTO, gh-158286 vs main:
These runners are noisier; the same-binary control there was −0.4% [−1.8, +0.9].
Raw data, configurations and scripts: https://github.com/matthiasgoergens/cpython/tree/clang19-dispatch-data
Affected builds
Checked against the compiler string in shipped binaries where possible:
--with-ltobuildpython-build-standalone/uv (January–February 2025) and conda-forge's macOS builds also used Clang 19 for a while, but have since moved on. Xcode 26.4 and later still merge some dispatch jumps (200 in
ceval.o, 328 with the option added by hand); gh-158286 leaves them alone.Proposed fix
gh-158286 passes
-mllvm -tail-dup-pred-size=1000when compilingceval.cwith Clang 19 or Apple clang 1700.x (Xcode 16.3–26.3), and to the linker with LTO.CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux, macOS