Skip to content

Computed-goto interpreter 9–13% slower when built with Clang 19 or Xcode 16.3–16.4: dispatch jumps are merged #158283

Description

@matthiasgoergens

When CPython is built with Clang 19, or with Apple clang from Xcode 16.3 or 16.4, the computed-goto interpreter ends up with a single shared indirect jump instead of one dispatch jump per opcode. Per-opcode branch prediction is the point of computed gotos; restoring it makes pyperformance 8–9% faster on Linux and 11% faster on macOS arm64 with Xcode 16.4. Put the other way round, current builds with these compilers are 9–13% slower.

LLVM 19 limited tail duplication of blocks that end in an indirect branch (llvm/llvm-project#78582). Clang lowers all computed gotos to one shared indirectbr block and relies on tail duplication to copy it back into each predecessor; the new limit stops that. LLVM 20.1.0 fixed this partially (llvm/llvm-project#116072) and 20.1.1 fully (llvm/llvm-project#114990). No 19.x release has the fix. Xcode 26.0–26.3 has the partial fix; it still merges most dispatch jumps, but the cost is only 1.4%.

GCC builds and the tail-calling interpreter (--with-tail-call-interp) are not affected.

Nelson Elhage found in March 2025 that this LLVM 19 regression accounted for most of the originally reported speedup of the 3.14 tail-calling interpreter. His issue about merged dispatch jumps, gh-129987, was closed after gh-132295 and gh-132530, which changed GCC's SLP vectorization; Clang 19 builds still have one dispatch jump.

Measurements

Indirect jumps in _PyEval_EvalFrameDefault, in Python/ceval.o built at -O3 (x86-64) and in the final PGO+LTO binary. PGO builds are not deterministic, so the PGO+LTO counts are ranges over the independent builds:

compiler -O3 object PGO+LTO binary
GCC 12 / 13 / 14 257 234
Clang 18 269 not measured
Clang 19.1.7 1 1–9
Clang 19.1.7 + -mllvm -tail-dup-pred-size=1000 269 356–360
Clang 21 268 354–362

pyperformance on x86-64 Linux GitHub runners, main at 6af40a6. Each experiment used a randomized block design with the builds interleaved per benchmark and a same-binary control; the 95% confidence interval is a bootstrap over independent builds. Negative is faster.

comparison build builds geomean 95% CI
Clang 19 + -mllvm -tail-dup-pred-size=1000 vs Clang 19 PGO+LTO 6 −8.6% [−9.6, −7.4]
Clang 19 vs GCC 13 PGO+LTO 6 +7.7% [+6.4, +8.7]
Clang 19 + flag vs GCC 13 PGO+LTO 6 −1.6% [−1.8, −1.3]
gh-158286 vs main, Clang 19 PGO+LTO 6 −8.4% [−9.5, −6.9]
gh-158286 vs main, Clang 19 thin LTO, no PGO (FreeBSD's configuration) 8 −8.7% [−9.2, −8.3]

The same-binary controls were within ±0.2%. The biggest single-benchmark gains are 18–24% (unpack_sequence, deepcopy_memo, nbody, scimark_sor); regex_effbot gets 6–9% slower. Under LTO the option also un-merges the regex engine's own computed-goto dispatch, and the regex_effbot slowdown depends strongly on code layout: with three randomised layouts per build it is 1–9% on the EPYC 7763 runners, and the three regex benchmarks together show no significant change (data; details in gh-158286).

macOS 15 arm64 GitHub runners, PGO+LTO, gh-158286 vs main:

compiler dispatch jumps before → after builds geomean 95% CI
Xcode 16.4 (Apple clang 1700.0.13) 1 → 356–357 5 −11.4% [−13.2, −9.8]
Xcode 26.3 (Apple clang 1700.6) 110–121 → 356–358 5 −1.4% [−2.2, −0.7]

These runners are noisier; the same-binary control there was −0.4% [−1.8, +0.9].

Raw data, configurations and scripts: https://github.com/matthiasgoergens/cpython/tree/clang19-dispatch-data

Affected builds

Checked against the compiler string in shipped binaries where possible:

  • FreeBSD 14 and 15 packages, python311 to python314 (base clang 19.1.7, thin LTO)
  • OpenBSD 7.8 and 7.9, python 3.12 and 3.13 (base clang 19.1.7)
  • OpenMandriva Lx 6.0, python 3.11 (clang 19.1.7)
  • macOS builds made with Xcode 16.3 or 16.4, for example MacPorts python313 and python314 on macOS 15: one dispatch jump left
  • macOS builds made with Xcode 26.0–26.3, for example Homebrew's macOS 15 bottles: 94 of ~300 (arm64) or 62 of ~276 (x86-64) dispatch jumps left in a --with-lto build

python-build-standalone/uv (January–February 2025) and conda-forge's macOS builds also used Clang 19 for a while, but have since moved on. Xcode 26.4 and later still merge some dispatch jumps (200 in ceval.o, 328 with the option added by hand); gh-158286 leaves them alone.

Proposed fix

gh-158286 passes -mllvm -tail-dup-pred-size=1000 when compiling ceval.c with Clang 19 or Apple clang 1700.x (Xcode 16.3–26.3), and to the linker with LTO.

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux, macOS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    buildThe build process and cross-buildperformancePerformance or resource usagetype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions