Skip to content

gh-130706: Add a fast path to _Py_Dealloc() for non-GC objects. - #158645

Open
corona10 wants to merge 2 commits into
python:mainfrom
corona10:gh-130706
Open

corona10 wants to merge 2 commits into
python:mainfrom
corona10:gh-130706

Conversation

@corona10

@corona10 corona10 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

@corona10

corona10 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member Author

@mpage @ZeroIntensity

Benchmark

Benchmark baseline opt
dealloc_only_float 5.66 ns 4.69 ns: 1.21x faster
dealloc_only_bigint 4.90 ns 4.48 ns: 1.09x faster
dealloc_only_str 5.94 ns 4.71 ns: 1.26x faster
dealloc_only_bytes 5.09 ns 3.89 ns: 1.31x faster
alloc_dealloc_float 14.1 ns 12.1 ns: 1.17x faster
alloc_dealloc_bigint 20.5 ns 18.2 ns: 1.13x faster
alloc_dealloc_str 53.9 ns 49.9 ns: 1.08x faster
alloc_dealloc_bytes 48.3 ns 46.9 ns: 1.03x faster
Geometric mean (ref) 1.16x faster

Script

import pyperf

INNER = 100_000
RANGE = range(INNER)


def _dealloc_only(loops, make):
    total = 0.0
    for _ in range(loops):
        objs = [make(i) for i in RANGE]
        t0 = pyperf.perf_counter()
        objs.clear()
        total += pyperf.perf_counter() - t0
    return total


def dealloc_only_float(loops):
    return _dealloc_only(loops, lambda i: i * 1.5)


def dealloc_only_bigint(loops):
    big = 2 ** 70
    return _dealloc_only(loops, lambda i: i + big)


def dealloc_only_str(loops):
    return _dealloc_only(loops, str)


def dealloc_only_bytes(loops):
    return _dealloc_only(loops, lambda i: bytes(3))


def alloc_dealloc_float(loops):
    t0 = pyperf.perf_counter()
    for _ in range(loops):
        for i in RANGE:
            x = i * 1.5
    return pyperf.perf_counter() - t0


def alloc_dealloc_bigint(loops):
    big = 2 ** 70
    t0 = pyperf.perf_counter()
    for _ in range(loops):
        for i in RANGE:
            x = i + big
    return pyperf.perf_counter() - t0


def alloc_dealloc_str(loops):
    t0 = pyperf.perf_counter()
    for _ in range(loops):
        for i in RANGE:
            x = str(i)
    return pyperf.perf_counter() - t0


def alloc_dealloc_bytes(loops):
    t0 = pyperf.perf_counter()
    for _ in range(loops):
        for i in RANGE:
            x = bytes(3)
    return pyperf.perf_counter() - t0


runner = pyperf.Runner()
for func in (dealloc_only_float, dealloc_only_bigint, dealloc_only_str,
             dealloc_only_bytes, alloc_dealloc_float, alloc_dealloc_bigint,
             alloc_dealloc_str, alloc_dealloc_bytes):
    runner.bench_time_func(func.__name__, func, inner_loops=INNER)

@corona10 corona10 added the performance Performance or resource usage label Oct 3, 2026
@corona10
corona10 requested a review from markshannon October 3, 2026 14:55
@corona10

corona10 commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

cc @markshannon @vstinner who recently touched this function.

@vstinner vstinner left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is Py_NO_INLINE really useful here? Why not just a fast-path at the _Py_Dealloc() entry point? Usually, I prefer to let the compiler decides how to inline or not.

Something like that:

diff --git a/Objects/object.c b/Objects/object.c
index c7aeba0cee2..b30332f5945 100644
--- a/Objects/object.c
+++ b/Objects/object.c
@@ -3312,6 +3312,13 @@ _Py_Dealloc(PyObject *op)
     PyTypeObject *type = Py_TYPE(op);
     unsigned long gc_flag = type->tp_flags & Py_TPFLAGS_HAVE_GC;
     destructor dealloc = type->tp_dealloc;
+#if !defined(Py_DEBUG) && !defined(Py_TRACE_REFS)
+    if (!gc_flag && _PyRuntime.ref_tracer.tracer_func == NULL) {
+        dealloc(op);
+        return;
+    }
+#endif
+
     PyThreadState *tstate = NULL;
     intptr_t margin = 0;
     if (gc_flag) {

@corona10

corona10 commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

Is Py_NO_INLINE really useful here? Why not just a fast-path at the _Py_Dealloc() entry point? Usually, I prefer to let the compiler decides how to inline or not.

Let me check with the benchmark.

@vstinner

vstinner commented Oct 3, 2026

Copy link
Copy Markdown
Member

There is also _Py_DECREF_SPECIALIZED() which skips completely Py_Dealloc(). For example, it's used by _Py_DECREF_INT() in longobject.c.

_Py_DECREF_SPECIALIZED() still calls _PyReftracerTrack(op, PyRefTracer_DESTROY);.

@corona10

corona10 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member Author

Is Py_NO_INLINE really useful here? Why not just a fast-path at the _Py_Dealloc() entry point? Usually, I prefer to let the compiler decides how to inline or not.

The suggested version is also faster, but with the GC path in the same function, Clang still saves the callee-saved registers even on the non-GC path, which never uses them. The current version avoids this by moving the GC path into a non-inlined function, so _Py_Dealloc doesn't need to save those registers (compiler magic...).
The gain is very small in my benchmarks, so I'm fine with either version if you prefer the simpler one.

x86-64, suggested version:

pushq   %rbp
movq    %rsp, %rbp
pushq   %r15
pushq   %r14
pushq   %r13
pushq   %r12
pushq   %rbx
pushq   %rax
...                             # flags / tracer check
addq    $8, %rsp
popq    %rbx
popq    %r12
popq    %r13
popq    %r14
popq    %r15
popq    %rbp
jmpq    *%rsi                   # tp_dealloc

x86-64, this PR:

pushq   %rbp
movq    %rsp, %rbp
...                             # flags / tracer check
popq    %rbp
jmpq    *64(%rax)               # tp_dealloc

AArch64, suggested version:

sub     sp, sp, #0x50
stp     x24, x23, [sp, #0x10]
stp     x22, x21, [sp, #0x20]
stp     x20, x19, [sp, #0x30]
stp     x29, x30, [sp, #0x40]
add     x29, sp, #0x40
...                             // flags / tracer check
ldp     x29, x30, [sp, #0x40]
ldp     x20, x19, [sp, #0x30]
ldp     x22, x21, [sp, #0x20]
ldp     x24, x23, [sp, #0x10]
add     sp, sp, #0x50
br      x1                      // tp_dealloc

AArch64, this PR:

stp     x29, x30, [sp, #-0x10]!
mov     x29, sp
...                             // flags / tracer check
ldp     x29, x30, [sp], #0x10
br      x1                      // tp_dealloc
Benchmark baseline this PR early return
dealloc_only_float 5.66 4.69 (1.21x) 4.82 (1.17x)
dealloc_only_bigint 4.90 4.48 (1.09x) 4.14 (1.18x)
dealloc_only_str 5.94 4.71 (1.26x) 4.89 (1.22x)
dealloc_only_bytes 5.09 3.89 (1.31x) 4.47 (1.14x)
Geometric mean 1.16x 1.12x

@corona10

corona10 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member Author

There is also _Py_DECREF_SPECIALIZED() which skips completely Py_Dealloc(). For example, it's used by _Py_DECREF_INT() in longobject.c.

_Py_DECREF_SPECIALIZED() still calls _PyReftracerTrack(op, PyRefTracer_DESTROY);.

I will handle this in a separate PR :)

@corona10

corona10 commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

The gain is very small in my benchmarks, so I'm fine with either version if you prefer the simpler one.

Okay, I 've verified this at the PGO + LTO build, and the result is the same.
When @colesbury piles this issue, he prefers the following assembly:

0x00000000001a7edb <+27>:    jmp    QWORD PTR [rax+0x40]  # The fast path ends here with the jump to `tp_dealloc()`

I would like to keep the current version.

Comment thread Objects/object.c
stack is shallower */
void
_Py_Dealloc(PyObject *op)
static Py_NO_INLINE void

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please explain in the comment the rationale for Py_NO_INLINE.

Comment thread Objects/object.c
}
}

void

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Or you may explain here the dealloc_general() split with Py_NO_INLINE.

Comment thread Objects/object.c
void
_Py_Dealloc(PyObject *op)
static Py_NO_INLINE void
dealloc_general(PyObject *op)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would prefer to rename the function to "py_dealloc()".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting core review performance Performance or resource usage type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants