From 77b9481c4cca4337023e6fdcec40a067151faceb Mon Sep 17 00:00:00 2001 From: Mark Shannon Date: Thu, 24 Sep 2026 15:44:36 +0100 Subject: [PATCH] PEP 805: various clarifications and elaborations. Add 'unprotect' --- peps/pep-0805.rst | 228 ++++++++++++---------- peps/pep-0805/appendix-examples.rst | 57 ++++++ peps/pep-0805/appendix-implementation.rst | 165 ++++++++++++++-- 3 files changed, 334 insertions(+), 116 deletions(-) diff --git a/peps/pep-0805.rst b/peps/pep-0805.rst index 3f183a66d86..775c54c4dbd 100644 --- a/peps/pep-0805.rst +++ b/peps/pep-0805.rst @@ -70,18 +70,15 @@ One CPython, not two CPython is currently split into two: the default build and the free-threading build. -Proponents of free-threading expect that free-threading will become the only +Proponents of free-threading expect that free-threading will become the default version of CPython in a few years. The authors feel that this will be very -challenging to achieve, and may be impossible. Removing the default build would -involve breaking vast numbers of applications and libraries that are not safe -to use with a free-threading build. Even though many libraries are marked as -supporting free-threading, it is unlikely that they are all completely safe -to use in a free-threading environment given the difficulty of eliminating -race conditions. +challenging to achieve, and may be impossible. Making the free-threading build +the default would break the very large number of applications and libraries +that are not safe to use with a free-threading build. The authors fear that without this PEP, or something like it, we will be stuck with two builds of Python forever: Users of free-threading will be unwilling to -give up parallelism, and users of the default build will be unable to risk +give up parallelism, and users of the with-GIL build will be unable to risk using the free-threading build. .. note:: @@ -152,7 +149,7 @@ safe to access that object from the current thread of execution. The core concept is that it is access to objects, rather than operations on those objects, that is controlled. If an object cannot be accessed by a thread, -then that thread cannot perform any unsafe operation on that object, since +then that thread cannot perform an unsafe operation on that object, since it cannot perform *any* operation on it. The motivation for this is both correctness and performance. Protecting @@ -286,6 +283,12 @@ sense that the object itself will not be corrupted, they are not generally thread safe. *Immutable* or *local* collections should be used where possible. +Sharability of objects and their classes +---------------------------------------- + +Making an object shareable, either by freezing or synchronization, will fail +if its class is not sharable (it is neither frozen nor synchronized). + Object dictionaries ------------------- @@ -294,6 +297,7 @@ Freezing an object will convert its ``__dict__`` into a ``frozendict``. Synchronizing a module (or any object that both supports synchronization and has a ``__dict__``) will convert the ``__dict__`` into a ``SynchronizedDict``. +Likewise, freezing an object's ``__dict__`` will freeze the object as well. .. _pep805-ThreadGroup: @@ -385,6 +389,12 @@ It is an error to call ``acquire`` or ``release`` on a *protective* lock. Such a lock can only get acquired by using a ``with`` statement with that lock, or a compound lock formed from it, as the context manager. +Unprotecting objects +'''''''''''''''''''' + +Objects can be unprotected, converting them back to local objects of the +current ``ThreadGroup``, by calling ``unprotect()``. + .. _pep805-new-api: New API @@ -398,13 +408,12 @@ This PEP proposes adding the following: * A builtin ``freeze(obj)`` function, which calls ``obj.__freeze__()`` * A ``protect(obj)`` method, added to ``Lock`` and ``RLock``, which returns a *protected* copy of ``obj``. +* A builtin ``unprotect(obj)`` function, to make *protected* objects *local* * The ``SynchronizedList``, ``SynchronizedDict`` and ``SynchronizedSet`` classes * A ``synchronize()`` method, added to ``list``, ``set`` and ``dict``, which returns the *synchronized* version of that object and clears the original - object. + object * A ``__shareable__`` read-only attribute for all objects -* The ``Channel`` and ``TransferBox`` classes for passing objects from one - ``ThreadGroup`` to another * The ``ThreadGroup`` class * The ``group`` parameter used when creating ``Thread``\s now has meaning and can be set to a ``ThreadGroup`` @@ -433,6 +442,15 @@ will gain a ``__freeze__()`` method, converting the object into a Note that freezing an object is a shallow operation; ``x.__freeze__()`` only freezes ``x`` and not any of the objects that ``x`` refers to. +The ``freeze`` function can be used as a decorator to freeze classes:: + + @freeze + class C: + """This class cannot be modified once constructed. + Instances of this class can still be mutated unless + explicitly frozen + """ + Freezing an object also freezes its dictionary: .. code-block:: pycon @@ -448,6 +466,7 @@ other developers. It is therefore recommended that freezing is done in a principled fashion, typically freezing all instances of a class, or none. For example:: + @freeze class ImmutablePoint: def __init__(self, x, y): @@ -482,18 +501,9 @@ finished:: merely a convention, it will be enforced by the VM. Once an object is frozen it cannot be unfrozen. -A ``__deep_freeze__`` method may be added as a +A ``deep_freeze`` function may be added as a :ref:`future enhancement`. -The ``freeze`` function can be used as a decorator to freeze classes:: - - @freeze - class C: - """This class cannot be modified once constructed. - Instances of this class can still be mutated unless - explicitly frozen - """ - Synchronization ''''''''''''''' @@ -501,66 +511,40 @@ Synchronization The ``synchronized`` state protects the internal state of an object, but is only available for some builtin and extension objects. -Passing mutable values between parallel threads -''''''''''''''''''''''''''''''''''''''''''''''' +Passing mutable objects between parallel threads +'''''''''''''''''''''''''''''''''''''''''''''''' -Two classes are provided to pass *local* objects between ThreadGroups. +Mutable objects can be passed between parallel threads of execution by +*protecting* them in one thread, storing a reference to them in a shareable +object, then extracting and unprotecting them in the second thread. -The ``TransferBox`` class provides a *synchronized* container -for moving objects from one ThreadGroup to another. +Using this protect then unprotect idiom, various helpers can be built to move +mutable objects from one ``ThreadGroup`` to another. -When creating a ``TransferBox`` from a *local* object, the object is -copied before boxing. The new *local* object is not attached to any -ThreadGroup. +For example, it is possible to build a ``Channel`` class for passing a +stream of objects from one ``ThreadGroup`` to another:: -When claiming the object from the box, the current ThreadGroup becomes -the owner of the object, if the box's ``sink`` is ``None`` or the current -ThreadGroup. + class Channel: -*Immutable*, *protected* and *synchronized* objects are passed uncopied:: + def put(self, obj): + "Put mutable object into channel" - EMPTY = sentinel('EMPTY') + def get(self): + "Get mutable object from channel" - class TransferBox[T]: - def __new__(cls, obj: T, sink: ThreadGroup | None=None): - self.sink = sink - self._obj = copy(obj) if obj.__state__ == LOCAL else obj +Or a ``TransferBox`` class for passing single mutable objects:: - def claim(self) -> T: - if self._obj is EMPTY: - raise ValueError(...) - if self.sink is not None and self.sink != current_ThreadGroup: - raise ValueError(...) - result = self._obj - self._obj = EMPTY - return result - - -The ``Channel`` class provides a higher level API for passing objects from one -ThreadGroup to another. Channel is equivalent to this Python class:: - - class Channel: - - def __init__(self): - self.mutex = Lock() - with self.mutex: - self.queue = self.mutex.protect(deque()) - self.__freeze__() - - def put(self, obj): - box = TransferBox(del obj) - with self.mutex: - self.queue.append(box) + class TransferBox[T]: - def get(self): - with self.mutex: - return self.queue.popleft().claim() + def __new__(cls, obj: T): + "Create a new box containing (a copy of) obj" + def claim(self) -> T: + "Claim the contents of the box" -Adding a "deep" ``put`` method might be added as a -:ref:`future enhancement`, if there is -sufficient demand for it. +See :ref:`pep805-examples-channel` and :ref:`pep805-examples-transfer-box` +for the full implementations. .. _pep805-GIL: @@ -607,6 +591,8 @@ Allowed operations +------------------------+-----------+-----------------+-----------------+---------------+----------------+ | ``protect()`` | No | Yes\ :sup:`2,3` | N/A | No | No | +------------------------+-----------+-----------------+-----------------+---------------+----------------+ +| ``unprotect()`` | No | No | N/A | Yes | No | ++------------------------+-----------+-----------------+-----------------+---------------+----------------+ | ``synchronize()`` | No | Yes\ :sup:`2` | N/A | No | No | +------------------------+-----------+-----------------+-----------------+---------------+----------------+ | All other operations | Yes | Yes | N/A | Yes | Yes | @@ -766,13 +752,40 @@ Backwards Compatibility Default build ------------- -Compared to the default build, the only incompatible change is that the -lifetimes of some objects (those of +Python code +''''''''''' + +Python code that currently runs on the default build will, with one obscure +exception, continue to run and run correctly. + +The only breaking change is that the ``__kwdefaults__`` of functions will +becomes a frozen dicts. Code that modifies a function’s ``__kwdefaults__`` will +now raise an exception. (The authors hope that no one writes code +that mutates the keyword defaults of functions, but it might happen) + +Compared to the default build, the lifetimes of some objects (those of :ref:`primitive types`) may be extended, possibly increasing memory use. -Free-threading build --------------------- +C extensions +'''''''''''' + +Due to the ABI breakage, C extensions will need to be recompiled. +Note, there is no API change, so no code changes are needed to support the +new ABI. + +C extensions that interact with CPython only through function calls will just +work. +Some care will be needed if calling functions provided by other extensions, +either through explicit APIs provided by those extensions or via function +pointers on ``PyTypeObject``\s from other extensions. In those cases explicit +calls to access control functions will need to be added. + +Porting code from free-threading +-------------------------------- + +Python code +''''''''''' The most obvious change is that sharing of mutable objects will raise an ``IllegalThreadAccessException`` instead of allowing data races. @@ -797,8 +810,8 @@ to it (e.g. by popping it from a shared list). Therefore, care must be exercised when transitioning dicts or lists into the synchronized state. To have threads running in parallel, without needing to explicitly set the -``ThreadGroup`` for each new thread, the environment variable -``PYTHON_PARALLEL`` should be set to 1. +``group`` for each new thread, the environment variable ``PYTHON_PARALLEL`` +should be set to 1. Safety ====== @@ -824,8 +837,8 @@ Take the example of indexing into a list: ``l[x]`` With the GIL, this can be done by first checking that ``l`` is a list, ``x`` is an int, and that ``x`` is in-bounds. Then the value can be read out of the list's array directly. However, in the free-threading build this approach -doesn't work as another thread may have mutated the list at the same time as it -was being indexed, meaning that additional synchronization is required. +doesn't work as another thread may have mutated the list at the same time as +it was being indexed, meaning that additional synchronization is required. The additional synchronization impairs performance but does not provide any useful protection against race conditions at the application level. @@ -838,9 +851,13 @@ However, additional checks will still be needed. Whenever a reference owned by a thread is created, then a check will be needed that it is legal. Since it is necessary to check that an object is *local* to the ThreadGroup, or that it is *immutable*, or that it is *synchronized* -or that it is *protected* and the correct lock is held, these checks could -be relatively expensive. However, the specializing adaptive interpreter or JIT -can specialize or eliminate these operations. +or that it is *protected* and the correct lock is held. A naive implementation +of these checks would be too expensive. + +Fortunately, the cost of the checks can be made as low as a few machine +instructions by combining them with reference counting operations, exploiting +common ownership between container objects and the objects they hold, and +specializations in the specializing adaptive interpreter and JIT. The general check:: @@ -855,8 +872,11 @@ The general check:: else: raise ... # Bad -is expensive, but by specializing for the expected case, the check can be made -cheap. +is expensive, but by embedding the state into the internal field for +``__owner__``, combining tests necessary for reference counting, and using +specialization to make sure that the most common case is tested first, the +check can be made very cheap. + For example, if we expect a *local* object, we can do a much cheaper check:: if obj.__owner__ == current_threadgroup_id: @@ -864,15 +884,15 @@ For example, if we expect a *local* object, we can do a much cheaper check:: else: do_general_check(obj) -Provided we make sure that ThreadGroup IDs and lock IDs are distinct. - +See :ref:`implementation appendix ` for more +details of how access checks can be optimized. The impact of parallelism on performance ---------------------------------------- If all threads belong to a single ``ThreadGroup`` then the JIT can eliminate checks for *local* objects (as these checks will always pass), -resulting in performance very close to the current with-gil build. +resulting in performance very close to the current with-GIL build. Depending on the amount of locking required, the performance impact of adding parallelism could range from close to zero, where only immutable objects are @@ -1003,24 +1023,21 @@ of locks, such as reader-writer locks. However, ensuring their correctness and maintaining the VM in a valid state is complex, so this is left for a future enhancement. -Deep freezing and deep transfers --------------------------------- +Deep freezing and deep protection +--------------------------------- Freezing a single object could leave a frozen object with references to -mutable objects, and transferring of single objects could leave an object local +mutable objects, and protecting a single objects could leave an object local to one thread, while other objects that it refers to are local to a different thread. Either of these scenarios are likely to lead to runtime errors. -To avoid that problem we need "deep" freezing. +To avoid that problem we need "deep" freezing and protection. Deep freezing an object would freeze that object and the transitive closure of -other mutable objects referred to by that object. Deep transferring an object -would transfer that object and the transitive closure of other local objects +other mutable objects referred to by that object. Deep protecting an object +would protect that object and the transitive closure of other local objects referred to by that object, but would raise an exception if one of those objects belonged to a different thread. -Similar to freezing, a "deep" put mechanism could be added to ``Channel``\ s -to move a whole graph of objects from one thread to another. - See also PEP 795, which proposes a deep freezing mechanism, although it is referred to as just "freezing" in that PEP. @@ -1038,24 +1055,25 @@ Open Issues Make ``del`` an expression -------------------------- -The functions ``protect``, ``Channel.put`` and creating a ``TransferBox`` -create a copy of the object passed as an argument. +The ``protect`` function, used to create thread-safe containers, +creates a copy of the object passed as an argument. + +By making ``del`` an expression, it can made clear that the +current thread is done with the object. Instead of writing:: -By making ``del`` an expression, it can be made clearer that the -current thread has done with the object. + y = mutex.protect(x) -Using ``del x`` as the argument clears ``x`` making it clear that the -current thread has done with the object. For example:: +using ``del`` makes the intent much clearer:: - channel.put(del x) + y = mutex.protect(del x) Doing this will also boost performance, as the copy can be avoided if the VM can determine, either by static analysis or reference counting, that the reference passed is unique. -The current way to do this is rather clunky:: +The current way to clear the variable is rather clunky:: - channel.put((x, x:=None)[0]) + y = mutex.protect((x, x:=None)[0]) Case of names for ``SynchronizedList``, etc. -------------------------------------------- @@ -1072,6 +1090,8 @@ parallelism, that could be added. But overwhelming developers with new additions to the standard library is not desirable. It is not clear yet which, if any, of these classes should be added: +* The ``Channel`` and ``TransferBox`` classes described above. Implementing them + in C allows the use of lock-free algorithms for better performance. * ``frozenlist`` * `AtomicRef `__ * ``SynchronizedProxy``, to proxy a *local* object (making it *protected*) diff --git a/peps/pep-0805/appendix-examples.rst b/peps/pep-0805/appendix-examples.rst index 20646b53ace..a9ec89b79b7 100644 --- a/peps/pep-0805/appendix-examples.rst +++ b/peps/pep-0805/appendix-examples.rst @@ -154,3 +154,60 @@ implemented:: for d in data: self._file.write(d) self._file.write(b"\n") + + +.. _pep805-examples-channel: + +Channel +------- + +A channel for passing mutable objects from one ``ThreadGroup`` to another:: + + class Channel: + + def __init__(self): + self.mutex = Lock() + with self.mutex: + self.queue = self.mutex.protect(deque()) + self.__freeze__() + + def put(self, obj): + with self.mutex: + if obj.__state__ == LOCAL: + obj = protect(del obj) + self.queue.append(obj) + + def get(self): + with self.mutex: + obj = self.queue.popleft() + if protected.__state__ == PROTECTED: + obj = unprotect(obj) + return obj + +.. _pep805-examples-transfer-box: + +TransferBox +----------- + +A box for passing a single mutable object from one ``ThreadGroup`` to another:: + + + class TransferBox[T]: + + def __new__(cls, obj: T): + self.mutex = Lock() + with self.mutex: + if obj.__state__ == LOCAL: + obj = protect(del obj) + self._obj = protect([obj]) + self.__freeze__() + + def claim(self) -> T: + with self.mutex: + obj = self._obj[0] + if obj is EMPTY: + raise ValueError(...) + self._obj[0] = EMPTY + if obj.__state__ == PROTECTED: + obj = unprotect(obj) + return obj diff --git a/peps/pep-0805/appendix-implementation.rst b/peps/pep-0805/appendix-implementation.rst index 3a1244bc06e..3c809cf5b8a 100644 --- a/peps/pep-0805/appendix-implementation.rst +++ b/peps/pep-0805/appendix-implementation.rst @@ -32,9 +32,9 @@ A possible object header: Reference counting ------------------ -The author expects that the biased reference counting mechanism from :pep:`703` -will be used. Like :pep:`703`, per-thread reference counting and deferred -reference counting will also be used where necessary to minimize contention. +Objects will have two reference counts, a local reference count and a shared +reference count. This is similar to the the implementation of biased +reference counting from :pep:`703`. Checking object states ---------------------- @@ -62,19 +62,23 @@ we can just check that the object is immutable. The JIT compiler can potentially remove redundant checks on the same object. +Some information about the object state can be folded into the ``owner_id`` +field. By making all thread group IDs positive and all lock IDs negative, the +state can be determined from the ``owner_id``. + Access control function ''''''''''''''''''''''' It is assumed that *local* objects will be the most likely, so if the thread state is available, that will be checked first:: - PyObject *PyObject_CheckAccessThread(PyObject *op, PyThread t) + PyObject *PyObject_CheckAccessThread(PyObject *op, PyThreadState *tstate) { - PyThreadState *tstate = PyThreadStateFromThread(t); if (op->owner_id == tstate->threadgroup_id) { return op; } - if (op->state >= SYNCHRONIZED) { + if (op->owner_id == 0) { + // Immutable or synchronized return op; } // Check for protected and stop the world cases... @@ -85,7 +89,8 @@ will be checked first:: PyObject *PyObject_CheckAccess(PyObject *op) { - if (op->state >= SYNCHRONIZED) { + if (op->owner_id == 0) { + // Immutable or synchronized return op; } PyThreadState *tstate = PyThreadState_GET(); @@ -100,6 +105,128 @@ held, as it is too easy to deadlock, so the set of held mutexes will be small and can be implemented as a LIFO array (stack). Typically the matching mutex for the object will be the first or second entry, so the check should be cheap. +In addition to the object states, *local*, *immutable*, *shared*, +and *protected*, the VM will also maintain two internal states, +*local-immutable* and *local-synchronized*, that are semantically equivalent to +*immutable* and *synchronized*, but will allow the performance advantages of +*local* reference counting. + +As far as the VM is concerned, *immutable* and *synchronized* are the same, as +both can be shared. + ++--------------------+------------+----------------+ +| State | owner_id | state | ++====================+============+================+ +| Local | > 0 | LOCAL | ++--------------------+------------+----------------+ +| Protected | < 0 | PROTECTED | ++--------------------+------------+----------------+ +| Synchronized | == 0 | SYNCHRONIZED | ++--------------------+------------+----------------+ +| Immutable | == 0 | IMMUTABLE | ++--------------------+------------+----------------+ +| Local-synchronized | > 0 | SYNCHRONIZED | ++--------------------+------------+----------------+ +| Local-immutable | > 0 | IMMUTABLE | ++--------------------+------------+----------------+ +| Inaccessible | INT_MAX | INACCESSIBLE | ++--------------------+------------+----------------+ + +Inaccessible objects can be used for testing and possibly as sentinels +internally. + + +Combining reference counting and access control +''''''''''''''''''''''''''''''''''''''''''''''' + +Like biased reference counting, each object will have two reference counts: +a local reference count, that will not need synchronization, and a shared +reference count that will need synchronization. + +In general, anywhere that an access control is needed a reference count +increment also necessary. By combining the two, no more tests are needed +and the lowest overhead increment can be performed. + +:: + + + static inline int + do_shared_incref(PyObject *op, uint32_t increment) { + int refcount = _Py_atomic_load_relaxed(op->ref_count_shared); + if (refcount >= IMMORTALITY_INCREMENT_THRESHOLD) { + return 0; + } + _Py_atomic_increment(&op->ref_count_shared, increment); + } + + static inline void + do_local_incref(PyObject *op) { + op->ref_count_local++; + if (op->ref_count_local != 0) { + return 0; + } + // Overflowed, so move 128 from local to shared. + op->ref_count_local = 128; + return do_shared_incref(op, 128); + } + + int + _PyObject_CheckAccessIncref(PyObject *op, PyThreadState *tstate) + { + uint32_t owner_id = op->owner_id; + if (owner_id == tstate->threadgroup_id) { + // This includes local-immutable objects. + return do_local_incref(op); + } + else if (owner_id == 0) { + // immutable or synchronized + return do_shared_incref(op, 1); + } + else if (owner_id < 0) { + // Protected + if (protected_check(op)) { + // unsynchronized refcount is allowed + return do_local_incref(op); + } + } + if (in_stop_the_world()) { + return do_local_incref(op); + } + PyErr_SetException(IllegalAccessException, op); + return -1; + } + + +Exploiting the state of the container object +'''''''''''''''''''''''''''''''''''''''''''' + +In practically all cases where a thread is getting a reference to an object +from the heap that reference is in a container object that we know to be +accessible. In many cases both objects will have the same accessibility, +so we can avoid expensive checks by comparing the ``owner_id`` fields. +If they are the same then the new reference is legal:: + + int + _PyObject_CheckAccessIncrefWithContainer(PyObject *op, PyObject *container) { + assert(PyObject_IsAccessible((container)); + if (op->owner_id != container->owner_id) { + return PyObject_CheckAccessIncref(op); + } + // Object could be in any state, but it is a legal one + int increment; + if (op->owner_id != 0) { + do_local_incref(op); + } else { + + } + int refcount = _Py_atomic_load_relaxed(op->ref_count_shared); + if (refcount >= IMMORTALITY_INCREMENT_THRESHOLD) { + return 0; + } + _Py_atomic_increment(&op->ref_count_shared, increment); + return 0; + } + } C API ----- @@ -118,7 +245,6 @@ implementation would first be renamed ``PyObject_FooUnchecked``, then return _PyObject_CheckAccessNullable(result); } - where ``_PyObject_CheckAccessNullable`` is an internal function providing the access control check. A ``_PyObject_CheckAccess`` variant would be provided for when the object reference was known to not be ``NULL``. @@ -126,13 +252,23 @@ provided for when the object reference was known to not be ``NULL``. This mechanical transformation is likely to leave some inefficiencies in the code base, so additional work will be needed to re-optimize later. +To take advantage of the combined access and incref function above, some +modification will be needed, for example:: + + PyObject * + PyObject_Foo(PyObject *op) + { + PyObject *result = PyObject_FooBorrowed(op); + return _PyObject_CheckAccessXIncref(result); + } + Since all API functions need to check against the current thread, new APIs taking a reference to the thread will be added to reduce the overhead of fetching the thread reference on every call. For example ``PyObject_GetAttr`` would gain a ``PyObject_GetAttrThread`` variant:: - PyObject *PyObject_GetAttrThread(PyObject *v, PyObject *name, PyThread t); + PyObject *PyObject_GetAttrThread(PyObject *v, PyObject *name, PyThreadState *t); Variants of ``_PyObject_CheckAccess`` that take a thread pointer will be added. @@ -142,6 +278,11 @@ always returns a ``str``, which is immutable, so no additional access check is needed. ``PyObject_SetItem`` does not return an object, so will need no additional check. +.. Note:: + Given we are adding new functions to the C API, the exact interface will + need further discussion. Particularly whether a new opaque value should + be used instead of ``PyThreadState *`` and what it should be called. + Interpreter ----------- @@ -278,15 +419,15 @@ Implementing ownership will require the ABI breakage discussed above. With that in mind, here is a possible order of implementation: -* ThreadGroups * One-time ABI breakage +* ThreadGroups for testing and development only +* Support parallel allocation and cyclic garbage collection * Port biased and deferred reference counting from the free-threading build * Simple ownership. Local and immutable only -* Support parallel allocation and cyclic garbage collection +* ThreadGroups API * ``__freeze__`` * Synchronized objects * Protected object state, including bytecode compiler support -* ``TransferBox`` and ``Channel`` * ``sys.monitoring.StopTheWorld`` * Performance work