diff --git a/_typos.toml b/_typos.toml index 8ed332a..ad3ef79 100644 --- a/_typos.toml +++ b/_typos.toml @@ -53,3 +53,5 @@ nd = "nd" ons = "ons" te = "te" writeable = "writeable" +BOCpy = "BOCpy" +Cpy = "Cpy" # BOCpy read as 'CPy'? diff --git a/content/posts/language-summit-2026-developer-in-residence-update-and-future/image.png b/content/posts/language-summit-2026-developer-in-residence-update-and-future/image.png new file mode 100644 index 0000000..27d989f Binary files /dev/null and b/content/posts/language-summit-2026-developer-in-residence-update-and-future/image.png differ diff --git a/content/posts/language-summit-2026-developer-in-residence-update-and-future/index.md b/content/posts/language-summit-2026-developer-in-residence-update-and-future/index.md new file mode 100644 index 0000000..af1b85b --- /dev/null +++ b/content/posts/language-summit-2026-developer-in-residence-update-and-future/index.md @@ -0,0 +1,44 @@ +--- +title: 'Developer-in-Residence Update & Future (Python Language Summit 2026)' +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Petr Viktorin gives an update on the Developer-in-Residence role and asks Python core developers for projects to prioritize' +tags: [language-summit, language-summit-2026] +published: true +--- + +The first three [Python Developers-in-Residence](https://www.python.org/psf/developersinresidence/) are Łukasz Langa, Petr Viktorin, and Serhiy Storchaka. Now that Developer-in-Residence Łukasz is [moving on from the role after five years](https://pyfound.blogspot.com/2026/04/reflecting-on-five-years-as-psfs-first.html), the now-“ironically named” *Deputy* Developer-in-Residence Petr came to the Language Summit to give an update on the role and ask core developers what’s next. + +![GitHub profile pictures for Łukasz, Petr, and Serhiy with Łukasz departing](image.png) + +Petr started the discussion by listing the responsibilities in his contract: + +* Keeping the core development workflow operational to ensure core developers and collaborators are not blocked from contribution. +* Improving tools and workflows, which are part of the core development experience, to enable a smoother contribution process. +* Code review of a subset of incoming changes to CPython to help collaborators receive timely communications on their changes. +* Authoring own changes to CPython to help clean the issue and PR backlog. +* Responding to requests for assistance from core developers in a timely manner. +* Regularly reporting on this work to the Steering Council and the wider community. +* Additional responsibilities as directed. + +Referencing the explicit “Responding to requests for assistance from core developers”, Petr noted that there “haven’t been that many requests for assistance” recently, so he wanted to ask core developers to communicate what they might want done, either at the Language Summit or later. + +Petr continued with an update on what he’d accomplished from 2024 to 2026, including maintaining [buildbots](https://devguide.python.org/testing/buildbots/), mentoring multiple people into becoming triagers or core developers, working on the [Stable ABI](https://docs.python.org/3/c-api/stable.html) for [free-threaded Python](https://docs.python.org/3/howto/free-threading-python.html), working on security vulnerability fixes, and managing CPython sprints at conferences. He compared this work to what Łukasz had accomplished in his five years of tenure, which included more “large-scale project management”, organizing events like the Language Summit, talking to sponsors, and overseeing the other Python Developers-in-Residence. + +Petr summarized his approach to the role as focusing on “important tasks, but leaving fun and glamorous ones to volunteers”. He also noted collaborating with people in other “Developer-in-Residence”-like roles, such as Hugo van Kemenade and Stan Ulbrych, who are Sovereign Tech Agency fellows focusing on CPython. Petr opened the floor for discussion by asking what challenges core developers would like the Developers-in-Residence to focus on. + +## Discussion + +“Thank you for doing all the boring stuff”, Ken Jin opened, appreciating Petr taking on grunt work that enables other core developers, and sharing that if he had to do this work as a volunteer he “would have quit a long time ago”. + +“You don’t have to do everything Łukasz did”, Thomas Wouters assured Petr, “we’re hiring a replacement for Łukasz and will sort out what that means for the team”, reminding everyone that the Steering Council members are all volunteers themselves, that managing multiple full-time employees is difficult, and that hiring takes time. + +Former Developer-in-Residence Łukasz Langa chimed in with a problem he and other core developers had noticed: the large volume of (likely) LLM-generated pull requests being submitted to CPython. One of Łukasz’s early goals for CPython [when he joined as Developer-in-Residence](https://lukasz.langa.pl/a072a74b-19d7-41ff-a294-e6b1319fdb6e/) was “addressing the PR backlog”, which was successful for some time, but that progress has now completely reversed with LLMs. Łukasz suggested either triaging these PRs somehow, noting he wasn’t sure “how many full-time jobs that task is”, or coming up with “some systematic solution”. +Petr confirmed triaging every LLM pull request wasn’t something he wanted to do, but he was “hired to also do tasks he didn’t want to do”. + +There was some discussion between Stefan Behnel and Mark Shannon about ensuring some amount of review “before a human looks at a pull request”, such as having an LLM provide a first pass on a pull request. Savannah Ostrowski was hesitant to engage with drive-by LLM contributions: “I’m not willing to give away my time because it won’t change their behavior”. She suggested that other core developers “nope out” if they didn’t want to engage. + +Mark Shannon recommended everyone [watch Pablo Galindo Salgado’s keynote on this topic](https://www.youtube.com/watch?v=e8uozuvRf7g). +There is also a Language Summit Lightning Talk by Gregory P. Smith about attempting to +steer or improve these contributions [using `AGENTS.md`](/2026/09/language-summit-2026-lightning-talks#agentsmd-for-cpython). diff --git a/content/posts/language-summit-2026-free-threading-post-era/image.png b/content/posts/language-summit-2026-free-threading-post-era/image.png new file mode 100644 index 0000000..6cf0521 Binary files /dev/null and b/content/posts/language-summit-2026-free-threading-post-era/image.png differ diff --git a/content/posts/language-summit-2026-free-threading-post-era/index.md b/content/posts/language-summit-2026-free-threading-post-era/index.md new file mode 100644 index 0000000..85a9b8d --- /dev/null +++ b/content/posts/language-summit-2026-free-threading-post-era/index.md @@ -0,0 +1,79 @@ +--- +title: Free-Threaded Python Post-Era (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Tobias Wrigstad, Fridtjof Stoldt, and Donghee Na propose a safe and performant, high-level concurrency model for free-threaded Python' +tags: [language-summit, language-summit-2026] +published: true +--- + +Tobias Wrigstad and Fridtjof Stoldt returned to the Python Language Summit, now joined by Donghee Na. Tobias and Fridtjof previously presented “[Fearless Concurrency](https://pyfound.blogspot.com/2025/06/python-language-summit-2025-fearless-concurrency.html)” to the Python Language Summit in 2025. This year the topic at hand was the “post-Free-Threading era of Python”, and what high-level concurrency primitives would be provided by Python. + +## Comfortable doesn’t mean “Good” + +Today the interface for accessing free-threading is, unsurprisingly, “threads”. Threads are how people typically first learn about true parallelism from school, textbooks, and other familiar materials. But what if threads as a user interface aren’t very good? Using threads means users need to care about deadlocks and race conditions. + +Python’s free-threading project has made considerable progress since [PEP 779](https://peps.python.org/pep-0779/)’s acceptance. But there was one aspect of PEP 779’s acceptance criteria that hadn’t been addressed yet: high-level concurrency primitives. The complete message from the Steering Council’s acceptance stated: + +> Preparation for high-level concurrency primitives. \ +> The Python core team should begin considering and proposing higher-level concurrency primitives that users can use safely and effectively, without requiring a deep understanding of the underlying threading mechanism. And the SC wishes that this task should be prioritized once the above tasks are stable. We recommend using the `concurrent` package in the stdlib for this, where appropriate. + +There is some prior art here; other programming languages provide primitives Python can be inspired by, such as Rust’s ownership and borrowing, Go with channels and goroutines, and Erlang with actors and sending messages. + + +## Fast, Simple, or Safe: what should Python concurrency be? + +The talk opened by contrasting approaches to implementing concurrency in different programming languages on three separate axes that are “sometimes at odds with each other”: speed, simplicity, and safety. + +C is fast and simple, offering direct access to memory through pointers, which also allows dangerous and unsafe operations; C relies on programmer discipline for program correctness. Erlang has multiple threads, but to communicate, data is copied over to the other “actor”, so it’s safe and simple, but performance suffers. Rust is performant and safe, but it’s not simple: “you’ll find yourself often fighting with the [borrow checker](https://doc.rust-lang.org/book/ch04-00-understanding-ownership.html)”. + +![](image.png) + +“Where would we put Python on this triangle?” + +Python using the [Global Interpreter Lock (GIL)](https://docs.python.org/3/glossary.html#term-global-interpreter-lock) (represented as a single lock on the graphic) is a simple model, but it’s not completely safe, as the GIL “only protects the runtime”. [Subinterpreters](https://docs.python.org/3/library/concurrent.interpreters.html) (represented as multiple locks on the graphic) as a model are safer due to the isolation, but this model isn’t simple or performant due to data copying. Finally, there is [free-threaded Python](https://docs.python.org/3/howto/free-threading-python.html) (racecar in the graphic), which is very similar to C: simple and performant, but it puts the burden on programmers for program correctness and safety. + +The three argued that for a high-level model “safety would be a priority and then performance”, with “simplicity as third”. “Simplicity is at odds with performance”. + +## Behavior-Oriented Concurrency (BOC) + +Behavior-Oriented Concurrency (BOC) is the model the trio is proposing for a high-level concurrency interface for Python. Programs written with the BOC model are task-based. Tasks are lightweight, can run on any core, and cannot deadlock. + +Each task owns data that is protected by mutexes, but unlike the [`threading.Lock`](https://docs.python.org/3/library/threading.html#lock-objects) objects that we’re used to in Python, these mutexes are aware of the data that they protect. This data awareness means that the mutexes can ensure that the data objects are only accessed by a single task at a time, providing isolation and preventing data races. BOC calls these mutexes “Cowns”, meaning “concurrent owner”. + +BOC allows defining which mutexes a task depends on using the `@when` decorator. One or more mutexes can be specified in this decorator, and tasks can only access data through mutexes that the task depends on. This mechanism is what allows BOC to create a dependency graph across all tasks and data. + +Providing a scheduler with a program written with the above constraints gets you three important properties: no deadlocks, no data races, and good “locking discipline”. The user “wouldn’t need to decide to spawn a thread or subinterpreter”; this would be handled by the runtime. The scheduler would be able to turn detected deadlocks or data races into exceptions instead, detecting these situations ahead of execution. + +Some examples of programs written using locks and threads were then transformed into programs using the BOC model, such as two bank accounts transferring money between them and printing the results: + +There are [many more examples available on GitHub](https://github.com/microsoft/bocpy/tree/main/examples). This proof of concept, implemented using subinterpreters, is available for anyone to try: [bocpy](https://microsoft.github.io/bocpy), available on the [Python Package Index](https://pypi.org/project/bocpy): + +```commandline +$ python -m pip install bocpy +``` + +Checking bocpy against their own criteria of simplicity, safety, and performance: bocpy is simple, as can be seen in the above examples. On the safety front, bocpy today runs in “stable Python”, and for this reason only isolation has been implemented so far. “We don’t yet have ownership, you can’t implement [ownership] as a third-party library”, as this would require changes to the runtime. Performance is “good for programs that aren’t communication dominated” due to “communication being expensive for subinterpreters”. + +The bocpy package provides three features: + +* “[Cowns](https://microsoft.github.io/bocpy/#cowns)”, or locks with ownership +* [Behaviors](https://microsoft.github.io/bocpy/#behaviors) (the tasks spawned with `@when` decorators) +* Scheduler (implemented with subinterpreters, with planned support for free-threading) + +Looking forward, the three have two other proofs of concept that modify the Python runtime “to allow for safe concurrency and create safe abstractions while keeping performance”. The first implements isolation by organizing the heap into isolated groups of objects, where ownership violations would raise an exception from Cowns. The second is for immutability of Python objects ([PEP 795](https://peps.python.org/pep-0795/)). + +The group made it clear that changes to core Python would be needed to support safe concurrency, asking whether Python would trade “some performance for safety”. What primitives do we want to provide in the standard library; are locks enough? And if we do provide primitives, should they be something like bocpy? + +## Discussion + +Thomas Wouters recalled that the core team “has experience trying to create universal interfaces for subprocesses, multiprocessing, threading”. In practice, there are always corner cases, and performance is suboptimal because of the constraints of the APIs. Thomas asked “how confident [the three] are that this isn’t the case for bocpy?” The three shared Thomas’s concern. “This is a question we’re working on”, answered Fridtjof. “If you have [the bocpy] ownership model it’s possible to treat subinterpreters and threads similarly, you can have communication, and you can share objects directly with the ownership model”. + +“Notion of tasks and schedulers immediately brings async to my head”, David Hewitt said, wondering “how does async fit into this picture?” He asked the trio whether bocpy “should be built on [`asyncio`](https://docs.python.org/3/library/asyncio.html)” and have “async mutex primitives instead of being [synchronous]”. Tobias confirmed that bocpy “could be” built using `asyncio`: “it depends on what backend infrastructure” is used and the trade-offs of each backend. “If you want to do some I/O, you tell the I/O library to put something into a Cown once data is available and schedule a task to run when the data is available to avoid blocking”. + +Larry Hastings was “happy to see this research going on” and welcomed more from the group, but didn’t see bocpy or any singular solution as “the” method to do concurrency in Python. “[Larry] would like Python to have all the tools that give you and other groups with competing ideas the ability to implement ideas and provide them to users”. Instead, Larry was wary of “anointing a single way”, to avoid locking Python into a particular implementation in case better options are discovered later. Donghee shared that the group wasn’t initially trying to force a single way, only to start the conversation. + +Fridtjof brought up Rust as an example of where not providing high-level primitives on top of `async` and `await` resulted in a “split” in the Rust ecosystem. “Sounds lively”, Larry responded. + +David Hewitt shared that a lot of the Rust community regards the “Rust async situation” as “slightly failed”, as the standard library lacked any standard interface for the runtime. As a result, the entire Rust ecosystem has been “forced” to converge on [Tokio](https://tokio.rs/). David agreed with Thomas that it would be “tough to come up with a good abstraction”, but that this “doesn’t mean we shouldn’t try”. diff --git a/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/image.svg b/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/image.svg new file mode 100644 index 0000000..292ed55 --- /dev/null +++ b/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/image.svg @@ -0,0 +1,1509 @@ + + + + + + + + 2025-02-14T08:48:45.406654 + image/svg+xml + + + Matplotlib v3.10.0, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/index.md b/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/index.md new file mode 100644 index 0000000..4e7aa32 --- /dev/null +++ b/content/posts/language-summit-2026-garbage-collection-generational-incremental-both/index.md @@ -0,0 +1,55 @@ +--- +title: 'Garbage Collection: Generational? Incremental? Both! (Python Language Summit 2026)' +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Mark Shannon proposes a future garbage collection strategy for Python following the revert of the incremental garbage collector in Python 3.14' +tags: [language-summit, language-summit-2026] +published: true +--- + +The third Language Summit talk was brought by Mark Shannon, who is the author of the [incremental garbage collector](https://github.com/python/cpython/issues/108362) implementation shipped in Python 3.14 that was [reverted back to the generational garbage collector](https://discuss.python.org/t/reverting-the-incremental-gc-in-python-3-14-and-3-15/107014) from Python 3.13 after reports of “significant memory pressure” in production environments. The original goal of the new incremental garbage collector was to reduce maximum pause times by an order of magnitude for larger heaps. + +Mark’s talk opened with a graph about where Python spends its time, split between the interpreter, lookups, modules, and the focus of the talk: garbage collection, which takes around 11.67% of execution time. Mark remarked that even if the [Just-in-Time (JIT) compiler](https://peps.python.org/pep-0744/) makes the interpreter faster (representing 30.6% of time), we’ll unfortunately still have to worry about memory management and garbage collection to make the runtime faster: Mark’s talk was about minimizing the time spent doing these tasks. + +![](image.svg) + +What does Python’s garbage collector do today? Funnily enough, this ~12% of time spent is not Python’s primary garbage collection mechanism; [reference counting](https://docs.python.org/3/glossary.html#term-reference-count) is. The garbage collector is a “backup” and “should be much faster”. Reference counting is already collecting dead objects, therefore the garbage collector should only have to find dead unreachable cycles. + +Mark emphasized pause times during garbage collection, a metric which the incremental garbage collector aimed to improve, as a major issue for servers and applications with a user interface. The incremental garbage collector reduced peak pause times during garbage collection to “tens of milliseconds”, down from “around 3 seconds for the current generational garbage collector”. + +So how would we know whether we’ve improved the garbage collector? Mark defined a term “effectiveness” along with other definitions that will be useful for thinking about different approaches to garbage collection: + +* **Scavenge**: This is the minimal garbage collector operation. Give the garbage collector a set of objects and all unreachable cycles are collected. At a high level, users don’t have to worry about how the garbage collector accomplishes this task. +* **Scavenge Effectiveness**: This is the metric being optimized for, calculated as “objects collected” divided by the “number of objects visited”. The effectiveness of the two Python garbage collectors is “pretty poor”, with the old generational GC at 0.3% effectiveness and the reverted incremental GC at ~1% effectiveness. +* **Generational hypothesis**: The assumption that “most objects die young”, which garbage collectors can use to be more effective. However, the exact profile depends on the program being executed. +* **Spaces**: Grouping of objects that are consecutively allocated. All new objects are added to the newest “space”. Spaces have a defined size, and once they are full, the GC can begin scavenging the space and a new empty space is created for newly created objects to be added to. The period when spaces are closed to new objects is a good opportunity to scavenge, as no scavenge operations can occur on an open space. +* **Generation**: A group of consecutive spaces. Generations are much fewer than spaces and typically baked into the garbage collector. Spaces only move in one direction through generations. + +Hood Chatham wondered about the “effectiveness” of other garbage collectors compared to Python’s. Mark noted that direct comparison between other programming language garbage collectors and Python would be “unfair”, because “nobody else does reference-counting *and* garbage collection”. + +Mark then turned to explaining the now-reverted incremental garbage collector. This garbage collector was a non-generational collector, meaning there was only a single generation that contained the whole heap. This made the incremental garbage collector “effective” because spaces that have existed for longer tend to have more objects collected. The trade-off is that non-generational GCs have a negative impact on memory usage. + +## Can we have the best of both worlds? + +Generational garbage collectors have lower memory usage overall but suffer from higher pause times, whereas incremental garbage collectors pause for less time but at the expense of more memory. What should Python do as a default? +Mark’s final proposal included how he’d interleave the concepts of generational and incremental garbage collectors to achieve the best of both: + +* Two generations: one “young” and one “old”. +* Alternate between young and old generation incremental scavengers. +* Fix the “young” generation as 20MB in size initially. +* Scavenge the old generation at twice the young generation survivor rate. + +Donghee Na was concerned that any changes to the garbage collector would cause issues for someone. Donghee wondered whether there was a way to implement a transition period between different garbage collectors. Mark answered that there’s nothing inherently “wrong” with the current garbage collector and that the primary issue would be that there would be “more code to maintain”, but he would prefer providing a GC that is better by default unless users are fine-tuning the GC themselves. + +Donghee asked whether configuration options similar to what is available for [JVM garbage collectors](https://docs.oracle.com/en/java/javase/21/gctuning/) could be made available so users could tweak settings to fit their needs. Mark didn’t think this should be necessary; in JVM languages the GC is “the whole thing”, compared to Python where GC is only the “backup” behind reference counting. + +Jukka signaled his interest in helping Mark and asked about existing applications that had fine-tuned their GC, noting that these fine-tunings “go stale” whenever the GC changes, even if the new GC is better on average. Mark hoped that if done correctly, Python could remove the need to fine-tune GCs. + +Gregory P. Smith lamented that adding [public APIs for the garbage collector](https://docs.python.org/3/library/gc.html) and allowing fine-tuning inhibits being able to improve the general case. Greg also didn’t want to support multiple GC implementations like the JVM does. Greg’s biggest takeaway from the incremental GC revert in Python 3.14 is that there are use-cases which are served well by one GC and not by another, and “those cases should be in our test suite”. + +Tobias Wrigstad suggested a potential fine-tuning mechanism that was already being adopted by Java: providing the garbage collector with an explicit “CPU budget”. The mechanism “seemed fairly intuitive”, but it also has the unfortunate side effect of letting users tune their GC “such that it cannot collect all the dead memory”. + +Pablo Galindo Salgado had run into issues with the GC and saw success with being able to fine-tune GC generations live in production to get instantaneous feedback. Thomas Wouters (shocking his fellow Steering Council members) agreed with Pablo on being able to fine-tune the garbage collector to “remediate pathological GC behavior”, also adding that this was a common blocker for being able to upgrade Python versions. Mark was interested in these numbers: “let’s get those before moving forward with any decisions”. + +Larry Hastings asked about concurrent or “lock-free” garbage collectors and whether there’s a possibility for this to happen in Python. “Memory back for free sounds like a wonderful sales pitch”. Unfortunately, in addition to being “really hard” to implement (which Larry countered with “I heard you were smart”), Mark shared that [C extensions](https://docs.python.org/3/extending/index.html) can “do whatever they want”, which would interfere with a lock-less concurrent collector. Tobias agreed that concurrent lock-less GCs would be “extremely invasive”, including needing to add metadata embedded in every pointer. diff --git a/content/posts/language-summit-2026-lightning-talks/index.md b/content/posts/language-summit-2026-lightning-talks/index.md new file mode 100644 index 0000000..c509ee4 --- /dev/null +++ b/content/posts/language-summit-2026-lightning-talks/index.md @@ -0,0 +1,106 @@ +--- +title: 'Lightning Talks (Python Language Summit 2026)' +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Lightning talks on a one-time ABI break, safer interruptions, EktuPy (Scratch but Python), an AGENTS.md file for CPython, and a call to read PEP 836.' +tags: [language-summit, language-summit-2026] +published: true +--- + +## One-time ABI breakage + +By Mark Shannon + +Why break the [Stable ABI](https://docs.python.org/3/c-api/stable.html)? Because there are 32 spare bits in the +[`PyObject`](https://docs.python.org/3/c-api/structures.html#c.PyObject) header that we currently can’t use. + +```c +struct _object { + Py_ssize_t ob_refcnt; // Mark wants the high bits in this struct. + PyTypeObject *ob_type; +}; +``` + +Since [PEP 683](https://peps.python.org/pep-0683/) was accepted in Python 3.12, +`ob_refcnt` has effectively been treated as a 32-bit integer, but these +bits are still inaccessible for other uses to prevent breakages. +If the Python core team could repurpose these +bits, Mark could implement a better garbage collector and faster allocations. + +Breaking the Stable ABI would be “easy to do... probably not painless though”, +in reference to users needing to port, recompile, and support multiple ABIs. + +“Well... it’s going to happen anyway”, Mark said, referencing `abi3` and `abi3t` for +[free-threading](https://docs.python.org/3/howto/free-threading-python.html), +and how Python packages depending on the Stable ABI would need to move to the +new `abi3t` ABI. + +Thomas Wouters provided an answer: wait until there is only a single Stable ABI, `abi3t`. +Free-threaded Python has not promised Stable ABI compatibility +for its object layout, unlike non-free-threaded Python, which exposed +details about `ob_refcnt` in the Stable ABI. + +It was at this point that it began slowly dawning on Mark, to his horror, that free-threaded +Python may be solving his problem. After free-threading becomes the new Python default, +the Stable ABI that exposes the object header internals will be no more, and Mark +and team can make changes to the object header more freely. We’ll just have to be patient! + +## Safer and Generic Interruptions + +By Daniele Parmeggiani + +Daniele presented a gap in implementing structured concurrency with threads in Python today: +the inability to safely interrupt tasks that failed or have been cancelled. + +For example, let’s say there were two +I/O-bound parallel database queries sent as a result of a web request (task A and task B). If task A returned early with +an error, the result of task B would no longer be needed, because an exception would be raised anyway. +However, right now there is no way to cancel task B, so instead the result of task B is waited for +and then thrown away before propagating the exception from task A. + +Interrupts are one way to implement this cancellation mechanism. +Daniele explained that the machinery to implement cancellations via interrupts is “already available in CPython, +but isn’t exposed at the Python level”. And why isn’t this functionality exposed? Because it’s a huge source of “footguns”. + +Instead of exposing the unsafe functionality directly, Daniele would like to offer users +a way to handle interruptions safely. To do this, Daniele proposes adding uninterruptible scopes, +where a [context manager’s](https://docs.python.org/3/reference/datamodel.html#context-managers) `__enter__()` and `__exit__()` methods are +“shielded” from being interrupted and are guaranteed to execute. + +“You can’t really rely on [context manager cleanups] with interruptions right now”. +If you read the documentation of the [`signal` module](https://docs.python.org/3/library/signal.html#note-on-signal-handlers-and-exceptions), +it tells you to turn off signals because they aren’t safe. “It’s a bit weird that the documentation for signals says that”. + +“Is this good for performance? No, it adds more bytecode”, Daniele said, closing his presentation, “but +I’ve made sure to make Mark Shannon happy”. Daniele invited anyone who is interested +in an implementation of this functionality to contact him. + +## EktuPy, Scratch but Python + +By Kushal Das + +Kushal Das brought a short presentation on a project he’d been working on to teach Python to the generation of young programmers who learned using Scratch. [Scratch](https://scratch.mit.edu/) is a programming language that is represented using blocks instead of text to create animations, games, and other media-focused programs. + +EktuPy brings many of the features that are beloved in Scratch, such as the focus on media like games and animations, and the “remix” concept to give new programmers a working base to start with instead of a daunting blank canvas. EktuPy tries to bridge the gap between block-based programming languages and text-based languages like Python. + +## AGENTS.md for CPython + +By Gregory P. Smith + +CPython, just like many other open source projects, has been seeing many contributions from folks using LLM agents. Gregory P. Smith and Łukasz Langa wanted to get a “vibe check” from core developers about adding a simple `AGENTS.md` file to the CPython repository. The hope was that even basic guidance about how to contribute to CPython (such as linking to the [Developer Guide](https://devguide.python.org/)) would improve the quality of the large number of pull requests made using agents. + +Greg asked for a show of hands for “who would be against adding an `AGENTS.md` to the repository”, and no core developer hands were raised. “Oh, that was easier than I thought”. Thomas Wouters remarked that he “wished it wasn’t necessary” to have an `AGENTS.md`. + +Another issue raised was that `AGENTS.md` files tend to receive many outside contributions from people wanting to add new things. Greg made it clear that the file would be kept very simple and “seldom change”, suggesting that the core team would “reject pull requests that modify the file”. + +Seth Larson asked about quantifying the improvement to contributions, citing that the Python Security Response Team had seen increased volumes of LLM-generated reports. Changes to the [security policy](https://www.python.org/dev/security/) and adding a threat model appeared to work, but this was difficult to prove. + +Łukasz shared his pessimistic view that CPython would someday have to stop accepting many outside contributions due to the large number of agent-driven contributions. + + +## Please read PEP 836 (JIT go brrr) + +By Ken Jin + +Closing out the Language Summit lightning talks, Ken Jin had a simple request: to [please read PEP 836](https://peps.python.org/pep-0836/). Ken had [only received 33 replies on the PEP discussion thread for PEP 836](https://discuss.python.org/t/pep-836-jit-go-brrr-the-path-to-a-supported-jit-compiler-for-cpython/108010), and encouraged core developers (and anyone reading this blog post) to read the PEP and “tell us what you disagree with”. diff --git a/content/posts/language-summit-2026-macos-python/index.md b/content/posts/language-summit-2026-macos-python/index.md new file mode 100644 index 0000000..289c696 --- /dev/null +++ b/content/posts/language-summit-2026-macos-python/index.md @@ -0,0 +1,53 @@ +--- +title: macOS and Python (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Ned Deily weighs whether Python should continue shipping macOS installers' +tags: [language-summit, language-summit-2026] +published: true +--- + +Ned Deily, the macOS expert and release manager, shared that this was the first time discussing this topic in 15 years of Python Language Summits. + +When CPython does a release, the main result of the release is source distributions: tarballs and ZIP archives. It’s then up to downstream packagers and distributions to build from these sources for their individual platforms. Historically, though, CPython has also provided some pre-compiled installers for Windows and macOS. This was done because these two platforms are “different” from most others. + +The question weighing on Ned was whether it was still worth shipping [macOS installers](https://docs.python.org/3/using/mac.html) as had been done in the past. “macOS has come a long way” and was no longer comparable to Windows “in terms of strangeness”. Ned admitted that the decision to continue was mostly on “autopilot”, that the macOS installers were “off in a dark corner with only a few people involved”, and that for certain stretches that were “too long”, he was the only one maintaining this functionality. + +Ned wanted to answer whether Python should “continue to [ship macOS installers]”? And if so, then it was “long-past due” to bring the knowledge about the macOS installer “more to the forefront” and inform other Python core developers what they need to know about changes that affect the macOS platform. “A lot of what happens for macOS also applies to iOS”, Ned continued, “now that we provide compiled binaries for [iOS], too”. + + +## What’s different about macOS? + +macOS has three different build types: “static” and “shared”, which are like Unix builds, and “framework”, which is unique to macOS and iOS. [Frameworks](https://developer.apple.com/library/archive/documentation/MacOSX/Conceptual/BPFrameworks/Concepts/WhatAreFrameworks.html) originated in [NeXTSTEP](https://en.wikipedia.org/wiki/NeXTSTEP), carried over into Mac OS X, and are used for system interfaces and APIs. macOS can build objects into files that support multiple processor types, such as ARM64 and x86-64, in a single file, which Ned called “fat files”. Apple developer tools (ADT) provide support for transparently building and linking multi-architecture files, including providing a version of Clang that handles this for you. iOS, watchOS, and tvOS use the same ADT, so they also benefit from this multi-architecture support. + +Apple Software Development Kit (SDK) files, such as header files and shared libraries, are not installed in traditional locations like `/usr`; instead, they live within the SDK. This allows using a single environment variable to switch between different SDK versions and makes building for any current Apple platform possible on a single macOS system (for example, building for iOS from macOS, or building for Intel x86-64 macOS on Apple Silicon). + +Ned noted that Python’s Framework builds make it easy to embed Python directly into a macOS app that “just works everywhere with no external dependencies”. + +But of course, it’s not all benefits; there are drawbacks, too. As-is, the macOS installer had accumulated “significant technical debt” over 25 years of support, didn’t follow macOS guidelines for file layouts, and didn’t support [macOS sandboxing](https://developer.apple.com/documentation/security/app-sandbox), which would be required to distribute Python via the App Store. Another drawback was that there was only one system-wide install location. The macOS Framework versioning scheme also didn’t align with the Python versioning scheme due to ABI compatibility issues. + +[Free-threaded Python](https://docs.python.org/3/howto/free-threading-python.html) (`python3t`) introduced a whole new separate Framework build instead of rolling support into the current Framework build, but Ned doesn’t think this distinction will last once free-threading is eventually made the default for Python. + + +## Who uses the official Python macOS distribution? + +“We don’t know, but we can make some guesses”, Ned said, sharing that his own mental profile of macOS installer users included “users on managed system environments such as centralized IT departments, public schools, engineering or scientific users, novice programmers, and experimenters”. [py2app](https://py2app.readthedocs.io/) uses the Framework builds, but “has to munge the build to make it work”. + +Ned listed a few other projects which build and distribute Python for macOS, including Homebrew, MacPorts, Conda, uv, ActiveState, and Apple themselves as part of Xcode for LLDB. These projects don’t use any of the builds available on [python.org](https://www.python.org); they all build the distribution themselves. + +Łukasz Langa asked whether the Python Developers Survey asked about macOS installers. Ned agreed that would be a good idea. + + +## What’s next? + +Ned closed the topic by listing his plans for macOS and Python. He noted that there was no section for macOS in [PEP 11](https://peps.python.org/pep-0011/), the PEP which details platform support for CPython, [whereas there was a section for Windows](https://peps.python.org/pep-0011/#microsoft-windows). He planned to propose such a section for macOS and get feedback from other core developers. More generally, Ned wanted a policy for “supporting macOS in general”, covering people who want to build Python themselves and detailing what is supported. +Currently, all three build types and architectures are considered the same in terms of “support” in PEP 11. Ned wondered whether each of these builds should be treated as a different PEP 11 target. This would allow macOS on ARM to be in a different support tier from macOS on x86-64. Ned would also like to leverage the large overlap in building and packaging for iOS and macOS. + +Ned would like to automate the build process, including the packaging and continuous integration, much like Windows already has as a part of the Python release process. In addition, Ned wants to modernize Python’s macOS support, including deprecating and removing the PPC universal builds and migrating to [XCFramework](https://developer.apple.com/documentation/xcode/creating-a-multi-platform-binary-framework-bundle) packaging instead of the legacy Framework format. “We’ve got a lot of cruft”. + +But not everything that’s no longer being used can be removed. Even though Apple is phasing out Intel Mac support entirely, there is still likely to be a need for at least the hooks for universal builds. “The universal fat-file build technology has been in use for over 30 years, and Apple has reached for it every time Apple migrates CPU architectures, you can almost bet that Apple will do something that requires universal builds again”. + +Nathan Goldbaum asked whether Ned planned to port [PyManager](https://github.com/python/pymanager) to macOS. Ned guessed that the internals would differ between Windows and macOS, but the user interface could be made as similar as possible for users. Ned shared that Russell Keith-Magee had created a small GUI app that may be a good jumping-off point for this work. + +Łukasz Langa referenced [MOPUp](https://github.com/glyph/MOPUp) (“**m**ac**O**S **P**ython.org **Up**dater”), a tool by Glyph Lefkowitz which updates the macOS installers from [python.org](https://www.python.org), and in particular noted an “advanced” feature of this tool: uninstalling Python versions. Ned agreed this was a desirable feature, noting that MOPUp was only a command-line tool, but that some years ago Glyph had created a manager that works in combination with the macOS installer. diff --git a/content/posts/language-summit-2026-memory-buffer-protocol/compat.png b/content/posts/language-summit-2026-memory-buffer-protocol/compat.png new file mode 100644 index 0000000..e7a796c Binary files /dev/null and b/content/posts/language-summit-2026-memory-buffer-protocol/compat.png differ diff --git a/content/posts/language-summit-2026-memory-buffer-protocol/image.png b/content/posts/language-summit-2026-memory-buffer-protocol/image.png new file mode 100644 index 0000000..0a0e4b7 Binary files /dev/null and b/content/posts/language-summit-2026-memory-buffer-protocol/image.png differ diff --git a/content/posts/language-summit-2026-memory-buffer-protocol/index.md b/content/posts/language-summit-2026-memory-buffer-protocol/index.md new file mode 100644 index 0000000..2ac30b3 --- /dev/null +++ b/content/posts/language-summit-2026-memory-buffer-protocol/index.md @@ -0,0 +1,77 @@ +--- +title: 'Memory Buffer Protocol (Python Language Summit 2026)' +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Nathan Goldbaum proposes safe concurrent access through buffer leases and custom data types for the Python Buffer Protocol.' +tags: [language-summit, language-summit-2026] +published: true +--- + +The [Python Buffer Protocol](https://docs.python.org/3/c-api/buffer.html) defines the semantics for accessing the underlying memory buffer of Python objects such as `bytes`, `bytearray`, and other types like `array.array`. The Python Buffer Protocol allows accessing an implementing object’s underlying memory directly, rather than only through higher-level APIs, which improves performance. + + +## Documenting the Buffer Protocol + +Nathan Goldbaum opened with some suggestions that were unlikely to be controversial, like improving the documentation and making the Buffer Protocol specification less fragmented. Buffers support a “struct-style” language that is similar to, but much more featureful than, Python’s [`struct` format strings](https://docs.python.org/3/library/struct.html#format-strings). “There are corners of the format language that are underspecified”. + +![Format string showing "T{>Q:t:(3)f:pos:T{H:id:xf:v:}:s:}"](image.png) + +Today, users are expected to read [PEP 3118](https://peps.python.org/pep-3118/) or implementations like [NumPy](https://numpy.org) to understand the full grammar, “including records, field names, subarrays, byte order, alignment, and complex numbers”. Documenting all features within the Python documentation would be a meaningful improvement. +Continuing with helping consumers of the Buffer Protocol, Nathan proposed creating a HOWTO guide for exporters and consumers of the Buffer Protocol, citing a lack of “complete examples” in C that implement validation, cleanup, ownership, and safe use of threads. + +## Safe concurrency + +Nathan was primarily looking for feedback on the upcoming sections of the talk, for which there are drafted PEPs addressing two shortcomings of the Buffer Protocol: [coordinating concurrent reads and writes](https://github.com/ngoldbaum/peps/blob/stable-views/peps/pep-9999.rst) and [custom data types](https://hackmd.io/@seberg/r1zB-s3tJl). + +The first issue was that the Python Buffer Protocol “didn’t support read and write coordination at the protocol level”. Storage addresses were stable, but contents were not. The Buffer Protocol supports a `readonly` attribute, but this value only applies to a single view, and writable access could not be made exclusive with existing APIs. + +Nathan’s proposal to add safe coordination of reads and writes was through new buffer “leasing” C APIs, `PyBufferLease` and `PyBufferAccessState`, similar to [Rust’s borrowing mechanism](https://doc.rust-lang.org/book/ch04-02-references-and-borrowing.html). These APIs would allow consumers to lease a buffer and declare their access pattern ahead of accessing a buffer view. Below is an example of how these new APIs would be used: + +```c +PyBufferLease *lease = PyObject_AcquireBufferLease(obj, + PyBUF_CONTIG_RO, + Py_BUFFER_ACCESS_SHARED_READ); + +if (lease == NULL) { + return -1; +} + +const Py_buffer *view = PyBufferLease_GetView(lease); + +Py_BEGIN_ALLOW_THREADS +consume_bytes(view->buf, view->len); +Py_END_ALLOW_THREADS + +PyBufferLease_Release(lease); +``` + +Leases could be opened in either `SHARED_READ` or `EXCLUSIVE_WRITE` mode, and calling `PyObject_AcquireBufferLease()` would not block, instead immediately either succeeding or failing if an incompatible lease already exists. `EXCLUSIVE_WRITE` would provide writable access and exclude every other access to overlapping memory. + +![Table showing which operations can co-exist. The current operation types (READ, WRITE) can coexist today, but are unsafe. SHARED_READ can coexist with itself and READ, all other operations can't coexist](compat.png) + +The existing [`PyObject_GetBuffer()`](https://docs.python.org/3/c-api/buffer.html#c.PyObject_GetBuffer), [`PyBuffer_Release()`](https://docs.python.org/3/c-api/buffer.html#c.PyBuffer_Release), and other related interfaces would remain unchanged under Nathan’s proposal. Buffer exporters would advertise explicit support for the new access modes, and exporters that don’t support the new access APIs would continue to use existing APIs as-is. Consumers would query the exporter using `PyObject_GetBufferAccessModes()` and use a lease if the desired access mode is available from the exporter. + +There is already an implementation being [worked on that is similar to this proposal](https://github.com/kumaraditya303/numpy/tree/view-tracking). [Kumar Aditya](https://github.com/kumaraditya303) is working on changes to NumPy allowing interoperability with the PEP, but Nathan noted that NumPy allowing access to raw pointers presented an “interesting” challenge. + +Nathan [published his PEP draft](https://github.com/ngoldbaum/peps/blob/stable-views/peps/pep-9999.rst), is requesting feedback, and will open a PEP discussion soon. + + +## Custom Data Types + +Finally, Nathan detailed a proposal that he and Sebastian Berg created that would add an extension point to the Buffer Protocol for custom data types. Today the Buffer Protocol format has a “fixed vocabulary”, meaning that to support custom data types, third parties either need to land a new type and format code in Python or use “out-of-band signaling and convention”. Instead of needing to block new types and format codes on a decision from the Python Steering Council, a namespace-based approach to custom data types could be used. + +In this proposed extension, a custom data type would be defined with square brackets (`[]`) with a `$` character delimiting the library namespace from the data type. Implementations that support the definition could then interpret known custom data types safely and reject any unknown custom data types without parsing the buffer. + +Nathan pointed to a discussion on [discuss.python.org](https://discuss.python.org/t/buffer-protocol-and-arbitrary-data-types/26256/13) and the [draft PEP](https://hackmd.io/@seberg/r1zB-s3tJl) for this proposal. + + +## Discussion + +Petr Viktorin asked whether a new API was needed for this proposal or whether the existing `PyObject_GetBuffer()` API could be extended. Nathan didn’t think the existing API, which is a part of Python’s [Stable ABI](https://docs.python.org/3/c-api/stable.html) and therefore can’t be changed, had enough space for a new field. “There’s an internal field, but it’s documented as being used by exporters, so we can’t use that one”. “We can’t smuggle information in the struct, we’d need a new struct to hold that information.” + +Thomas Wouters asked how Nathan’s proposal compares to Rust: when asking for exclusive write access, should the request block or error out? Nathan replied that erroring out would be the simplest implementation. David Hewitt asked about the composability of the implementation, and whether it could accommodate additional, likely desirable features like “blocking, asynchronously blocking, or more complicated structures like writing to slices”. + +Larry Hastings wondered whether this new API assumes that users are being diligent and using the API correctly. Nathan confirmed that memory corruption would be possible if the APIs were used incorrectly, but that pure-Python users wouldn’t be able to do this, only authors of [C extensions](https://docs.python.org/3/extending/index.html). Leaking memory wouldn’t be possible. + +Hood Chatham asked whether buffer exporters would be allowed to reject consumers using the old buffer access pattern without specifying either `SHARED_READ` or `EXCLUSIVE_WRITE`. Doing this would enable using a Rust buffer as the backing storage. diff --git a/content/posts/language-summit-2026-memory-snapshots/image.png b/content/posts/language-summit-2026-memory-snapshots/image.png new file mode 100644 index 0000000..71b768c Binary files /dev/null and b/content/posts/language-summit-2026-memory-snapshots/image.png differ diff --git a/content/posts/language-summit-2026-memory-snapshots/index.md b/content/posts/language-summit-2026-memory-snapshots/index.md new file mode 100644 index 0000000..b3ba787 --- /dev/null +++ b/content/posts/language-summit-2026-memory-snapshots/index.md @@ -0,0 +1,45 @@ +--- +title: Memory Snapshots for CPython (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Hood Chatham proposes memory snapshots and an initialization phase for speedier Python startups' +tags: [language-summit, language-summit-2026] +published: true +--- + +If you open up a new Python interpreter, close it, and then open a new Python interpreter again… how different are these two processes? Hood Chatham’s talk at the Language Summit aimed to provide a safe alternate “bootstrapping” path for the Python interpreter that forgoes running the same initialization from scratch each time and instead uses a “snapshot” of the Python process memory after initialization has completed. + +In tests of snapshots in [Pyodide](https://pyodide.org/), executing a simple “Hello, world” program was around 4 times faster (0.353 seconds versus 1.406 seconds) when loading the Python interpreter from a memory snapshot compared to loading the Python interpreter from scratch without a snapshot. Node.js implemented V8 snapshots in v4.2.4 and saw ~33% improvements to startup time, but shortly after, in v4.8.4, snapshots were disabled due to the problem that Hood wanted to raise for discussion. + + +## Randomization where you don’t expect + +The issue was one that was familiar to Python and many other languages: [denial-of-service through hash algorithm collisions](https://ocert.org/advisories/ocert-2011-003.html). This was a security issue reported to many programming languages at the same time, including Java, Ruby, PHP, and many other projects. + +Python fixed this issue in Python 3.3, as did many other programming languages, by adding a random [“salt” to `hash()` calculations](https://docs.python.org/3/reference/datamodel.html#object.__hash__). The salt is randomized at program startup to prevent remote attackers from controlling the hash values of inputs, and thus causing a CPU denial-of-service by creating extremely unbalanced data structures that rely on object hash values being evenly distributed, like [`dict`](https://docs.python.org/3/tutorial/datastructures.html#dictionaries) and [`set`](https://docs.python.org/3/tutorial/datastructures.html#sets). + +The danger of snapshotting and restoring the memory of a Python process post-initialization is captured by Randall Munroe’s [xkcd “Random Number”](https://xkcd.com/221/) (image below licensed CC-BY-NC 2.5), where the fixed value of “4” was at some indeterminate time in the past chosen by a fair dice roll and is now reused whenever a new random number is requested. + +![](image.png) + +If a Python process were to begin “restoring from a memory snapshot”, effectively this is what would happen for hash seed randomization: a value that is expected to be randomized every time would become deterministic and shared across all snapshots. Hood noted that there were likely other places, especially in libraries and programs, where this assumption would break existing use-cases, too. + +Hood proposed that the concept of an “initialization phase” be added to the Python language model, citing [RPython](https://rpython.readthedocs.io/en/latest/) and [SPy](https://github.com/spylang/spy) as two runtimes that have introduced an “explicit entrypoint to an initialization phase”. These entry points were useful for tree-shaking and pre-evaluation, respectively, but could be reused for the purpose of fixing issues related to adding memory snapshots to CPython. This initialization phase could be called again on re-initialization, allowing the runtime and third-party libraries to safely reintroduce randomness into the program if memory snapshotting were implemented. + + +## Discussion + +Stefan Behnel referenced some similar functionality that Python already provides: the [`atexit` module](https://docs.python.org/3/library/atexit.html), which provides an [API for registering a callback function](https://docs.python.org/3/library/atexit.html#atexit.register) that is called when the Python process is exiting. Stefan wondered whether a `reinit` callback function could serve the needs Hood described. + +Peter Bierma wondered whether improvements to initialization time could speed up startup enough that snapshotting would no longer be needed. Hood answered that there’s a “ton of work” to do when initializing Python, such as importing modules, and that a proposal to speed up initialization is “an order of magnitude more complex” than memory snapshotting and reinitialization. “I know how to implement snapshotting, I don’t know how to speed up the interpreter”. + +Thomas Wouters wondered whether [lazy imports, defined in PEP 810](https://peps.python.org/pep-0810/) and coming to Python 3.15, would solve some of the issues with slow initialization. Hood was cautious, stating that lazy imports might make his specific use-case “worse before it gets better”, especially if memory snapshots were implemented for Python. Lazy imports mean that imported modules aren’t parsed and executed until a member of the module is accessed, so if an import isn’t required for a particular execution of a Python program, then its cost isn’t paid by the interpreter. + +Snapshots, however, benefit performance because the cost to parse and execute all imports is paid a single time, the state is saved, and then it is “restored” in future executions to skip this part of the process. If imports aren’t resolved initially (such as when lazy imports are used), then the imports won’t be saved in the snapshot, and their cost will be paid on each Python execution, even if memory snapshots are implemented. + +Hood also shared some details about the use-case he was interested in: using Python with [WebAssembly](https://webassembly.org/) and edge compute services. Edge compute services don’t necessarily have access to a filesystem at runtime, which lazy imports would require. + +The conversation switched to [freezing modules](https://docs.python.org/3/using/cmdline.html#cmdoption-X). Pablo Galindo Salgado referenced some prior art at Alibaba for snapshots using a “memory dump and metadata table”. Guido van Rossum, speaking from experience on the Faster CPython team, noted that deep-freezing modules “didn’t affect startup that much” while adding complexity, and so “wasn’t worth it” from the team’s perspective. The team “didn’t spend much time trying to figure out why” this didn’t affect startup times. + +Eric Snow added that another approach considered was freezing the memory of objects and introducing a “form of copy-on-write (COW)”, acknowledging that many objects aren’t going to change, so copying data can be avoided. diff --git a/content/posts/language-summit-2026-namespaces/index.md b/content/posts/language-summit-2026-namespaces/index.md new file mode 100644 index 0000000..8cea959 --- /dev/null +++ b/content/posts/language-summit-2026-namespaces/index.md @@ -0,0 +1,126 @@ +--- +title: One namespace to namespace them all (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Pablo Galindo Salgado proposes a top-level `std` namespace for the Python standard library to prevent module shadowing and free up module names.' +tags: [language-summit, language-summit-2026] +published: true +--- + +Steering Council member and Release Manager Pablo Galindo Salgado opened the Language Summit this year to propose solutions to a problem that everyone who’s used Python has encountered at least once before. +The issue manifests as seemingly random `AttributeError` exceptions from modules like `json`, `math`, or `http` for names you know are correct. Why is the standard library suddenly raising errors? + +```terminaloutput +$ cat game.py +import random +if int(input("Pick a number 1-10")) == random.randint(1, 10): + print("You win!") + +$ python game.py +AttributeError: module 'random' has no attribute 'randint' (???) + +# Why is random.randint() not available??? +``` + +The issue is usually that a module is “shadowing” the standard library module from a higher-precedence location in `sys.path`, such as your current working directory or a directory on `PYTHONPATH`. There is probably a file named `json.py` or `math.py` in your project, and that module takes precedence over the standard library module of the same name, which is only noticed once you start using the module elsewhere in your project. + +```terminaloutput +$ ls +game.py random.py + +# random.py is being imported first, not the stdlib random + +$ python game.py +AttributeError: module 'random' has no attribute 'randint' +``` + +This issue is caused by Python’s flat module namespace; there is no delineation between what is part of the standard library and what are third-party modules, either dependencies or project code. + +Pablo noted that there is a low-cost option that solves this issue, available since Python 3.11: [the `-P` option](https://docs.python.org/3/using/cmdline.html#cmdoption-P), which, according to the documentation, “doesn’t prepend a potentially unsafe path to `sys.path`”, such as the current working directory. The problem is that most users are not running Python with this option enabled, or are depending on the current working directory being importable because they don’t install their own project code into the environment. + +## New standard library module names are constrained + +The problem goes beyond being confusing to users. Pablo explained that the flat namespace means that core developers often need to choose “awkward” names for new modules due to what’s already in use on the [Python Package Index (PyPI)](https://pypi.org/). This restriction applies even when a module is being adopted from PyPI into the standard library, such as PyYAML, which provides the `yaml` module. This is the reason that many new standard library modules have been suffixed with “lib” (such as `tomllib` or `graphlib`) or are unexpectedly named (such as `zoneinfo` instead of `timezone`). + +What is one potential solution? Adding a new top-level namespace +for all standard library modules. Pablo suggested the name `std` +for this namespace, so users who want to guarantee they are importing +a standard library module would `import std.json` or `from std import json`, for example. + +Pablo assured everyone that `import json` without the `std` namespace +would keep working basically forever (“We can’t break the world, that would be bad”). + +```python +>>> import std.json, json +>>> std.json is json +True +>>> json.__name__ +'json' # unchanged +``` + +But Pablo didn’t exclude the idea of introducing *new* modules *exclusively* +in the `std` namespace as a method for encouraging users to begin adopting +the new `std` namespace: + +```python +>>> import std.new_stdlib_module +# This would work for new modules. + +>>> import new_stdlib_module +ImportError +# But new modules without 'std.' wouldn't work... +``` + +Python wouldn’t be alone, either. Many other programming languages have +already namespaced their standard library (or equivalent): + +| Language | Namespace | Notes | +|----------|-----------|-------| +| Rust/C++ | `std::` | Reserved from day one | +| Java | `java.*` | Enforced at the classloader | +| Go | `"fmt"` | Standard library known by path shape | +| Node.js | `node:fs` | Retrofitted, unspoofable | +| Python | `import os` | **(⚠) Flat, unreserved** | + +## A tempting door opens: unbundling the standard library + +Pablo shared that a potential side effect of reserving a new top-level namespace +for the Python standard library is that the standard library +could then more easily be “unbundled”. Unbundling the standard library +would mean that the stdlib modules would be upgradeable separately +from the Python interpreter and independently of each other, likely +distributed through PyPI. + +Unbundling the standard library would provide a few benefits: + +* Ship standard library modules via PyPI, versioned independently of the interpreter. +* Security fixes out-of-band, not waiting for a Python release. +* Faster iteration on modules that move quicker than Python core. + +But there are trade-offs to unbundling, and we know of one concrete example. +Ruby “gemified” its standard library and ran into issues, such as long +deprecation periods (`require` took three release cycles), warnings from +transitive dependencies, and a CVE in the `uri` gem that went unpatched because the gem wasn’t +listed in [`Gemfile`](https://bundler.io/guides/gemfile.html), the Ruby equivalent of a package manifest. + +## Discussion + +Kushal Das shared that the shadowing issue was a “big problem for newcomers”, especially first-time Python users writing code doing arithmetic in a file named `math.py`. +Jukka Lehtosalo shared that he had also “personally encountered this problem”, and asked if there was “data about how often this happens”. Pablo didn’t have concrete data and shared that the error message has improved in recent Python versions. +Pablo didn’t want to over-focus on the shadowing issue and instead wanted to focus on what he believed was the larger issue: how the flat namespace affects how core developers choose standard library module names. + +David Hewitt wondered whether there is a “security edge” to this proposal, positing that core developers are “more familiar with what is in the standard library” compared to a beginner, and asked whether this change could help learners know what is included in Python and what isn’t. Pablo pushed back on the security angle: “the standard library is huge, we could do [standard library module] trivia and we’d all fail”. He concluded that this could be another positive reason to adopt the proposal, but he didn’t want to oversell this aspect, either. + +Guido van Rossum asked whether every stdlib module would eventually need to move under this proposal, which Pablo confirmed. As a follow-up, Guido asked whether this would mean touching imports across the entire standard library, which Pablo also confirmed, noting that this migration could be “mostly mechanical” and could include freezing all modules. + +Peter Bierma asked whether the new `std` namespace would add performance costs, to which Pablo answered that the cost would be “effectively zero”. + +Thomas Wouters imagined an incremental rollout of the new `std` namespace, proposing that new modules would land under the `std` namespace, with the possibility of a future mode that disables top-level shadowing entirely once enough of the ecosystem has moved. + +Stefan Behnel felt that Python “should provide a way to confidently import from the standard library”. He offered a potential solution that would keep most code the same: limiting the proposal to `from` imports. The syntax would be `from std import random`, so the module name would be the same and the `std` namespace wouldn’t appear in user code. This would also avoid the issue of [`sys.modules`](https://docs.python.org/3/library/sys.html#sys.modules) duplicating the module. Guido concurred, noting that `from std` could be magic, a “special keyword”, Pablo added. + +David Hewitt noted the precedent already set by the [`lazy` keyword](https://peps.python.org/pep-0810/) (new in Python 3.15) for changing the `import` statement without having to introduce a new module. +After a show of hands to get a temperature check on the idea of using a keyword, no one attending “hated the idea” and many attendees “loved the idea”. + +Gregory P. Smith provided guidance with his “PEP hat” on, noting that the proposal was potentially trying to solve multiple issues at once. “Keep all the ideas on the table, but be particular about what you want to tackle”. diff --git a/content/posts/language-summit-2026-pep-827-type-manipulation/index.md b/content/posts/language-summit-2026-pep-827-type-manipulation/index.md new file mode 100644 index 0000000..d254d93 --- /dev/null +++ b/content/posts/language-summit-2026-pep-827-type-manipulation/index.md @@ -0,0 +1,128 @@ +--- +title: 'PEP 827: Type Manipulation (Python Language Summit 2026)' +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Michael J. Sullivan presents PEP 827 and discusses a key design decision: how to store type annotations?' +tags: [language-summit, language-summit-2026] +published: true +--- + +Michael J. Sullivan came to the Language Summit to discuss [PEP 827](https://peps.python.org/pep-0827/), a PEP that includes many proposed improvements to Python type annotations. For the Python Language Summit, Michael wanted to focus in particular on one aspect: how type annotations are stored. + +Michael began by introducing PEP 827 “by looking at a completely different programming language”. He brought up an example from [Prisma](https://github.com/prisma/orm), an object-relational mapper (ORM) library written in [TypeScript](https://www.typescriptlang.org/docs/). Calling +`findMany` on a `User` model infers a return type based on the input parameters’ values, such as when selecting a subset of fields: + +```ts +const user = await prisma.user.findMany({ + select: { + name: true, + email: true, + posts: true, + }, +}); +``` + +This query derives a type that looks like this in TypeScript: + +```ts +{ + email: string; + name: string | null; + posts: { + id: number; + title: string; + content: string | null; + authorId: number | null; + }[]; +}[] +``` + +With PEP 827, Python’s type annotations could work similarly. Return types +could be inferred from a function’s input parameter types. + +Another use-case PEP 827 serves is generating new types programmatically, such as for CRUD (Create, Read, Update, Delete) web applications, where a single database table typically has a separate ORM model definition for each CRUD operation. Create would require all properties without defaults, Read would include all the public properties of a model, and Update would make all properties optional (supporting the `PATCH` HTTP method). Delete doesn’t typically need a model, as only the primary key is needed to delete a record. + +The example given in PEP 827 uses a `Hero` model, which has a `secret_name` and an optional `age` parameter, and creates three other models for the “public get”, create, and update operations: + +```python +class HeroBase(SQLModel): + name: str = Field(index=True) + age: int | None = Field(default=None, index=True) + +class Hero(HeroBase, table=True): + id: int | None = Field(default=None, primary_key=True) + secret_name: str + +class HeroPublic(HeroBase): + id: int + +class HeroCreate(HeroBase): + secret_name: str + +class HeroUpdate(HeroBase): + name: str | None = None + age: int | None = None + secret_name: str | None = None +``` + +Keeping all of these models and parameters up-to-date with the database model requires lots of diligence and can’t usually be done programmatically. So what if there was a way to instead contain the complexity within a type that transforms the overall model type for each of the CRUD operations? PEP 827 would allow writing Python types like the example below, containing the complexity in the `Public`, `Create`, and `Update` annotations: + +```python +class Hero(SQLModel, table=True): + id: int | None = Field(default=None, primary_key=True) + name: str = Field(index=True) + age: int | None = Field(default=None, index=True) + secret_name: str = Field(hidden=True) + +type HeroPublic = Public[Hero] +type HeroCreate = Create[Hero] +type HeroUpdate = Update[Hero] +``` + +## Type manipulation primitives in PEP 827 + +The core of the idea is introducing a handful of primitives to Python’s type system: + +* Conditional types (`true_type if bool_type else false_type`). +* Comprehension types (`*[t for t in Iter[iter_t]]`). Comprehension types can have conditional `if` clauses, just like regular comprehensions. +* Member types (`typing.Members[t]` would return an iterable of the type’s members). `member.name` and `member.type` would be properties of the member. + +The PEP was “inspired by TypeScript, but importantly [the PEP] is not modeled after TypeScript”. Michael noted that Python and JavaScript are “very different languages”, so the system ended up looking “quite a bit different”. The PEP was also designed so that no new keywords would be required to adopt it. + +PEP 827 contains many other proposed features that Michael didn’t have time to cover, including new operators, callable special construction, tuple slicing and length, iterating over unions, getting types of an attribute, inspecting initializers, and ways to handle `__init_subclass__` decorators. The PEP contains a [mypy-based prototype](https://github.com/msullivan/mypy-typemap/tree/typemap) for all the features being proposed. + +## How to store type annotations? + +Having introduced the PEP, Michael then returned to his original motivation: discussing how to store type annotations for runtime use. Popular libraries like [Pydantic](https://docs.pydantic.dev/) and [FastAPI](https://fastapi.tiangolo.com/) use type annotations at runtime to drive their behavior. “This was complicated by our desire to use Python-y syntax”, like `if`, `for`, and dot notation. +In concrete terms, Python will itself *evaluate* the types at runtime. In order to implement everything in PEP 827, there would need to be a way to access the *unevaluated* type even after Python had +executed the conditionals or iterations. As currently drafted, “the PEP is carefully designed to support an evaluator that works by calling the [`__annotate__()`](https://docs.python.org/3/reference/datamodel.html#object.__annotate__) functions”. + +There’s already a module, `annotationlib`, which has a “string format” ([`Format.STRING`](https://docs.python.org/3/library/annotationlib.html#annotationlib.Format.STRING)). In theory, if you call `annotationlib.get_annotations(..., format=Format.STRING)`, then you’d get the type annotation back unevaluated. “But this isn’t how this function is implemented today”; instead, the function does “dreadfully clever stuff” using proxy objects that overload every dunder method. + +Imogen has [drafted a PEP proposing a new “AST” format](https://discuss.python.org/t/draft-pep-more-expressive-type-expressions/108856) that is related to this problem, but Michael thought it would be smaller and more straightforward to make the `STRING` format work with conditional expressions. “There’s a surprisingly large design space here with trade-offs between time, space, and complexity”. + +One of the original ideas considered by the PEP that moved to `__annotate__()` functions was “can we just store the strings?” The most straightforward option would be for `__annotate__()` to support `STRING` in addition to `VALUE` and `VALUE_WITH_FAKE_GLOBALS`. The `__annotate__()` function would get a little bigger, but if `STRING` is passed as the format, it would return the strings “instead of requiring the library to do something complicated to get there”. + +This approach is fast, but uses more memory by representing annotations redundantly. The team considered whether to do the “even more obvious thing” and evaluated storing *only the strings*. But this isn’t enough; type annotations can refer to local variables, so evaluating them requires more context. “This is the whole reason we moved away from `from __future__ import annotations` ([PEP 563](https://peps.python.org/pep-0563/)), did [PEP 649](https://peps.python.org/pep-0649/), and added complexity to generating `__annotate__` functions”. + +Michael detailed a proposed alternate approach that saves memory at the expense of computation speed: “If we have the string for a type annotation, and we have the environment closure for `__annotate__`, then we can `eval()` the string in that closure and it will work. What would that look like? +Generate an `__annotate__` function that only implements the format for `STRING`, but is set up to only contain the free variables that the annotation used. If the type annotation refers to local +variables, the annotation will still depend on [the variables], [the type annotation] wouldn’t directly do anything with [the variables]. [The type annotation] would just build the dictionary with the strings. +We’d have a function in `annotationlib` that knows +how to pull [the variables] out, combine that with the strings, and evaluate them [both] properly. We could make `__annotate__()` call that [annotationlib] function when you ask for an evaluated type annotation”. +This approach would use a lot less memory than using strings, but would be much slower because you’d need to use [`eval()`](https://docs.python.org/3/library/functions.html#eval). + +## Discussion + +Stefan Behnel asked whether accepting PEP 827 would make the type system Turing-complete with the proposed conditionals and comprehensions. Michael replied that the Python type system “is already Turing-complete, but only *accidentally* Turing-complete”. Java had run into this issue with generics with subtyping bounds, and “Python does the same Generic stuff as Java”. “There was never an intention for the type system to be Turing-complete”, therefore accepting PEP 827 would only change the type system to be “intentionally Turing-complete”. + +Łukasz Langa commented on the aesthetics of the proposed syntax, pointing out the example of the comprehensions with “stars”. He cautioned core developers against seeing this syntax and having a knee-jerk negative reaction, such as thinking: “Oh [f-string], they’re going to make type hints look even worse”. Łukasz was clear that “...these features are for [frameworks] to implement internally [features] that allow type checkers to magically generate the correct types”. “Users are not meant to write types like these or only use them very rarely”. + +Łukasz also commented on the performance of using strings: when he first wrote `from __future__ import annotations`, one of the motivations was that strings are “[interned](https://docs.python.org/3/library/sys.html#sys.intern)”. Using string interning for types would take “less memory”. The problem was that the alphabet for interning strings doesn’t allow for square brackets (`[]`), which are used often in Python type annotations. There was a question whether this could be changed in Python, to which Łukasz replied that “we could, but then we ruin Python for everybody else” because this behavior has been “relied on for over 30 years”. “However…”, Łukasz continued, “what if `__annotate__()` interned [strings] all the time?” We expect many type annotations to be the same or similar, like `None`, `str`, and `str | None`. Łukasz suggested that this method “wouldn’t cost much memory”, which Michael agreed with. + +David Hewitt shared his experience working with TypeScript, where there are very complicated types. “When you’re a TypeScript user and you’re trying to figure out an API, you start clicking down through the API layers”, and you often hit the issue (“as Łukasz described”, with an accompanying “oh [f-string]”). David was concerned about whether this new syntax would be something that users would encounter frequently. + +Michael confessed that users clicking into “the select method” on an ORM would “see a wild type” and “there’s no getting around that”. “...but hopefully they’ll see the docstring first. If they look at the return type they’ll see something nice”. “I think there are cases where you’re doing `select(User)` you might be able to populate specific properties”, but this would be “potentially finicky” and dependent on the language server being used. + +“There will be moments where users discover that the world is complex”, Yury Selivanov said, acknowledging the need for complex type annotations. “Right now the world is not typed”. “The Python type system does not match the expressiveness of the language”. “Certainly some people might be confused”, he added, listing a few ways to mitigate this issue, such as better IDEs, language servers, and documentation. diff --git a/content/posts/language-summit-2026-rust-for-cpython/index.md b/content/posts/language-summit-2026-rust-for-cpython/index.md new file mode 100644 index 0000000..0bc1196 --- /dev/null +++ b/content/posts/language-summit-2026-rust-for-cpython/index.md @@ -0,0 +1,104 @@ +--- +title: Rust for CPython (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'David Hewitt shares a status update, first module, and potential acceptance criteria for the Rust for CPython project' +tags: [language-summit, language-summit-2026] +published: true +--- + +# Rust for CPython + +“No one said ‘don’t do this’ last year”. After [testing the waters at PyCon US 2025](https://pyfound.blogspot.com/2025/06/python-language-summit-2025-what-do-core-developers-want-from-rust.html), David Hewitt returned to the Python Language Summit asking what Python core developers want from Rust, along with proposed timelines, phases, and success criteria for how the Rust for CPython project might proceed and become a permanent fixture within the CPython project. + +David is acting as an “ambassador” for the Rust for CPython project team, which is currently led by core developers [Kirill Podoprigora](https://github.com/eclips4) and [Emma Smith](https://github.com/emmatyping) as [authors of the Rust for CPython PEP draft](https://discuss.python.org/t/pre-pep-rust-for-cpython/104906). Emma also [spoke at PyCon US 2026](https://www.youtube.com/watch?v=42kibVnUHYE) about the Rust for CPython project. The team itself is around 60 developers in a Discord channel, among them a “few [Python] core developers” and a “delegation from the Rust project”. The team has experience with previous projects integrating Rust into existing codebases, such as Android and the Linux kernel, and is “excited by the work and keen to support [the project] if we proceed”. + + +## Why Rust? + +![Graph showing the number of 'type-crash' issues increasing over time](type-crash.png) + +Showing how adopting Rust may specifically help CPython, David noted how the number of issues labeled with “[`type-crash`](https://github.com/python/cpython/issues?q=is%3Aissue%20state%3Aopen%20label%3Atype-crash)” has been steadily rising over time. “We’ve been making some big technical bets”, he said, referencing the [new parser](https://peps.python.org/pep-0617/), the [JIT](https://peps.python.org/pep-0744/), and [free-threading](https://peps.python.org/pep-0703/). These large, complex features may be one of the reasons more crash reports are being opened on GitHub, and Rust could be a potential solution here. David explained that Jeff Vander Stoep described Rust in Android as “move fast and fix things”, and that “fewer revisions for patches of the same size” was the experience Android has had since adopting Rust. + +Rust and Python are already working together in Python’s ecosystem of packages thanks to [PyO3](https://pyo3.rs) and [Maturin](https://pypi.org/project/maturin). David also emphasized that many technology companies were “choosing Rust as a bet” instead of only as the new shiny tool. + +But adopting Rust into CPython would not all be smooth sailing; there were still concerns that David and the team are aware of. + +One of the biggest concerns from a year ago was Rust’s lack of platform support compared to CPython, which at the time of writing [officially supports 20 different architectures and platforms](https://peps.python.org/pep-0011) at either Tier 1, 2, or 3. David shared that Rust’s support of different platforms has “widened since last year” and was becoming “less and less of a concern”. The proposal would be to ask Python distributors to attempt using the optional Rust support in Python 3.16 (October 2027) and report platform-specific issues upstream in time to be resolved around the Python 3.17 timeline (October 2028). + +David acknowledged that adopting Rust into a project is a social challenge as much as a technical challenge. Rust knowledge is “not universal amongst core developers”, which would be partially mitigated by “targeting small portions” of CPython and an incremental approach, as “1 million lines of C code can’t be ported all at once”. + +David was clear that using Rust in itself does not necessarily mean that ported code would be free of bugs or security issues. Although Rust does mitigate classes of issues that commonly create bugs, the code can still have correctness issues. The team proposes mitigating this by adopting property-based testing, fuzzing, and using “prudent engineering practices” during the porting process. + + +## The proposed first Rust module: zlib + +Below are the proposed timelines for the Rust for CPython project making a “significant improvement” to CPython, with a PEP defining the success criteria for Rust expected in late 2026. Under this timeline, the first Rust code to ship in Python would be in Python 3.16, where it would be completely optional, with the existing C code kept as a fallback. The earliest that Rust would become required to build CPython is Python 3.18 in 2029, at least three years away. + +!['Rust for CPython' proposed timeline](timeline.png) + +The timeline includes a build system and CI, Rust API proof-of-concept happening +in the Summer 2026, a PEP defining success criteria in late 2026, +an optional Rust backend for the zlib module and private Rust API in Python 3.16 (October 2027), +resolving platform issues and Rust in more places (json, xml, memoryview, parser) +for Python 3.17 (October 2028). Finally, in some distant Python version (October 2029+) +the Rust build would be made required and a public Rust API would be published. + +The Rust for CPython team has selected the [`zlib` module](https://docs.python.org/3/library/zlib.html) as the first module to be given an optional Rust implementation because they “wanted to achieve a significant improvement” with a “small scope”. The proposed Rust implementation will use [zlib-rs](https://crates.io/crates/zlib-rs), which is “heavily tested and used by the Firefox, uv, and [Cargo](https://doc.rust-lang.org/cargo/)” projects and “faster than zlib and zlib-ng on many platforms”. The module was also selected because it’d require using an external Cargo package, meaning this aspect of the build process would need to be designed and exercised. + +This small change would have an impact: the zlib compression algorithm is “widely used by Python packaging”, meaning that (almost) “every `pip install` in Python 3.16 will be sped up” if the proposal is accepted. + + +## Rust API Sketch + +David provided an example of some Rust code calling a hypothetical Rust API for Python. The API would use a Rust attribute (`#[pyfunction]`) and be similar to [Argument Clinic](https://devguide.python.org/development-tools/clinic/), a development tool for automatically generating blocks of code for handling C function arguments from Pythonic function syntax. Rust functions would always be passed the thread and interpreter state (`Python<'_>`), use smart pointers around objects (`Py<...>`), and lean into Rust error handling with the [`Result`](https://doc.rust-lang.org/std/result/) enum, returning either a result or an error that was raised. + +```rust +#[pyfunction(signature = ( + data, + /, + wbits=MAX_WBITS, + bufsize=DEF_BUF_SIZE, +))] +fn decompress( + py: Python<'_>, + data: Py, + wbits: c_int, + bufsize: isize +) -> PyResult, PyErrRaised> { + + // buf will be cleaned on scope exit + let buf = PyObject::get_buffer(py, &data)?; + let decoded = /* ... */; + + Ok(PyBytes::new(py, &decoded)) +} +``` + +The example above shows a buffer being allocated and automatically cleaned up on scope exit, rather than being cleaned up manually as would be required when writing the same function in C. + + +## Proposed success criteria + +David moved on to success criteria: what would the Rust for CPython project need to show to move into new phases of the roadmap and eventually become a required part of building CPython? “For previous big changes like the JIT and free-threading, we’ve explicitly defined success criteria that would need to be met in order for the added complexity to be accepted”, David explained, “we expect we’d need to do the same here. What should those criteria be?” + +David had these suggestions for core developers: + +* **Critical:** A majority of active core developers are *open* to using the Rust API to implement functionality. +* No meaningful slowdown to CPython performance benchmarks. +* All tiered platforms must be supported by Rust. The experience of distributors building CPython should generally indicate that adding Rust support is manageable. + +David noted that the first point reads “open”, not “familiar”, and shared a plan to survey Python core developers about how they’ve used Rust while contributing to CPython when deciding whether to move Rust out of experimental stages. + +## Discussion + +On the topic of designing the new Rust API so that it’s “familiar” to users of the C API, Thomas Wouters advised against “making compromises for the dinosaurs”, including himself in the subset, instead asking whether the Rust API should be designed from first principles. David Hewitt responded that there are “places to lean into Rust”, such as dropping resources on scope exit, but there are also idiomatic Rust designs which “won’t be the best fit”. David noted that it would be reasonable for core developers to be looking at both the C and Rust API at the same time while working. “We should be mindful of our audience, which is also ourselves”. + +Larry Hastings asked why the Rust for CPython project wasn’t a “rewrite”, suggesting the team “display your success as a fork”. David acknowledged that “[RustPython](https://github.com/RustPython/RustPython) already exists” and that the Rust for CPython team had already spoken with the contributors of the project. “RustPython isn’t as performant as CPython, but could be used to inform what APIs we design”. + +Larry also shared that he “wasn’t super excited to learn Rust to work on CPython”. David assured him that there are many areas of CPython that would not be considered for writing in Rust: “CPython should not be written in Rust for the sake of Rust”. “CPython will be a dual-language project for a meaningful amount of time”. However, David cautioned that “it would be disingenuous to say that Rust would be optional forever”, as one of the aforementioned roadmap items for the Rust for CPython project is to become a required part of the build process and provide a public Rust API. + +Pablo Galindo Salgado was more concerned about the future, which was “reaching for our dependencies from Cargo”, noting that this would be a “huge problem” and a potential “showstopper” for the project. “We vendor our dependencies, and we have a very selective set”, he said, noting that each time a vulnerability is published for one of those projects, the release managers need to make new releases, which can be “tiresome”. “Right now we’re only focusing on the APIs and the basics”, he added, highlighting that the challenge of taking on many Rust dependencies hasn’t been addressed yet. + +David Hewitt answered that the team should “select as few external dependencies as possible”; “zlib-rs is only one dependency”, and the Rust sources would be vendored so that “building CPython would not require Cargo”. David added that “Cargo has a relatively clean system for vendoring dependencies” and that the “Rust for CPython proof-of-concept uses this system”. The vendored sources “wouldn’t live in the CPython tree”. Łukasz Langa agreed that “it’s better for dependencies to live separately”, referencing the [cpython-source-deps repository](https://github.com/python/cpython-source-deps), with Thomas reminding everyone that this would be a new usage of that repository; today it is only used for binary installers of CPython. diff --git a/content/posts/language-summit-2026-rust-for-cpython/timeline.png b/content/posts/language-summit-2026-rust-for-cpython/timeline.png new file mode 100644 index 0000000..d8d0eb5 Binary files /dev/null and b/content/posts/language-summit-2026-rust-for-cpython/timeline.png differ diff --git a/content/posts/language-summit-2026-rust-for-cpython/type-crash.png b/content/posts/language-summit-2026-rust-for-cpython/type-crash.png new file mode 100644 index 0000000..c6d8fc9 Binary files /dev/null and b/content/posts/language-summit-2026-rust-for-cpython/type-crash.png differ diff --git a/content/posts/language-summit-2026-spicycrab/index.md b/content/posts/language-summit-2026-spicycrab/index.md new file mode 100644 index 0000000..69c1137 --- /dev/null +++ b/content/posts/language-summit-2026-spicycrab/index.md @@ -0,0 +1,53 @@ +--- +title: Spicycrab (Python Language Summit 2026) +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'Kushal Das shows off Spicycrab, a Python-to-Rust transpiler for Python users who need performance without learning Rust or leaving Python' +tags: [language-summit, language-summit-2026] +published: true +--- + +Kushal Das brought a project that fills a niche for Python users who hit a performance wall that can’t be solved by scaling horizontally, but who also don’t want to learn a new programming language. [Spicycrab](https://github.com/kushaldas/spicycrab/) is a Python-to-Rust transpiler named for Kushal’s love of spicy food. The demonstration showed compiling a simple Python script to Rust source code and then into an executable binary. + +```commandline +$ crabpy transpile greet.py -o greet +``` + +This simple demo produces Rust source code at `greet/src/main.rs`, which can be run with `cargo run`. + + +## Who is this for? + +After introducing the tool, Kushal shared more about who could benefit from a tool like Spicycrab. + +Back when Kushal was working with Django, there would often be a moment where the web service could no longer scale on the single virtual machine that was allocated to the team. At that time in 2008, the only option was to learn to write [C extensions](https://docs.python.org/3/extending/index.html) for the “hot paths”, but this approach meant you weren’t writing Python and you had to contend with crashes and security issues from writing C. + +Now that Rust is here, some of the challenges like security and crashing are addressed, but Rust’s “syntax breaks [his] brain” (although [PyO3](https://pyo3.rs/) helps). Kushal wanted to get more people taking advantage of Rust’s performance without needing to learn Rust. + +Kushal shared that many smaller organizations he knows in this exact resource-constrained situation would move away from a programming language like Python towards a language like Go for web backends. Kushal wanted to provide a tool that offers easy performance without needing to leave Python behind. “I hope that Python can become that fast one day, but what can we do until then?” + +Next, Kushal demonstrated [an async web service using actix-web](https://spicycrab.readthedocs.io/en/latest/actix_web.html), which includes [type annotations](https://docs.python.org/3/library/typing.html). This mechanism works by transpiling [actix-web](https://actix.rs/) and its dependencies to Python and then installing the resulting code as a Python package behind the scenes. + +Kushal installed actix-web using [Cargo](https://doc.rust-lang.org/cargo/), with a few chuckles as 180 dependencies were downloaded and compiled, a callback to some concerns from the “Rust for CPython” project regarding third-party dependencies. After installation and transpiling completed, the following web server was built using Spicycrab: + +```python +from spicycrab_actix_web import App, HttpServer, HttpResponse, get + +async def hello() -> HttpResponse: + return HttpResponse.Ok().body("Hello World!") + +async def main() -> None: + HttpServer.new(App.new().route("/", get().to(hello))).bind("127.0.0.1:8080").run() +``` + +After running the transpiler and compiling again, the resulting binary served the “Hello World!” response over HTTP. + +## Discussion + +Ken Jin asked why this approach would be chosen over [mypyc](https://mypyc.readthedocs.io/), Cython, or [SPy](https://github.com/spylang/spy). Kushal answered that +“SPy only supports a subset of Python” but admitted that “mypyc is good and in many cases that [mypyc] be used directly”. Kushal’s primary +motivation for going this route was to solve multiple problems at once, and one of the problems was getting +Python users to use Rust “without being scared of the syntax”. + +Gregory P. Smith provided a “spicier” thought that was “existential for Python”. Greg imagined it becoming more commonplace for state-of-the-art large language models to rewrite code from one programming language to another, such as from Python to Rust for performance. Kushal noted that this solves the issue “one time”, but doesn’t answer how the code is maintained long-term, and that the cost of doing so would be prohibitive for many. As an example, the [rewrite of Bun from Zig to Rust cost $165,000 USD](https://bun.com/blog/bun-in-rust). diff --git a/content/posts/language-summit-2026/index.md b/content/posts/language-summit-2026/index.md new file mode 100644 index 0000000..9f07a72 --- /dev/null +++ b/content/posts/language-summit-2026/index.md @@ -0,0 +1,45 @@ +--- +title: Python Language Summit 2026 +publishDate: '2026-09-30' +updatedDate: '2026-09-30' +author: Seth Larson +description: 'The 2026 Python Language Summit was hosted in Kraków, Poland before EuroPython 2026. There were 15 talks covering free-threading, Rust, garbage collection, type annotations, and more.' +tags: [language-summit, language-summit-2026] +published: true +--- + +On July 14th, 2026, 47 Python core developers and special guests sat down +at the [Python Language Summit](https://ep2026.europython.eu/language-summit/), this year +held in Kraków, Poland at EuroPython 2026, to discuss many topics about the future +of the Python programming language, including free-threading, Rust, garbage collectors, +type annotations, and namespacing. + +This marked the first time the Python Language +Summit had been hosted in Europe since 2011, when the event was +[held in Florence on June 19th, 2011](https://lwn.net/Articles/449710/). +Going forward, the Python Language Summit will alternate between +PyCon US and EuroPython on a yearly basis. + +The summit was organized by Emily Morehouse, Hugo van Kemenade, Lysandros Nikolaou, +and Łukasz Langa and blog posts were written by Seth Larson. + +Below are summaries of the 10 full-length +talks and 5 lightning talks that were presented at the 2026 Python Language Summit. +I hope you enjoy them, and thank you for your patience. + +* “[One namespace to namespace them all](/2026/09/language-summit-2026-namespaces)” by Pablo Galindo Salgado +* “[macOS and Python](/2026/09/language-summit-2026-macos-python)” by Ned Deily +* “[Garbage Collection: Generational? Incremental? Both!](/2026/09/language-summit-2026-garbage-collection-generational-incremental-both)” by Mark Shannon +* “[Memory Buffer Protocol](/2026/09/language-summit-2026-memory-buffer-protocol)” by Nathan Goldbaum +* “[Memory Snapshots for CPython](/2026/09/language-summit-2026-memory-snapshots)” by Hood Chatham +* “[Rust for CPython](/2026/09/language-summit-2026-rust-for-cpython)” by David Hewitt +* “[Spicycrab](/2026/09/language-summit-2026-spicycrab)” by Kushal Das +* “[Developer-in-Residence Update & Future](/2026/09/language-summit-2026-developer-in-residence-update-and-future)” by Petr Viktorin +* “[Free-Threaded Python Post-Era](/2026/09/language-summit-2026-free-threading-post-era)” by Donghee Na, Tobias Wrigstad, and Fridtjof Stoldt +* “[PEP 827: Type Manipulation](/2026/09/language-summit-2026-pep-827-type-manipulation)” by Michael J. Sullivan +* [Lightning Talks](/2026/09/language-summit-2026-lightning-talks) + * “One-time ABI breakage” by Mark Shannon + * “Safer and Generic Interruptions” by Daniele Parmeggiani + * “EktuPy, Scratch but Python” by Kushal Das + * “`AGENTS.md` for CPython” by Gregory P. Smith + * “Please read PEP 836 (JIT go brrr)” by Ken Jin