The 30-Year Evolution of Python's GIL
Throughout more than thirty years of Python history, the Global Interpreter Lock (GIL) has remained its most debated design choice. It kept the interpreter simple and single-threaded execution blisteringly fast, but it also locked CPU-bound multithreaded programs onto a single core for decades.
Since the GIL’s introduction in 1992, the community has attempted to remove it many times, but each effort stalled under unacceptable single-threaded performance penalties. It was not until PEP 703 proposed a radically redesigned memory architecture that Python introduced an experimental free-threaded build in 3.13, followed by official support in Python 3.14.
This overhaul is far more than simply “removing a lock”—it reaches deep into CPython’s low-level memory management and reshapes both Python’s concurrency model and its C extension ecosystem.
The 1992 Starting Point: Three Decades of Prosperity at the Price of a Lock
To understand the GIL, start with the core of CPython’s memory management model: Reference Counting. In CPython, every object tracks a counter of how many references currently point to it; once this count hits zero, the memory is freed immediately.
+----------------------------------------------------+
| CPython Process |
| |
| Thread 1 (Running) Thread 2 (Waiting) |
| +--------------------+ +--------------------+ |
| | Holds the GIL | | Blocked on GIL | |
| | Executes Bytecode | | Cannot Execute | |
| +--------------------+ +--------------------+ |
| | |
| v |
| +----------------------------------------------+ |
| | Shared Object Memory | |
| | [Object A: refcount] [Object B: refcount] | |
| +----------------------------------------------+ |
+----------------------------------------------------+
In a multithreaded environment, if two threads mutate the same object’s refcount concurrently, they risk a race condition that could deallocate memory prematurely (causing a crash) or never deallocate it (causing a memory leak).
Why the GIL Was the Sensible Choice at the Time
To protect these counters, an interpreter has two primary choices:
- Fine-grained Locks: Allocate a separate mutex for every single object, acquiring and releasing it on every reference increment and decrement.
- Coarse-grained Global Lock (GIL): Place a single mutex over the entire interpreter, allowing only one thread to execute bytecode at any given moment.
On August 4, 1992, Guido van Rossum introduced threadmodule.c in commit 1984f1e and added interpreter_lock inside ceval.c. In an era dominated by single-core CPUs, a global lock kept single-threaded execution fast at a very low runtime cost while drastically lowering the barrier to writing C extension modules.
This design lowered the barrier to developing C extensions like NumPy and helped the ecosystem flourish—but once multi-core processors became mainstream, it also became the limitation for CPU-bound multithreaded programs.
Debunking the Myth: The GIL Protects the Interpreter, Not Your Code
Many developers mistakenly believe that “because Python has the GIL, multithreaded code is automatically thread-safe.” In reality, the GIL only protects CPython’s internal state and memory consistency, offering zero protection for user application logic.
Take a common counter += 1 statement. At the bytecode level, it unpacks into three distinct operations: read the value, increment it, and write it back. If a thread switch occurs between these steps, shared updates get clobbered.
import sys
import threading
counter = 0
def work(n):
global counter
for _ in range(n):
counter += 1 # Read -> Increment -> Write back (not atomic)
N, THREADS = 1_000_000, 4
threads = [threading.Thread(target=work, args=(N,)) for _ in range(THREADS)]
for t in threads:
t.start()
for t in threads:
t.join()
print(f"Expected: {N * THREADS:,}, Actual: {counter:,}")
Running this benchmark across different runtimes reveals stark behavioral differences:
- Python 3.9 (GIL Enabled): four threads performing 4 million increments in total yield only ~1.4M to 2.6M, losing more than half of the updates.
- Python 3.14 (GIL Enabled): The result happens to hit exactly 4,000,000, because commit 4958f5d in 3.10 changed when thread switches are checked—only after function calls and at loop back-edges—so a plain loop never triggers a switch point. This is a CPython implementation detail, not a language guarantee; replace the increment with a function call and the race condition comes right back.
- Python 3.14t (Free-threaded, GIL Disabled): With true multicore parallelism, the counter drops to ~1.1M, losing nearly 70% of the data.
This empirical test shatters the illusion of GIL-provided safety: shared mutable state across threads must always be synchronized with threading.Lock or queue.Queue.
Why the GIL Barely Affects I/O Work
The GIL usually has little impact on I/O-bound work. Such work spends most of its time waiting on the network, disk, or a database rather than executing Python code. When a thread enters an I/O wait or time.sleep(), the interpreter releases the GIL and lets another thread take over. Multiple threads still cannot execute Python bytecode at the same time, but they can overlap one another’s waiting time.
Pure-Python CPU-bound work is the opposite: it spends most of its time executing bytecode, so threads can only take turns holding the GIL and never put multiple cores to use. Some C extensions are the exception—they release the GIL themselves while computing.
A micro-benchmark on an Apple M4 Max (16 cores) running macOS and Python 3.14.7 reflects this split: four I/O tasks each waiting 0.5 seconds dropped from 2.0 seconds sequentially to 0.51 seconds on four threads. Pure CPU-bound computation saw almost no speedup with the GIL enabled (0.33 s sequential vs 0.32 s on four threads); with the GIL disabled, runtime fell from 0.26 s to 0.06 s.
async Solves a Different Problem Than the GIL
The GIL and async operate at different layers:
- The GIL limits: within a single CPython process, multiple threads usually cannot execute Python bytecode at the same time.
asyncsolves: when huge numbers of tasks are all waiting on the network, a database, or files, how to manage that concurrency efficiently with a small number of threads.
Traditional threads already release the GIL while waiting on I/O. The value of async is avoiding one thread per task: asyncio switches between coroutines on an event loop, which typically brings:
- Lower memory and scheduling overhead.
- Easier support for very large numbers of connections.
- More explicit cancellation, timeout, and flow control.
Even if Python removes the GIL entirely in the future, async will remain valuable. It cannot solve CPU-bound computation—running heavy computation directly on the event loop blocks every other coroutine. That kind of work still belongs in multiple processes, in GIL-releasing C extensions, or in free-threaded Python.
NOTE
Further reading: The 20-Year Evolution of Python Async
Thirty Years of Siege: The Cost of Past No-GIL Attempts
Starting in the 1990s, the community launched several attempts to remove the GIL, but all foundered on severe performance penalties:
“I’d welcome a set of patches into Py3k only if the performance for a single-threaded program (and for a multi-threaded but I/O-bound program) does not decrease.” — Guido van Rossum (2007)
| Year / Version | Project & Lead | Architecture | Outcome & Performance Impact |
|---|---|---|---|
| 1996 / 1999 | Greg Stein’s Patch | Fine-grained locks on interpreter containers and refcounts | Single-threaded performance slowed by 2x to 4x; rejected |
| 2007 | Guido’s published criterion | Defined prerequisites for removing the GIL | Established the zero-regression rule for single-threaded code |
| 2011 (3.2) | Antoine Pitrou (New GIL) | Replaced instruction count checks with fixed time intervals (5ms) | Reduced thread contention overhead, but retained the GIL itself |
| 2016 | Larry Hastings (Gilectomy) | Fine-grained locking refactor for CPython 3.x | Atomic memory operations caused severe single-thread regression; stalled |
| 2017 | PEP 554 / Eric Snow | Explored subinterpreter isolation model | Laid groundwork for per-interpreter GIL and multiple interpreter paths |
| 2021 | Sam Gross (nogil) | Biased refcounting, immortal objects, and mimalloc integration | Reduced single-thread overhead to 5–6%; became the PEP 703 prototype |
Why Removing the GIL Is So Expensive
The common bottleneck across decades of attempts is the hardware cost of atomic operations and cache-line bouncing.
In Python, even the simplest variable assignment or attribute access mutates a refcount. Without a global lock, turning every refcount update into an atomic instruction synchronized across CPU cores makes each core’s cache invalidate constantly, causing severe single-threaded slowdown.
The Breakthrough: The Five Pillars of PEP 703
In 2021, Sam Gross, then at Meta, published the nogil fork; in 2023, Meta committed three engineer-years to help land it. PEP 703 succeeded not by blindly turning every counter into an atomic operation, but by orchestrating a defense-in-depth suite of memory innovations:
PEP 703: Defense-in-Depth Memory Architecture
+------------------------------------------------------+
| 1. Biased Reference Counting |
| Local Thread -> Plain add/sub (Fast Path) |
| Shared Thread -> Atomic add/sub (Slow Path) |
+------------------------------------------------------+
| 2. Immortal Objects (PEP 683) |
| None/True/Small Ints -> Refcount Never Changes |
+------------------------------------------------------+
| 3. Deferred Reference Counting |
| Functions/Modules/Code -> Counted during GC |
+------------------------------------------------------+
| 4. mimalloc Allocator + Per-Object Critical Sections |
| Thread-Safe Alloc + Fine-Grained Object Locks |
+------------------------------------------------------+
These five pillars form a layered filtering hierarchy:
- Biased Reference Counting: Most Python objects are only ever accessed by the thread that created them. Reference counts are split into “owner-local” and “shared” fields. The owner thread executes plain, non-atomic increments/decrements (the fast path), and only other threads fall back to atomic instructions.
- Immortal Objects (PEP 683): Pervasively shared constants (
None,True, small integers, interned strings) have their refcounts marked with a fixed bit pattern. Increments and decrements on these objects become no-ops, completely eliminating cache contention for common values. - Deferred Reference Counting: Long-lived, frequently accessed objects (functions, modules, code objects) skip immediate refcount updates during execution and are instead reconciled in one pass during garbage collection (GC).
- mimalloc Allocator: The thread-safe mimalloc allocator lets threads allocate memory without contending for a global lock, and lets built-in data structures like
dictandlistbe read without locking. - Critical Sections (Per-Object Locks): Fine-grained critical sections guard individual container objects; when a thread hits a potentially blocking call, the object locks are temporarily released to prevent deadlocks.
As recorded in the official Python 3.14 documentation, the average pyperformance overhead of the free-threaded build comes to about 1% on macOS aarch64 and about 8% on x86-64 Linux.
Dual Tracks: Free-Threading vs. Subinterpreters
Besides removing the GIL outright, CPython has been advancing a second architectural track at the same time: Subinterpreters.
- Subinterpreters (PEP 684 / PEP 734): Runs multiple mutually isolated interpreter instances inside a single OS process, each with its own GIL. This model combines isolation of Python state with resource costs close to threads; data is passed mainly by copying or through cross-interpreter queues—suited to architectures whose tasks are independent and that want to strictly limit shared state.
- Free-Threading (PEP 703): Removes the GIL from the interpreter entirely, letting all threads share the same memory address space. It delivers true native parallelism for workloads that need to share large amounts of memory tightly, such as machine-learning matrices and graph algorithms.
How Subinterpreters Differ from subprocess
Both isolate Python state and both sidestep the single GIL; the difference is the isolation boundary. A subinterpreter is an independent Python runtime inside the same OS process, while a subprocess is a complete OS process of its own.
| Aspect | Subinterpreter | subprocess |
|---|---|---|
| Execution boundary | Lives in the same process as the main program | A separate OS process |
| Startup & memory cost | Usually lower | Usually higher |
| Resource isolation | Python state is separated, but some process resources are still shared | Each process has its own address space; resources must be explicitly inherited or passed |
| Failure blast radius | A low-level error can take down the whole process | A child process crash usually leaves the parent unaffected |
| Package compatibility | Some C extensions are not yet supported | Generally more mature |
A subinterpreter isolates only Python state—it is not a security boundary, and it cannot directly share arbitrary Python objects. If the goal is to run Python functions in parallel, the closer comparison is InterpreterPoolExecutor versus ProcessPoolExecutor: the former hosts multiple interpreters on threads, while the latter uses multiple processes.
CAUTION
A subinterpreter is not a security sandbox: it isolates only Python object state, OS-level resources remain shared, and low-level faults can still take down the entire process. Never run untrusted code inside a subinterpreter assuming it is contained.
The Three-Phase Roadmap for Free-Threading
The Steering Council advances free-threading in three phases:
Free-Threading Roadmap
[ Phase I: Python 3.13 ]
- Experimental Build (python3.13t)
- Specializing interpreter disabled (~40% overhead)
[ Phase II: Python 3.14 - Current State ]
- Officially Supported Build (python3.14t)
- Single-thread overhead reduced to 5-10%
- Default download STILL HAS the GIL
[ Phase II Extension: Python 3.15 (2026-10) ]
- Introduces abi3t for stable C extension binaries
[ Phase III: Future (Unscheduled) ]
- Free-threading becomes the DEFAULT build
- Decision depends on ecosystem maturity & community vote
Phase I: Experimental Builds in Python 3.13
Python 3.13 shipped the first experimental python3.13t build. The specializing adaptive interpreter was not yet enabled at that point, and the average pyperformance overhead came to about 40%.
Phase II: Official Support in Python 3.14
Python 3.14 is currently in Phase II (officially supported). The official installers offer a python3.14t binary as an opt-in choice, but the default Python download still ships with the GIL.
Phase II Extension: ABI3T in Python 3.15
C extension compatibility is the biggest challenge on the road to no-GIL. If a module does not explicitly declare free-threading support, the interpreter prints a warning at load time and automatically re-enables the GIL.
WARNING
A service may appear to run on a free-threaded build while still operating in GIL mode: an undeclared C extension triggers a warning at load time and re-enables the GIL, so multi-core parallelism never engages. Use sys._is_gil_enabled() to confirm the GIL is actually disabled.
Python 3.15, scheduled for release in October 2026, will formally introduce the abi3t stable ABI (the binary interface between C extensions and the interpreter; PEP 803), letting C extension authors compile a single wheel (tagged abi3.abi3t) that runs on both the standard and free-threaded builds.
Phase III: Becoming the Default
“Any decision to transition to Phase III, with free-threading as the default or sole build of Python is still undecided, and dependent on many factors both within CPython itself and the community.” — Python Steering Council
The officials have even kept a clause allowing the effort to be halted or withdrawn entirely if the transition proves too disruptive to the ecosystem. The community expects to keep watching how quickly third-party packages add compatible support before deciding whether to retire the GIL for good.
NOTE
Further reading: Django’s Async Evolution: Six Years of Overhaul, Architectural Bottlenecks, and the No-GIL Reversal
Which Python Parallelism Approach Should You Choose?
When choosing an approach, first ask whether the work is waiting on I/O, burning CPU, or sharing large amounts of memory:
| Concurrency Model | Best Suited For | Key Constraints & Tradeoffs | Minimum Supported Version |
|---|---|---|---|
| threading (with GIL) | I/O-bound tasks (API calls, DB queries, file I/O) | Cannot parallelize pure Python CPU work; requires manual locking | Any Python 3.x |
| asyncio | High-concurrency I/O services (Web servers, API gateways) | Requires async ecosystem; CPU compute blocks the event loop | Python 3.4+ |
| multiprocessing | Stable CPU-bound workloads in production today | Process startup and pickle serialization costs; no shared memory | Python 2.6+ |
| C Extensions (NumPy, etc.) | Matrix math, image processing, cryptography | Only low-level C code releases the GIL; object arrays do not | Package dependent |
| Subinterpreters | Parallel workloads that want to strictly limit shared state | Data must be copied across interpreters; some C extensions unsupported | Python 3.14 (Official Python API) |
| Free-Threading | CPU-bound parallel workloads with shared in-memory data | Slight single-threaded slowdown; all C extensions must support it | Python 3.14 (Official Support) |
Practical Migration Guide for Developers
For teams planning to try or migrate to free-threaded Python, the following practices help:
- Verify Interpreter Capabilities: Use
sysconfig.get_config_var("Py_GIL_DISABLED")in code to determine whether the binary was built for free-threading, andsys._is_gil_enabled()to check whether the GIL is currently inactive (ensuring no extension forced it back on). - Avoid Premature Overhauls: Existing I/O-bound applications should continue using
asyncioorthreading. For production CPU-bound tasks in the 3.14 era,ProcessPoolExecutorremains the safest default. - Use
PYTHON_GILfor Root-Cause Triage: When encountering unexpected crashes or data anomalies in a free-threaded environment, setPYTHON_GIL=1and re-run. If the issue disappears, you are chasing a race condition in your application logic. - Add
3.14tto CI Test Matrices: Installpython3.14tand run your existing unit tests on it to surface thread-safety bugs hiding in your application logic early.
Removing the GIL was never about chasing multi-threading benchmarks; it is a thirty-year engineering balance between preserving single-threaded speed and unlocking multicore hardware. With PEP 703 officially supported in 3.14 and abi3t arriving in 3.15, the most pragmatic move today is to add 3.14t validation to CI and prepare for the era of native multicore Python.