The 30-Year Evolution of Python's GIL

Throughout more than thirty years of Python history, the Global Interpreter Lock (GIL) has remained its most debated design choice. It kept the interpreter simple and single-threaded execution blisteringly fast, but it also locked CPU-bound multithreaded programs onto a single core for decades.

Since the GIL’s introduction in 1992, the community has attempted to remove it many times, but each effort stalled under unacceptable single-threaded performance penalties. It was not until PEP 703 proposed a radically redesigned memory architecture that Python introduced an experimental free-threaded build in 3.13, followed by official support in Python 3.14.

This overhaul is far more than simply “removing a lock”—it reaches deep into CPython’s low-level memory management and reshapes both Python’s concurrency model and its C extension ecosystem.


The 1992 Starting Point: Three Decades of Prosperity at the Price of a Lock

To understand the GIL, start with the core of CPython’s memory management model: Reference Counting. In CPython, every object tracks a counter of how many references currently point to it; once this count hits zero, the memory is freed immediately.

 +----------------------------------------------------+
 |                  CPython Process                   |
 |                                                    |
 |  Thread 1 (Running)       Thread 2 (Waiting)       |
 |  +--------------------+    +--------------------+  |
 |  | Holds the GIL      |    | Blocked on GIL     |  |
 |  | Executes Bytecode  |    | Cannot Execute     |  |
 |  +--------------------+    +--------------------+  |
 |           |                                        |
 |           v                                        |
 |  +----------------------------------------------+  |
 |  |            Shared Object Memory              |  |
 |  | [Object A: refcount]   [Object B: refcount]  |  |
 |  +----------------------------------------------+  |
 +----------------------------------------------------+

In a multithreaded environment, if two threads mutate the same object’s refcount concurrently, they risk a race condition that could deallocate memory prematurely (causing a crash) or never deallocate it (causing a memory leak).

Why the GIL Was the Sensible Choice at the Time

To protect these counters, an interpreter has two primary choices:

  1. Fine-grained Locks: Allocate a separate mutex for every single object, acquiring and releasing it on every reference increment and decrement.
  2. Coarse-grained Global Lock (GIL): Place a single mutex over the entire interpreter, allowing only one thread to execute bytecode at any given moment.

On August 4, 1992, Guido van Rossum introduced threadmodule.c in commit 1984f1e and added interpreter_lock inside ceval.c. In an era dominated by single-core CPUs, a global lock kept single-threaded execution fast at a very low runtime cost while drastically lowering the barrier to writing C extension modules.

This design lowered the barrier to developing C extensions like NumPy and helped the ecosystem flourish—but once multi-core processors became mainstream, it also became the limitation for CPU-bound multithreaded programs.


Debunking the Myth: The GIL Protects the Interpreter, Not Your Code

Many developers mistakenly believe that “because Python has the GIL, multithreaded code is automatically thread-safe.” In reality, the GIL only protects CPython’s internal state and memory consistency, offering zero protection for user application logic.

Take a common counter += 1 statement. At the bytecode level, it unpacks into three distinct operations: read the value, increment it, and write it back. If a thread switch occurs between these steps, shared updates get clobbered.

import sys
import threading

counter = 0

def work(n):
    global counter
    for _ in range(n):
        counter += 1  # Read -> Increment -> Write back (not atomic)

N, THREADS = 1_000_000, 4
threads = [threading.Thread(target=work, args=(N,)) for _ in range(THREADS)]
for t in threads:
    t.start()
for t in threads:
    t.join()

print(f"Expected: {N * THREADS:,}, Actual: {counter:,}")

Running this benchmark across different runtimes reveals stark behavioral differences:

This empirical test shatters the illusion of GIL-provided safety: shared mutable state across threads must always be synchronized with threading.Lock or queue.Queue.

Why the GIL Barely Affects I/O Work

The GIL usually has little impact on I/O-bound work. Such work spends most of its time waiting on the network, disk, or a database rather than executing Python code. When a thread enters an I/O wait or time.sleep(), the interpreter releases the GIL and lets another thread take over. Multiple threads still cannot execute Python bytecode at the same time, but they can overlap one another’s waiting time.

Pure-Python CPU-bound work is the opposite: it spends most of its time executing bytecode, so threads can only take turns holding the GIL and never put multiple cores to use. Some C extensions are the exception—they release the GIL themselves while computing.

A micro-benchmark on an Apple M4 Max (16 cores) running macOS and Python 3.14.7 reflects this split: four I/O tasks each waiting 0.5 seconds dropped from 2.0 seconds sequentially to 0.51 seconds on four threads. Pure CPU-bound computation saw almost no speedup with the GIL enabled (0.33 s sequential vs 0.32 s on four threads); with the GIL disabled, runtime fell from 0.26 s to 0.06 s.

async Solves a Different Problem Than the GIL

The GIL and async operate at different layers:

Traditional threads already release the GIL while waiting on I/O. The value of async is avoiding one thread per task: asyncio switches between coroutines on an event loop, which typically brings:

Even if Python removes the GIL entirely in the future, async will remain valuable. It cannot solve CPU-bound computation—running heavy computation directly on the event loop blocks every other coroutine. That kind of work still belongs in multiple processes, in GIL-releasing C extensions, or in free-threaded Python.


Thirty Years of Siege: The Cost of Past No-GIL Attempts

Starting in the 1990s, the community launched several attempts to remove the GIL, but all foundered on severe performance penalties:

“I’d welcome a set of patches into Py3k only if the performance for a single-threaded program (and for a multi-threaded but I/O-bound program) does not decrease.” — Guido van Rossum (2007)

Year / VersionProject & LeadArchitectureOutcome & Performance Impact
1996 / 1999Greg Stein’s PatchFine-grained locks on interpreter containers and refcountsSingle-threaded performance slowed by 2x to 4x; rejected
2007Guido’s published criterionDefined prerequisites for removing the GILEstablished the zero-regression rule for single-threaded code
2011 (3.2)Antoine Pitrou (New GIL)Replaced instruction count checks with fixed time intervals (5ms)Reduced thread contention overhead, but retained the GIL itself
2016Larry Hastings (Gilectomy)Fine-grained locking refactor for CPython 3.xAtomic memory operations caused severe single-thread regression; stalled
2017PEP 554 / Eric SnowExplored subinterpreter isolation modelLaid groundwork for per-interpreter GIL and multiple interpreter paths
2021Sam Gross (nogil)Biased refcounting, immortal objects, and mimalloc integrationReduced single-thread overhead to 5–6%; became the PEP 703 prototype

Why Removing the GIL Is So Expensive

The common bottleneck across decades of attempts is the hardware cost of atomic operations and cache-line bouncing.

In Python, even the simplest variable assignment or attribute access mutates a refcount. Without a global lock, turning every refcount update into an atomic instruction synchronized across CPU cores makes each core’s cache invalidate constantly, causing severe single-threaded slowdown.


The Breakthrough: The Five Pillars of PEP 703

In 2021, Sam Gross, then at Meta, published the nogil fork; in 2023, Meta committed three engineer-years to help land it. PEP 703 succeeded not by blindly turning every counter into an atomic operation, but by orchestrating a defense-in-depth suite of memory innovations:

 PEP 703: Defense-in-Depth Memory Architecture

 +------------------------------------------------------+
 | 1. Biased Reference Counting                         |
 |    Local Thread -> Plain add/sub (Fast Path)         |
 |    Shared Thread -> Atomic add/sub (Slow Path)       |
 +------------------------------------------------------+
 | 2. Immortal Objects (PEP 683)                        |
 |    None/True/Small Ints -> Refcount Never Changes    |
 +------------------------------------------------------+
 | 3. Deferred Reference Counting                       |
 |    Functions/Modules/Code -> Counted during GC       |
 +------------------------------------------------------+
 | 4. mimalloc Allocator + Per-Object Critical Sections |
 |    Thread-Safe Alloc + Fine-Grained Object Locks     |
 +------------------------------------------------------+

These five pillars form a layered filtering hierarchy:

As recorded in the official Python 3.14 documentation, the average pyperformance overhead of the free-threaded build comes to about 1% on macOS aarch64 and about 8% on x86-64 Linux.


Dual Tracks: Free-Threading vs. Subinterpreters

Besides removing the GIL outright, CPython has been advancing a second architectural track at the same time: Subinterpreters.

How Subinterpreters Differ from subprocess

Both isolate Python state and both sidestep the single GIL; the difference is the isolation boundary. A subinterpreter is an independent Python runtime inside the same OS process, while a subprocess is a complete OS process of its own.

AspectSubinterpretersubprocess
Execution boundaryLives in the same process as the main programA separate OS process
Startup & memory costUsually lowerUsually higher
Resource isolationPython state is separated, but some process resources are still sharedEach process has its own address space; resources must be explicitly inherited or passed
Failure blast radiusA low-level error can take down the whole processA child process crash usually leaves the parent unaffected
Package compatibilitySome C extensions are not yet supportedGenerally more mature

A subinterpreter isolates only Python state—it is not a security boundary, and it cannot directly share arbitrary Python objects. If the goal is to run Python functions in parallel, the closer comparison is InterpreterPoolExecutor versus ProcessPoolExecutor: the former hosts multiple interpreters on threads, while the latter uses multiple processes.

CAUTION

A subinterpreter is not a security sandbox: it isolates only Python object state, OS-level resources remain shared, and low-level faults can still take down the entire process. Never run untrusted code inside a subinterpreter assuming it is contained.


The Three-Phase Roadmap for Free-Threading

The Steering Council advances free-threading in three phases:

 Free-Threading Roadmap

 [ Phase I: Python 3.13 ]
 - Experimental Build (python3.13t)
 - Specializing interpreter disabled (~40% overhead)

 [ Phase II: Python 3.14 - Current State ]
 - Officially Supported Build (python3.14t)
 - Single-thread overhead reduced to 5-10%
 - Default download STILL HAS the GIL

 [ Phase II Extension: Python 3.15 (2026-10) ]
 - Introduces abi3t for stable C extension binaries

 [ Phase III: Future (Unscheduled) ]
 - Free-threading becomes the DEFAULT build
 - Decision depends on ecosystem maturity & community vote

Phase I: Experimental Builds in Python 3.13

Python 3.13 shipped the first experimental python3.13t build. The specializing adaptive interpreter was not yet enabled at that point, and the average pyperformance overhead came to about 40%.

Phase II: Official Support in Python 3.14

Python 3.14 is currently in Phase II (officially supported). The official installers offer a python3.14t binary as an opt-in choice, but the default Python download still ships with the GIL.

Phase II Extension: ABI3T in Python 3.15

C extension compatibility is the biggest challenge on the road to no-GIL. If a module does not explicitly declare free-threading support, the interpreter prints a warning at load time and automatically re-enables the GIL.

WARNING

A service may appear to run on a free-threaded build while still operating in GIL mode: an undeclared C extension triggers a warning at load time and re-enables the GIL, so multi-core parallelism never engages. Use sys._is_gil_enabled() to confirm the GIL is actually disabled.

Python 3.15, scheduled for release in October 2026, will formally introduce the abi3t stable ABI (the binary interface between C extensions and the interpreter; PEP 803), letting C extension authors compile a single wheel (tagged abi3.abi3t) that runs on both the standard and free-threaded builds.

Phase III: Becoming the Default

“Any decision to transition to Phase III, with free-threading as the default or sole build of Python is still undecided, and dependent on many factors both within CPython itself and the community.” — Python Steering Council

The officials have even kept a clause allowing the effort to be halted or withdrawn entirely if the transition proves too disruptive to the ecosystem. The community expects to keep watching how quickly third-party packages add compatible support before deciding whether to retire the GIL for good.


Which Python Parallelism Approach Should You Choose?

When choosing an approach, first ask whether the work is waiting on I/O, burning CPU, or sharing large amounts of memory:

Concurrency ModelBest Suited ForKey Constraints & TradeoffsMinimum Supported Version
threading (with GIL)I/O-bound tasks (API calls, DB queries, file I/O)Cannot parallelize pure Python CPU work; requires manual lockingAny Python 3.x
asyncioHigh-concurrency I/O services (Web servers, API gateways)Requires async ecosystem; CPU compute blocks the event loopPython 3.4+
multiprocessingStable CPU-bound workloads in production todayProcess startup and pickle serialization costs; no shared memoryPython 2.6+
C Extensions (NumPy, etc.)Matrix math, image processing, cryptographyOnly low-level C code releases the GIL; object arrays do notPackage dependent
SubinterpretersParallel workloads that want to strictly limit shared stateData must be copied across interpreters; some C extensions unsupportedPython 3.14 (Official Python API)
Free-ThreadingCPU-bound parallel workloads with shared in-memory dataSlight single-threaded slowdown; all C extensions must support itPython 3.14 (Official Support)

Practical Migration Guide for Developers

For teams planning to try or migrate to free-threaded Python, the following practices help:

  1. Verify Interpreter Capabilities: Use sysconfig.get_config_var("Py_GIL_DISABLED") in code to determine whether the binary was built for free-threading, and sys._is_gil_enabled() to check whether the GIL is currently inactive (ensuring no extension forced it back on).
  2. Avoid Premature Overhauls: Existing I/O-bound applications should continue using asyncio or threading. For production CPU-bound tasks in the 3.14 era, ProcessPoolExecutor remains the safest default.
  3. Use PYTHON_GIL for Root-Cause Triage: When encountering unexpected crashes or data anomalies in a free-threaded environment, set PYTHON_GIL=1 and re-run. If the issue disappears, you are chasing a race condition in your application logic.
  4. Add 3.14t to CI Test Matrices: Install python3.14t and run your existing unit tests on it to surface thread-safety bugs hiding in your application logic early.

Removing the GIL was never about chasing multi-threading benchmarks; it is a thirty-year engineering balance between preserving single-threaded speed and unlocking multicore hardware. With PEP 703 officially supported in 3.14 and abi3t arriving in 3.15, the most pragmatic move today is to add 3.14t validation to CI and prepare for the era of native multicore Python.