P32 · Concurrency & Parallelism
threads, processes, the GIL & synchronisation
Make Python do many things at once - correctly. Meet the GIL and learn when to reach for threads, when for processes, and how Locks, Queues and Semaphores tame the data races that shared mutable state invites.
Take a function that spends its time computing, run it in four threads, and measure. It will not be four times faster. It will usually be slightly slower than the single-threaded version, and if you have never met the global interpreter lock this result is baffling enough to send people to Stack Overflow for a decade. This pillar makes the GIL something you can reason about rather than a rumour: one lock, one thread executing Python bytecode at a time, released around waiting.
That single mechanism decides the whole shape of the answer. If your program is
waiting — on a socket, a disk, a database — threads work well, because the
lock is released for the duration of the wait and other threads make progress.
If your program is computing, threads cannot help and processes can, since
each process brings its own interpreter and its own lock; ProcessPoolExecutor
and fork are how you reach them, at the cost of pickling what crosses the
boundary. You will run both and watch the timings confirm it, which is more
convincing than being told.
The second half is the correctness price of sharing state. Two workers
incrementing the same counter is the textbook race, and in CPython the naive
version often appears to work, which is worse than failing — so the exercises
widen the window deliberately until the race is reliable and the fix is
demonstrably necessary. Lock for mutual exclusion, Queue for handing work
between workers without sharing anything, Semaphore for bounding how many run
at once. The pillar closes on the senior version of the question: given a
workload, is the right tool threads, processes, or the event loop of P33 — a
decision framework rather than a preference.