Concurrency & Performance
Make Python fast and concurrent: meet the GIL, run real threads and processes, reach for asyncio when I/O dominates, and always profile before you optimise.
Add threads to a CPU-bound Python function and it does not get faster. It often gets slower. That single fact — the global interpreter lock — explains most of the confusion around Python performance, and it also explains why the answer is never "use threads" or "use asyncio" but "what is this workload actually waiting on?". This domain teaches the diagnosis before the remedy, because the remedies only make sense once you can classify the problem.
Three pillars, in the order the decision is made. Concurrency comes first: real
threads and real processes, why the GIL makes one useless for computation and
fine for waiting, and the synchronisation primitives — Lock, Queue, Semaphore —
that keep shared mutable state correct when several workers touch it. Then
asyncio, built before it is used: you model an event loop out of generators, so
await is a suspension you have already implemented, then run the real thing
with TaskGroups that give concurrent work a lifetime it cannot outlive, and
bridge back out to blocking or CPU-bound code without stalling the single
thread everything shares. Performance closes the domain with a discipline rather
than a trick: profile to find the hot function, measure allocations, then fix
the algorithm before the constant factor, and prove the win against a budget
instead of a feeling.