P34 · Performance & Profiling
profiling, memory & optimisation
Make slow Python fast, on evidence not instinct: profile to find the hot spot, measure memory with tracemalloc, then optimise - complexity before constant factors - and confirm the win against a real budget.
Ask any group of engineers which line of their program is slowest and most will answer confidently. Profile it and most will be wrong — usually because the expensive thing is not the clever algorithm they worried about but a lookup inside a loop, or a function called far more often than anyone realised. This pillar's entire premise is that optimisation without measurement is guesswork with extra steps, and it enforces that order.
Measurement comes first and has two halves, because "slow" and "heavy" are
different complaints. cProfile answers where the time went — which function,
how many calls, cumulative against own time, and how to read the difference,
since a function that is slow because it calls something slow needs a different
fix from one that is slow itself. tracemalloc answers where the memory went,
which is how you find both the peak that will fail on a smaller machine and the
leak that grows quietly across a long-running process.
Only then do you change anything, and the order there matters just as much:
complexity before constant factors. Removing a quadratic beats micro-optimising
the inner loop of that quadratic, every time, and no amount of vectorisation
rescues an algorithm that should not have been written. Once the complexity is
right, the constant factors become worth attention — when numpy pays for
itself and when it does not, when caching genuinely helps and when it just moves
the cost — and the capstone closes the loop by requiring the win to be
demonstrated against a metered budget rather than asserted.