0e3c2fcda5
Karatsuba multiplication (splitting 8 limbs into 4+4 halves for 48 inner products instead of 64) was implemented and tested but reverted because the overhead of extra additions, carry propagation, and 5 temporary array allocations per call negates the product-count savings at only 8 limbs. The crossover point where Karatsuba beats schoolbook is typically ~32+ limbs on hardware with fast multiply. https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg