8291987145
Analyzed lazy fe_sub: on underflow, the unsigned-wrapped value (a - b + 2^256) differs from (a - b + P) by C = 2^32 + 977. When multiplied: (a-b+2^256)*x ≠ (a-b)*x mod p (off by x*C mod p). The P-add-back on underflow is mandatory for correctness. This is a fundamental difference from 5x52 lazy reduction where magnitude tracking keeps values representable. With 4x64 fully-packed limbs, sub MUST add P back, but add CAN skip reduceSelf since values in [P, 2^256) differ from [0, C) which is handled by mulWide+reduceWide. Also evaluated: - WINDOW_G 12→14: only 1.2µs savings for 4x memory (128→512KB). Not worth it on phones where L1 cache is 128-256KB. - fe_sqrt optimization: only 831ns overhead over theoretical minimum of 5.85µs. The addition chain is already near-optimal. https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY