4dea4b5db5
Remove fe_normalize/reduceSelf from the end of field multiply and square. After reduceWide, the output is in [0, 2^256) which may include values in [P, P+C) where C = 2^32+977. This is the same "unreduced" range that lazy fe_add produces, and is safe because: - mul/sqr: mulWide handles any 256-bit input via reduceWide ✓ - add: carry fold handles overflow past 2^256 ✓ - sub: P-add-back on underflow produces correct field element ✓ - neg/half: already normalize input via reduceSelf ✓ - isZero/cmp/toBytes: caller normalizes before use ✓ Native C-to-C results (x86_64, vs ACINQ): verifyFast: 0.95x → 0.99x (essentially tied with ACINQ!) sign (cached): 1.18x → 1.24x faster ECDH: 1.05x → 1.06x faster batch(200)/event: 7.2µs → 6.4µs Kotlin JVM results (vs previous lazy-add-only): Kotlin numbers stable — reduceSelf in reduceWide was already cheap on JVM since the branch is almost never taken. https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY