ae77935625
Remove reduceSelf from FieldP.add, making it lazy (same approach as the C implementation). Values may be in [0, 2^256) between additions. Safety analysis: - mul/sqr: reduceWide handles any 256-bit input ✓ - neg: added reduceSelf before P - a (prevents underflow) ✓ - half: added reduceSelf before conditional add P ✓ - sub: already adds P back on borrow (self-normalizing) ✓ - isZero/cmp: called on sub outputs (normalized) or after mul ✓ - verifySchnorrCore: added reduceSelf before Jacobian x-check ✓ - toBytes: only called on toAffine outputs (from mul, normalized) ✓ JVM benchmark results (ops/sec, HotSpot C2): pubkeyCreate: 32,096 → 37,211 (+15.9%) signXOnly: 31,149 → 34,507 (+10.8%) sign: 16,137 → 17,822 (+10.4%) verify: 12,733 → 14,027 (+10.2%) verifyFast: 14,734 → 15,449 (+4.9%) ECDH: 10,975 → 12,102 (+10.3%) batch(200): 91,611 → 102,354 (+11.7%) The improvement is larger than expected (10-16% vs estimated 6%) because HotSpot's branch prediction overhead for reduceSelf (virtual dispatch + 4 Long comparisons) is higher than the ~2.5ns estimated for native code. https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY