39c97f7c32
Replace fe_half's normalize-then-branch approach with a branchless mask-based conditional add of P. This eliminates the fe_normalize_full call and branch prediction penalty. Note: dedicated fe_sqr with cross-product doubling was attempted but reverted — with 4x64 limbs, each 64x64 product is 128 bits and doubling overflows uint128. The 5x52 representation wouldn't have this issue (104-bit products, 105 bits doubled) but was rejected earlier for having more total products (25 vs 16). This is a fundamental tradeoff. Performance (x86_64 standalone, µs/op): verifyFast: 51.5 µs (19,422 ops/s) pubkeyCreate: 16.7 µs (59,925 ops/s) signSchnorr: 35.5 µs (28,164 ops/s) https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY