4fd4ad63d1
Replace MULQ with MULX in the inline assembly for fe_mul on x86_64. MULX advantages over MULQ: - Uses RDX as implicit input (not RAX), outputs to two arbitrary regs - Does NOT clobber flags, enabling better instruction scheduling - Enables the compiler to interleave multiplies with carries Requires BMI2 (available on Haswell+ / Zen+). Added -mbmi2 to CMakeLists.txt for x86_64 builds. Note: ADCX/ADOX (ADX extension) for dual carry chains was investigated but the compiler doesn't auto-generate them from __int128 code, and hand-encoding in inline ASM requires restructuring the entire multiply to express two independent carry chains. The MULX-only approach still gives a measurable improvement. fe_mul: 17.2ns → 15.5ns (10% faster) pubkeyCreate: 16.7µs → 15.5µs (7% faster) verifyFast: 36.5µs → 36.1µs (1% faster, 27,700 ops/s) https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY