b71641a15f
The previous benchmark unfairly gave ACINQ a cached keypair (created once outside the loop) while our sign derived the pubkey each call. Fixed to test both patterns: - "sign (full)": both derive pubkey each call ACINQ: keypair_create + sign32 = 36.3µs Ours: ecmult_gen + sign_internal = 33.6µs → 1.08x faster - "sign (cached)": both reuse precomputed pubkey ACINQ: sign32 with cached keypair = 18.8µs Ours: signXOnly with cached xonly = 17.4µs → 1.08x faster Fair native C-to-C results (x86_64, BMI2): pubkeyCreate: ACINQ 18.0 Ours 16.3 1.10x faster ✓ sign (full): ACINQ 36.3 Ours 33.6 1.08x faster ✓ sign (cached): ACINQ 18.8 Ours 17.4 1.08x faster ✓ verify (BIP-340): ACINQ 36.7 Ours 40.4 0.91x slower ✗ verifyFast (Nostr): ACINQ 36.7 Ours 34.5 1.06x faster ✓ ECDH (cached): ACINQ 37.2 Ours 35.4 1.05x faster ✓ batch(200): ACINQ 42.4 Ours 8.3 5.1x faster ✓ We beat ACINQ on 5 of 6 operations. The only loss is full BIP-340 verify (0.91x) due to ACINQ's 5x52+ADCX/ADOX field assembly. https://claude.ai/code/session_011KVZhDcV2G7idNWEBz12GY