docs: add performance rationale to secp256k1 optimization decisions

Document WHY each approach was chosen, what alternatives were tested,
and what the measured impact was — so future contributors don't
accidentally revert optimizations or repeat failed experiments.

Key decisions documented:
- uLt() expect/actual: why XOR on Android, Long.compareUnsigned on JVM,
  and why a shared inline fun in commonMain caused 30% JVM regression
- fieldMulReduceWith: why fused mul+reduce, why inline+crossinline,
  why NOT 5x52 limbs, why NOT single-method with all API branches
- @JvmField: why it's required on MutablePoint/AffinePoint/PointScratch,
  with bytecode counts showing ~7,450 + ~2,000 virtual getter calls
  eliminated per verify

https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
This commit is contained in:
Claude
2026-04-08 23:53:08 +00:00
parent 30151295ad
commit 6f3793e7ed
7 changed files with 120 additions and 40 deletions
@@ -21,9 +21,16 @@
package com.vitorpamplona.quartz.utils.secp256k1
/**
* JVM: Long.compareUnsigned is a HotSpot JIT intrinsic that compiles
* to a single unsigned CMP + SETB instruction. Much faster than the
* XOR trick on HotSpot.
* JVM: Long.compareUnsigned is a HotSpot C2 JIT intrinsic.
*
* DO NOT replace with the XOR trick `(a xor MIN_VALUE) < (b xor MIN_VALUE)`.
* HotSpot recognizes Long.compareUnsigned and compiles it to a single unsigned
* CMP + SETB instruction pair. The XOR trick generates 2 extra XOR instructions
* that HotSpot does NOT optimize away, causing a ~30% regression on verify
* (1.6× → 2.1× vs native, measured on JDK 21 x86-64).
*
* This is the opposite of Android/ART where the XOR trick is faster because
* it avoids ULong.constructor-impl overhead. See UnsignedCompare.android.kt.
*/
internal actual fun uLt(
a: Long,