c7b164b845
In mulDoubleG (verify), the P-side wNAF-5 affine table (8 odd multiples of P plus 8 GLV λ counterparts) is rebuilt from scratch on every call. This costs ~437 field ops: 1 doublePoint + 7 addPoints + 8 β-multiplies + 1 batch inversion with full field inversion (~270 ops alone). Add a 256-entry direct-mapped cache that stores these affine tables keyed by the point's x-coordinate. On cache hit, the table build is skipped entirely — the cached arrays are used directly (zero-copy). This saves ~27% of mulDoubleG cost (~20% of total verify) for repeated pubkeys. Combined with the liftX cache from the previous commit, repeated-pubkey verifications now skip ~33% of the work. JVM benchmark: verify ratio improved from 2.0x to 1.6-1.8x vs native C (with the benchmark's 100% cache hit rate). Original baseline was 3.0x. https://claude.ai/code/session_017UbWduFi1sLUsgVUMUH2nx