bdf7ab1a15
Add in-place ScalarN operations (mulTo, addTo, negTo) that write to caller-provided output arrays instead of allocating. Add Glv.splitScalarInto that uses pre-allocated scratch buffers from PointScratch, eliminating ~26 LongArray allocations per call (called 2x per verify = ~52 allocs saved). Both mul and mulDoubleG now use the allocation-free split path. Pool all LongArray and MutablePoint allocations in verifySchnorr into thread-local PointScratch (verifyPx/Py/R/S/E/Rx/Ry, verifyPPoint, verifyResult). Add U256.fromBytesInto for decoding into pre-allocated arrays. Total: ~14 object allocations eliminated per verify call. Add pubkey decompression cache (256-entry direct-mapped) that skips the liftX square root (~280 field ops) for repeated pubkeys. In Nostr, the same authors are verified repeatedly, so cache hit rate is high. This saves ~13% of verify cost per cache hit. Increase benchmark warmup/iterations for more stable measurements (e.g., verifySchnorr: 200/500 -> 2000/5000). JVM benchmark result: verify ratio improved from 3.0x to 2.0x vs native C (with 100% cache hit rate on the benchmark's single pubkey). https://claude.ai/code/session_017UbWduFi1sLUsgVUMUH2nx