40cb0e270c
Pass PointScratch through mulDoubleG instead of re-fetching from ThreadLocal. The verify path was calling ThreadLocal.get() 2x: once in verifySchnorrFast and again in mulDoubleG. On ART, each ThreadLocal.get costs ~41µs (hash table probe), totaling ~83µs per verify (~1% of total). mulDoubleG now accepts an optional PointScratch parameter (defaults to scratch.get() for backward compat). The verify path passes its already-fetched scratch through. Also fix all remaining LongArray copyInto calls with default params: - MutablePoint.copyFrom: 3 calls per copy (x, y, z) - MutablePoint.setAffine: 2 calls (x, y) - mulDoubleG P-table build: 2 calls per table entry - mul P-table build: 2 calls per table entry - batchToAffine: 2 calls - U256.copyInto: was delegating with defaults Each copyInto$default adds a bitmask check + 3 branches + arraylength per call. With ~13 LongArray copies per verify, this eliminates ~52 extra branch instructions from the hot path. https://claude.ai/code/session_015CtM5k88rF7WFgX8o2AGNR