547be89577
FieldP.mul/sqr were calling ThreadLocal.get() for every invocation (~500+ times per scalar multiplication, ~20-30ns each on JVM). Point operations (doublePoint, addMixed, addPoints) each did an additional ThreadLocal.get() for their scratch buffers. Fix: add overloads that accept a pre-fetched wide buffer (LongArray(8)) and PointScratch. Top-level entry points (mulG, mul, mulDoubleG) fetch the ThreadLocal once and thread it through all inner calls. Results (ops/s, vs native JNI): - pubkeyCreate: 19,163 → 29,205 (+52%, 3.0x → 2.2x) - signSchnorr cached: 13,007 → 18,397 (+41%, 2.1x → 1.5x) - signSchnorr: 5,365 → 7,490 (+40%, 5.7x → 3.7x) - verifySchnorr: 3,840 → 4,873 (+27%, 7.2x → 5.4x) - ECDH: 5,569 → 7,870 (+41%, 5.5x → 3.8x) https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg