41e54ca33e
The fused fieldMulReduceWith + crossinline lambda was designed for ART which struggles with deep call chains. On HotSpot C2, it creates a 2351-bytecode method (exceeding FreqInlineSize=325) that can't be inlined into FieldP.mul, and wastes 180 bytecodes on lambda param shuffling that C2 must clean up. The unfused path (U256.mulWide + FieldP.reduceWide) produces a tiny 40-bytecode fieldMulReduce that HotSpot easily inlines. HotSpot's 8+ level inlining depth handles the full chain down to the Math.unsignedMultiplyHigh intrinsic (single MULQ on x86-64). Benchmark on JVM 21 x86-64: equivalent performance (1.5× verify, 0.9× sign-cached vs native C). Cleaner bytecode with no lambda waste. https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY