Add verifySchnorrBatch to Secp256k1InstanceOurs wrapper and add
Android benchmark tests for batch(8) and batch(16).
To get per-event cost: ns/op ÷ batchSize.
JVM benchmark showed 4-7× speedup over individual verify.
Android results TBD.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
The previous commit replaced ALL uLt calls with uLtInline (XOR trick),
which regressed JVM verify from 1.6× to 1.9× vs native C. The XOR trick
is slower than Long.compareUnsigned on HotSpot.
Fix: uLtInline is used ONLY inside the fused FieldMulPlatform.kt inline
functions (which JVM doesn't use — it uses the unfused path). All other
code (U256.addTo/subTo, FieldP.add/sub/half, ScalarN, Glv) keeps the
expect/actual uLt which uses Long.compareUnsigned on JVM.
JVM: verifySchnorrFast 1.3× (restored, improved)
Android: uLt overhead remains for non-fused paths (~11K calls/verify)
but the fused mul/sqr path (the dominant cost) is fully inlined.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Trace profiling showed unsignedMultiplyHighFallback at 16.9% of verify
time (32,524 calls × 82ns = 2.674ms). Although called from inside an
inline crossinline lambda, the function itself was a regular dispatch.
Adding @Suppress("NOTHING_TO_INLINE") inline makes the Kotlin compiler
embed the 4-multiply arithmetic directly at each call site, eliminating
all function dispatch overhead.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Trace profiling on Pixel 8 revealed two massive overhead sources:
1. ExternalSyntheticBackport0.m — 17.5% of verify (3.45ms)
D8/R8 desugaring wraps Math.unsignedMultiplyHigh in a synthetic
backport because minSdk=26 < 35. The backport adds 139ns per call
× 24,875 calls. The actual UMULH intrinsic is ~1ns, but the
wrapper adds ~138ns.
FIX: Use pure-Kotlin unsignedMultiplyHighFallback on Android instead
of Math.unsignedMultiplyHigh. The fallback (4 Long multiplies +
shifts, ~10-20ns) is FASTER than the backported intrinsic (139ns).
Removed all API-level dispatch — a single fallback path for all
Android versions.
2. uLt function calls — 20.7% of verify (4.08ms)
The expect/actual uLt (can't be inline) was called 49,924 times
per verify at 82ns each. Most calls came from the fused
fieldMulReduceWith/fieldSqrReduceWith inline expansions.
FIX: Add uLtInline — a private inline function using the XOR trick
directly. Since it's @Suppress("NOTHING_TO_INLINE") inline, the
Kotlin compiler inlines it at every call site (not the JIT).
Replaces all 55 uLt calls in the fused inline functions.
JVM is unaffected (uses unfused path with Long.compareUnsigned).
Combined: eliminates ~38% of verify overhead on Android.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Wire verifySchnorrFast (x-check only, no y-parity inversion) into:
- Secp256k1InstanceOurs wrapper for app-level usage
- Android benchmark as verifySchnorrFastOurs for Pixel 8 measurement
Expected: ~15% faster than verifySchnorrOurs (~100μs vs ~120μs on Pixel 8)
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
BIP-340 verification checks both R.x == r AND R.y is even. The y-parity
check requires a full field inversion (~270 field ops, ~14% of verify).
verifySchnorrFast skips the y-parity check, verifying only the
x-coordinate in Jacobian coordinates (2 field ops, no inversion).
WHY THIS IS SAFE FOR NOSTR:
For a given x on secp256k1, there are exactly 2 points: (x, y_even) and
(x, y_odd). A signature producing the correct x but wrong y-parity would
require solving the discrete log — equivalent to forging the signature.
The y-parity check is defense-in-depth, not a distinct security boundary.
DO NOT use for Bitcoin/financial protocols — use verifySchnorr for strict
BIP-340 compliance.
JVM benchmark:
verifySchnorrFast: 20,706 ops/s (1.4× vs native C)
verifySchnorr: 18,038 ops/s (1.6× vs native C)
Improvement: ~15% faster
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Allocation audit from Android benchmark showed 19 allocs in signSchnorr
and 4 in verifySchnorr. Most were intermediate ByteArray/LongArray that
can be replaced with pre-allocated scratch buffers.
Changes:
- Add sha256Into() (expect/actual) that writes digest into existing buffer
instead of allocating a new ByteArray(32) per call. Uses
MessageDigest.digest(buf,off,len) on JVM/Android, CC_SHA256 on Apple.
- Add scratch byte buffers to PointScratch: hashBuf(256), bytesTmp1/2(32),
scalarTmp1/2/3 for intermediate scalar results.
- Add ScalarN.reduceTo() allocation-free variant.
- Rewrite signSchnorrInternal to reuse scratch buffers for:
- dBytes serialization (bytesTmp1 instead of U256.toBytes alloc)
- AUX_PREFIX+auxrand hash (hashBuf instead of array concatenation)
- auxHash XOR (scalarTmp1/2/3 instead of U256.fromBytes allocs)
- nonce/challenge hash inputs (hashBuf instead of ByteArray alloc)
- nonce scalar (scalarTmp1 instead of ScalarN.reduce alloc)
- challenge scalar (scalarTmp3 instead of allocs)
- e*d and k+e*d (splitK1/entryTmp2 instead of ScalarN.mul/add allocs)
- Rewrite verifySchnorr to use hashBuf and sha256Into for challenge hash.
Only the 64-byte output signature is allocated per sign call.
Verify allocates nothing for 32-byte messages (hashBuf is large enough).
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
BIP-340 public keys always have even y, so the y-parity prefix byte
(0x02) is redundant when the caller already has the 32-byte x-only
pubkey. This new overload takes the x-only pubkey directly, avoiding:
1. The expensive G multiplication to derive the pubkey (~20μs on Android)
2. The 33→32 byte array copy that signSchnorrWithPubKey does internally
Added to:
- Secp256k1.signSchnorrWithXOnlyPubKey (core implementation)
- Secp256k1Instance (expect/actual, falls back to C lib's signSchnorr
since the native C lib always derives pubkey internally)
- Secp256k1InstanceOurs (uses the optimized pure-Kotlin path)
- Nip01Crypto.signWithPubKey (app-level convenience)
The app's KeyPair already stores the 32-byte x-only pubkey. Callers
like EventAssembler.hashAndSign can pass it through to skip the
G multiplication entirely.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Kotlin generates Intrinsics.checkNotNullParameter at the entry of every
function taking non-null reference types. Bytecode audit showed:
Before: 128 checkNotNullParameter calls across secp256k1 classes
After: 8 (only expression-value checks in non-hot paths)
Per Schnorr verify, this eliminates ~4,000+ invokestatic calls.
On ART (~2-3ns each): saves ~8-12μs per verify.
On HotSpot: neutral (C2 already optimizes null checks to fast branches).
These flags are safe for this module: all internal secp256k1 functions
use non-null LongArray/MutablePoint parameters that are never null.
Applied at the module level (compilerOptions) so all targets benefit.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Two optimizations for verifySchnorr:
1. Jacobian x-coordinate check before toAffine:
Instead of converting to affine first (1 inversion = ~270 field ops),
check X == r·Z² in Jacobian coordinates (1 sqr + 1 mul = 2 ops).
Invalid signatures (x mismatch) are rejected immediately without
paying the inversion cost. Valid signatures still need inversion for
the y-parity check. For Nostr, ~all sigs are valid so this mainly
helps adversarial/spam rejection.
2. Inline wNAF zero-checks in mulDoubleG inner loop:
~70% of wNAF digits are zero. Previously, each still called
addWnafMixedPP (function call overhead: null checks, frame setup).
Now the zero-check is inlined before the call, avoiding ~364
function calls per verify. On ART (5-8ns per call), this saves
~2-3μs. On HotSpot, the JIT already optimized this — no change.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
The fused fieldMulReduceWith + crossinline lambda was designed for ART
which struggles with deep call chains. On HotSpot C2, it creates a
2351-bytecode method (exceeding FreqInlineSize=325) that can't be
inlined into FieldP.mul, and wastes 180 bytecodes on lambda param
shuffling that C2 must clean up.
The unfused path (U256.mulWide + FieldP.reduceWide) produces a tiny
40-bytecode fieldMulReduce that HotSpot easily inlines. HotSpot's 8+
level inlining depth handles the full chain down to the
Math.unsignedMultiplyHigh intrinsic (single MULQ on x86-64).
Benchmark on JVM 21 x86-64: equivalent performance (1.5× verify,
0.9× sign-cached vs native C). Cleaner bytecode with no lambda waste.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Document WHY each approach was chosen, what alternatives were tested,
and what the measured impact was — so future contributors don't
accidentally revert optimizations or repeat failed experiments.
Key decisions documented:
- uLt() expect/actual: why XOR on Android, Long.compareUnsigned on JVM,
and why a shared inline fun in commonMain caused 30% JVM regression
- fieldMulReduceWith: why fused mul+reduce, why inline+crossinline,
why NOT 5x52 limbs, why NOT single-method with all API branches
- @JvmField: why it's required on MutablePoint/AffinePoint/PointScratch,
with bytecode counts showing ~7,450 + ~2,000 virtual getter calls
eliminated per verify
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Bytecode analysis showed MutablePoint.x/y/z, AffinePoint.x/y, and all
PointScratch properties compile to invokevirtual getter calls instead
of direct field reads (getfield). Per verify:
- MutablePoint getX/Y/Z: ~7,450 invokevirtual → getfield
- PointScratch getT/getW/etc: ~2,000+ invokevirtual → getfield
Each invokevirtual has ~3-5ns overhead on ART vs ~1ns for getfield.
@JvmField eliminates the getter method entirely, compiling property
access to a direct field read. On non-JVM targets (iOS), @JvmField
is silently ignored.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Bytecode analysis revealed that every `a.toULong() < b.toULong()`
comparison generates 2 invokestatic calls to ULong.constructor-impl
(NOOPs that return the input unchanged) plus Long.compareUnsigned.
Across all secp256k1 hot paths, this produced 554 NOOP invokestatic
calls in the bytecode, translating to ~18,000 wasted calls per verify.
Replace all toULong() comparisons with an inline uLt() helper that
uses the XOR-with-MIN_VALUE trick directly:
(a xor Long.MIN_VALUE) < (b xor Long.MIN_VALUE)
This produces pure arithmetic bytecode (lxor, lcmp, ifge) with ZERO
method calls, eliminating all ULong.constructor-impl overhead.
Before: lload, invokestatic ULong.constructor-impl, lload,
invokestatic ULong.constructor-impl, invokestatic
Long.compareUnsigned, ifge (6 bytecodes, 3 method calls)
After: lload, ldc MIN_VALUE, lxor, lload, ldc MIN_VALUE, lxor,
lcmp, ifge (8 bytecodes, 0 method calls)
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
The previous approach inlined all 3 API-level branches into a single
fieldMulReduce method (~600 DEX instructions). ART's register allocator
produces suboptimal code for such large methods, causing stack spills
that negate the benefit of eliminating wrapper calls.
Split each API-level path into its own private function (~200 DEX
instructions each). ART JIT-compiles only the hot one (fieldMulApi35
on API 35+) with full optimization, while the tiny dispatch function
(fieldMulReduce) is easily devirtualized and inlined.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
On Android (ART), the unsignedMultiplyHigh wrapper function has per-call
branching for API level detection that prevents ART's JIT from inlining
the intrinsic into the hot loop. This creates ~10,000 extra function
calls per signature verify (20 wrapper calls × 500 field muls).
This commit introduces fieldMulReduce/fieldSqrReduce as expect/actual
functions that fuse U256.mulWide + FieldP.reduceWide into a single
compilation unit. The multiply-high intrinsic is passed as a crossinline
lambda and inlined at each call site, producing platform-specific code
with zero wrapper overhead:
- Android API 35+: Math.unsignedMultiplyHigh inlined directly (UMULH)
- Android API 31-34: Math.multiplyHigh + correction inlined (SMULH)
- Android API <31: pure-Kotlin fallback inlined
- JVM: Math.unsignedMultiplyHigh inlined directly
- Native: pure-Kotlin fallback inlined
The API level check happens ONCE per fieldMulReduce call (outermost
branch) rather than per multiply-high call (innermost loop), so ART
profiles and JIT-compiles only the hot path.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Port the NWC connection buttons from the ZapSetup screen (UpdateZapAmountDialog)
to the AddWalletScreen used by the multi-wallet flow. Users can now connect a
local wallet app via deep link, paste an NWC URI from clipboard, or scan a QR
code — matching the same options already available in the zap settings.
https://claude.ai/code/session_01QFunqjudjjrrn6pnsUSPCS
- AddWalletScreen: dedicated screen with name input + NWC URI paste,
replacing the generic NIP47Setup flow for adding new wallets
- Rename: inline dialog on wallet cards to rename wallets
- Reorder: up/down arrow buttons on wallet cards to reorder the list
- AccountSettings: add renameNwcWallet() and moveNwcWallet() methods
- WalletViewModel: add addWallet(), renameWallet(), moveWallet()
- Fix client.send -> client.publish after main rename
- New Route.WalletAdd for the dedicated add wallet screen
https://claude.ai/code/session_013sWHKeE1uNfAZJWD4FU9mT
Migrate from a single NWC wallet connection to supporting multiple wallets.
Users can now add, remove, and manage multiple NWC wallets, view balance
and transactions per wallet, and select a default wallet for zaps.
Key changes:
- New NwcWalletEntry/NwcWalletEntryNorm data models for wallet storage
- AccountSettings: replace single zapPaymentRequest with nwcWallets list
and defaultNwcWalletId, with backward-compatible changeZapPaymentRequest
- LocalPreferences: new persistence keys with automatic migration from
legacy single-wallet format
- NwcSignerState: derives default wallet URI from multi-wallet settings,
adds sendNwcRequestToWallet for targeting specific wallets
- Account: adds sendNwcRequestToWallet method
- WalletViewModel: supports wallet list, per-wallet balance/info fetching,
wallet selection, set default, and remove operations
- WalletScreen: shows wallet cards with balance, default indicator, and
management actions (set default, remove with confirmation)
- WalletDetailScreen: per-wallet detail view with balance, send/receive
- New Route.WalletDetail for per-wallet navigation
https://claude.ai/code/session_013sWHKeE1uNfAZJWD4FU9mT
In chess note cards (NoteCompose), player hex keys were displayed as
raw strings. Now uses LoadUser + ClickableUserPicture + UsernameDisplay
to show proper profile pictures and display names for:
- Challenge cards (incoming/outgoing)
- Game end cards (both players)
- PGN metadata in game viewers (white/black players)
Added playerContent composable slot to PGNMetadata and ChessGameViewer
so callers can inject platform-specific user rendering.
https://claude.ai/code/session_0171mKrVEfQnNRabmT7Kv4gf
Previously, selecting a community from the home screen top nav filter
would navigate to the community page. Now it filters the home feed
in-place, applying the community filter to all tabs (New Threads,
Conversations, Live) and adjusting subscriptions accordingly.
The in-place filtering infrastructure already existed (SingleCommunityTopNavFilter,
FilterByListParams, HomeOutboxEventsEoseManager) but was bypassed because
community FeedDefinitions had a route that triggered navigation.
https://claude.ai/code/session_0139NCoHmqvp3aRD3FUjuues
All three saves (zap settings, lightning address, payment targets) now
run sequentially inside one accountViewModel.launchSigner block,
so the signer is only invoked once per save instead of three times.
- Extract sendPostSuspend() from UpdateZapAmountViewModel.sendPost()
- Extract saveLnAddressSuspend() from WalletViewModel.saveLnAddress()
- Extract savePaymentTargetsSuspend() from PaymentTargetsViewModel.savePaymentTargets()
- NIP47SetupScreen.onPost calls all three suspend fns in one launchSigner
https://claude.ai/code/session_017NPhwCnSzCuMn9sTX3ca37
- Remove separate Save/Manage buttons; both are now saved via the
screen's existing SavingTopBar save action
- Lightning address: plain OutlinedTextField, saved on confirm
- Payment targets: inline editable list (Column, not LazyColumn) with
add/delete directly in the scrollable content area
- Wire walletViewModel.saveLnAddress() and
paymentTargetsViewModel.savePaymentTargets() into onPost
- Wire paymentTargetsViewModel.refresh() into onCancel
https://claude.ai/code/session_017NPhwCnSzCuMn9sTX3ca37
- Add trailingContent lambda to UpdateZapAmountContent so callers can
inject sections inside its scrollable Column
- Inject lightning address field and payment targets row into
NIP47SetupScreen via trailingContent
- Add WalletViewModel to NIP47SetupScreen to load/save the lightning address
- Add settings icon button to WalletScreen TopAppBar when wallet is configured,
navigating back to NIP47SetupScreen
- Restore WalletScreen to its original layout (no ln/payment sections)
https://claude.ai/code/session_017NPhwCnSzCuMn9sTX3ca37
Same approach as relay filtering: selecting a hashtag from the top nav
filter now filters the home feed in-place across all tabs, with
subscriptions scoped to that hashtag.
https://claude.ai/code/session_01PZLLjE7MWh7PPz1woVqmxM
- Add lightning address field with save button to WalletScreen
- Add payment targets navigation row to WalletScreen
- Remove lightning address field from profile editor (NewUserMetadataScreen)
- Remove payment targets entry from AllSettings
- Extend WalletViewModel to load/save the lightning address via userMetadata
https://claude.ai/code/session_017NPhwCnSzCuMn9sTX3ca37