Skip to Content
GuideNative Overhead Evidence

Native overhead evidence

This historical source-build measurement builds on db835af (the Node compatibility PR). It changes three existing native paths: compatible glob patterns share a walk, workers collect glob results locally, and Buffer writes avoid redundant copies. Public signatures and native/Node routing are unchanged.

Method and scope

  • Public api.js entry, including N-API crossings, release builds (pnpm build).
  • Node 22.22.0 and 24.21.0; macOS arm64, Apple M4 Pro, 14 logical CPUs.
  • Two warmups and ten measured samples per implementation/case. Tables show the median in milliseconds. Node, old Vooya and new Vooya rotate execution order each iteration. Raw samples retain outliers; there are no CI speed assertions.
  • Glob fixtures: 8, 2,000 and 20,000 files, up to 50 files per directory, with four extensions. Counts include 2, 41 and 401 directories respectively. Workers: 4. Single pattern is **/*; four patterns are **/*.ts, **/*.js, **/*.md, **/*.txt. Single pattern also returns directories. Node’s async iterator is fully collected so both APIs return the same batch.
  • Writes: 64 B, 64 KiB and 8 MiB Buffer subarrays at byte offset 16; concurrency 1. Each destination is reset to an empty file outside timing. No explicit flush or fsync in any implementation. A synchronous result also passes through the harness’s await, adding the same small continuation overhead to all versions.
  • Reused fixtures and warm filesystem cache on local storage. Setup, validation, GC before timing and cleanup are outside timing. Exact names/Dirent predicates and output bytes are checked after every sample, including warmups.
  • RSS deltas are sampled immediately before/after each operation, not peak memory. They include allocator retention and reclamation and do not quantify the number of payload copies. Native output ownership and JS conversion are included in latency. No platform-wide or cold-cache performance claim is made.

Glob results

NodeFilesPattern/resultNode msBefore msAfter msBefore / after
v22.22.08single0.4531.8081.7501.03x
v22.22.08four0.3787.4691.7834.19x
v22.22.08four-dirents0.4547.5261.7834.22x
v22.22.02,000single6.6513.8663.7141.04x
v22.22.02,000four13.58914.3053.7923.77x
v22.22.02,000four-dirents13.13114.7974.2293.50x
v22.22.020,000single62.16613.98314.1760.99x
v22.22.020,000four118.27449.86014.0703.54x
v22.22.020,000four-dirents118.76556.95518.2093.13x
v24.21.08single0.3621.7311.7450.99x
v24.21.08four0.4706.8641.8883.64x
v24.21.08four-dirents0.4748.2531.8534.45x
v24.21.02,000single6.1033.8243.7721.01x
v24.21.02,000four12.05014.1283.8163.70x
v24.21.02,000four-dirents12.11914.8584.2203.52x
v24.21.020,000single58.81413.70013.0571.05x
v24.21.020,000four110.09747.22012.5903.75x
v24.21.020,000four-dirents113.35453.45818.8962.83x

Grouping removes repeated traversal and worker setup; the strongest benefit is multiple compatible patterns. Single-pattern results are broadly comparable to the old implementation: this report does not establish an independent speedup from local collection alone. Tiny trees remain slower than Node despite the large improvement over the old four-walk implementation.

The planner keeps distinct literal roots, hidden traversal policies and terminal ** root inclusion in separate groups. It does not merge overlapping roots into one broader scan. A single walker yields each path once, so cross-walk hash-based deduplication is skipped when only one group remains. Different groups still perform deduplication. Workers merge their vectors once when they finish.

Local vectors can retain spare capacity, and merging briefly holds local and combined allocations at the same time. Results remain fully materialized; this is not a streaming or bounded-memory API.

Buffer write results

NodeOperationBytesNode msBefore msAfter msBefore / after
v22.22.0writeFile640.1450.0580.0690.84x
v22.22.0appendFile640.1180.0520.0640.82x
v22.22.0writeFileSync640.0560.0540.0451.20x
v22.22.0appendFileSync640.0550.0460.0500.93x
v22.22.0writeFile65,5360.1310.0780.0810.96x
v22.22.0appendFile65,5360.1310.0730.0711.02x
v22.22.0writeFileSync65,5360.0540.0570.0501.14x
v22.22.0appendFileSync65,5360.0490.0500.0451.11x
v22.22.0writeFile8,388,6080.9330.8980.7781.15x
v22.22.0appendFile8,388,6081.0270.9480.8291.14x
v22.22.0writeFileSync8,388,6080.7490.8140.7481.09x
v22.22.0appendFileSync8,388,6080.7800.9090.7671.19x
v24.21.0writeFile640.1470.0700.0720.97x
v24.21.0appendFile640.1300.0560.0580.96x
v24.21.0writeFileSync640.0430.0360.0351.04x
v24.21.0appendFileSync640.0370.0370.0321.14x
v24.21.0writeFile65,5360.1610.0690.0661.05x
v24.21.0appendFile65,5360.1240.0610.0620.98x
v24.21.0writeFileSync65,5360.0410.0430.0411.05x
v24.21.0appendFileSync65,5360.0660.0480.0520.92x
v24.21.0writeFile8,388,6080.9910.9780.8461.16x
v24.21.0appendFile8,388,6081.0361.0020.9021.11x
v24.21.0writeFileSync8,388,6080.7160.8070.7021.15x
v24.21.0appendFileSync8,388,6080.6950.7870.6711.17x

The 8 MiB cases improve over the previous implementation across both runtimes in these final runs. Small inputs show mixed results, including regressions, and sync writes do not consistently beat Node. An exploratory run also showed noisy 8 MiB regressions under filesystem writeback; this is a CPU-copy reduction, not a guarantee against disk scheduling or cache effects. Do not select this API solely for a tiny-write speed claim.

Synchronous Buffer input is borrowed for the duration of the call. Promise calls still copy once at entry, guaranteeing an owned snapshot until completion; the worker writes that snapshot directly without rebuilding a N-API Buffer or copying it again. The owned snapshot is released on the worker after I/O, even if the JS event loop is busy. There is no new unsafe code. String encoding paths are unchanged and have no new performance claim. The removed allocation is one payload-sized Vec per Buffer operation; actual process peak memory is not established here.

Raw evidence and reproduction

Each report records CPU/runtime, source and binary SHA-256 hashes, the baseline revision, timings and per-sample RSS deltas. The source hashes identify the exact native patch before the documentation/evidence commit.

Build the baseline separately, preserving its public wrapper and native loader:

baseline_dir=$(mktemp -d) git archive db835afca93871cb91cc3696a884eab9648bca95 | tar -x -C "$baseline_dir" (cd "$baseline_dir" && pnpm install --frozen-lockfile && pnpm build) pnpm build node --expose-gc scripts/benchmark-native-overheads.mjs "$baseline_dir" .perf/native-overheads-node22.json npm exec --yes --package=node@24 -- node --expose-gc scripts/benchmark-native-overheads.mjs "$baseline_dir" .perf/native-overheads-node24.json

Run the first measurement with Node 22.22.0 and the second with Node 24.21.0 to match the recorded runtimes; the node@24 tag may later resolve to another patch. Run without concurrent builds/tests/benchmarks. The two releases load through distinct public entry paths in the same process. The fixture is temporary and removed on successful completion or an exception inside measurement.

Compatibility validation

Local macOS arm64 release checks: 451 tests on each of Node 22 and 24, TypeScript, Oxlint, Rust formatting, Clippy and documentation build. Regression coverage tests public sync and Promise glob results against Node, including overlapping roots, root exclusions, ignored files, symlinks and Dirent predicates. Buffer tests cover nonzero-offset/empty slices, unchanged input bytes, concurrent entry snapshots and sequential append ordering. Snapshot tests verify Vooya’s existing ownership policy; they do not claim Node snapshots mutable caller buffers identically.

The PR runs Node 22/24 tests on Linux, Windows and macOS x64/arm64. See the PR checks for actual platform results; local macOS evidence is not Windows/Linux performance validation. CI now runs on all PR base branches to validate stacked changes as well as changes targeting main.

Last updated on