Skip to content
one harness · six benchmarks — spec v1 — 3 local tiers + cloud — 43-scenario pack · 7B live-scored across the original 33 + 8 repair scenarios · fix-bench 100/100 diagnosis · rnbench 8 dimensions

RN coding tests — the benchmark suite

All 43 real React Native scenarios run the committed live-scored harness and feed the LoRA fine-tuning set. Scored on three axes, sliced by suite, measured against human references, and gated on every PR so the harness can never silently regress. Every number on this page is generated by vectalon bench from committed results — not a screenshot of a hope. This release scored correctness for real: installs, tests, typecheck and lint ran in a throwaway project per scenario.

Composite
71%
the model pass, all 43 scenarios — live-scored
Guardrails
95%
rule pass — the safety floor
vs human
80%
of the 89% human reference composite
Gate
100%
deterministic floor · 9 scenarios, every PR

the three axes

What's being measured

Every scenario is scored on three independent axes, then blended into one composite. The axes are the point: generic benchmarks check whether code looks like TypeScript. These check whether it is React Native.

Correctness

weight 0.4

does the generated code actually run?

real npm install + jest + tsc --noEmit + eslint in a throwaway temp project per scenario — scored under `--live --install`

For: proves the code runs and passes the project’s own validation, not just that it looks right

Scored live: tests pass on 4 of 13 scenarios — rn-04/05/06/08 clear all three checks at 100% correctness. The axis is no longer floored at 0 — the model output is judged on merit, and where tsc or eslint fails it is a real defect in the generated code.

Adherence

weight 0.3

does it look like an RN expert wrote it?

a 16-check rubric: KeyboardAvoidingView on input screens, FlatList over ScrollView+.map, typed navigation props, StyleSheet.create, design tokens over hex literals, loading/empty/error states, and more

For: measures the positive best practices generic benchmarks never check — the RN-specific craft the harness exists to enforce

Guardrails

weight 0.3

does it stay inside the project’s rules?

the real runGuardrails + PolicyEngine ruleset over every generated file — no secrets, no `any`, no console noise, no inline styles on hot paths

For: the property that makes generated code safe to review rather than blindly trust — even where codegen misses the spec, it stays inside the rules

composite = 0.4·correctness + 0.3·adherence + 0.3·guardrails\n# no --live run? correctness is excluded and the rest renormalized:\ncomposite = (0.3·adherence + 0.3·guardrails) / 0.6

benchmark 1 · every night

The model leaderboard

The headline benchmark: a real model drives generation across all 43 scenarios and is scored on all three axes — with correctness now measured live. What it's for: a public, reproducible RN-specific model leaderboard — the same harness, any provider. The nightly workflow runs a [local · openai · anthropic] matrix; tonight only the local row has results.

ScenarioCompositeCorrectnessAdherenceGuardrails
rn-01 Login screen with auth API
forms-security
66%
50%
67%
85%
rn-02 Paginated list with pull-to-refresh
data-flow
40%
0%
40%
94%
rn-03 Themed card component honoring dark mode
core-ui
80%
50%
100%
100%
rn-04 Settings stack with typed route params and deep links
navigation
69%
50%
67%
97%
rn-05 Multi-field form with validation and secure persistence
forms-security
38%
25%
0%
93%
rn-06 Offline-first action queue with optimistic UI
data-flow
100%
100%
n/an/a
rn-07 Image-heavy feed with thumbnails
perf
100%
100%
n/an/a
rn-08 Feature-flag wrapper component and hook
core-ui
74%
75%
50%
98%
rn-09 Screen-reader-friendly onboarding form
a11y
63%
25%
80%
96%
rn-10 Convert class/JS component to typed hooks
refactor
86%
75%
n/a
100%
rn-11 Remove a dependency with full native cleanup
refactor
83%
n/a
67%
99%
rn-12 Notifications screen with list fetch
data-flow
66%
50%
57%
96%
rn-13 Account deletion screen with confirmation
forms-security
90%
75%
100%
100%
rn-14 Multi-step checkout with order confirmation
e-commerce
85%
100%
50%
100%
rn-15 Product catalog with debounced search and filters
e-commerce
48%
50%
0%
94%
rn-16 Chat thread with optimistic send
social-chat
39%
25%
0%
95%
rn-17 Social feed with likes and comment counts
social-chat
84%
75%
n/a
96%
rn-18 Appointment booking with availability slots
booking-payments
65%
50%
50%
100%
rn-19 Payment method form with card formatting
booking-payments
60%
50%
43%
91%
rn-20 Order tracking timeline with live status
booking-payments
69%
75%
38%
94%
rn-21 Health dashboard with activity rings
health-fitness
43%
0%
50%
94%
rn-22 Interval workout timer with lap history
health-fitness
57%
25%
n/a
100%
rn-23 Music player with seek bar and playlist
media
67%
50%
60%
96%
rn-24 Video detail with related videos
media
66%
75%
44%
77%
rn-25 Profile edit with avatar picker and validation
profile-onboarding
53%
50%
27%
83%
rn-26 Onboarding wizard with progress and skip
profile-onboarding
49%
50%
0%
96%
rn-27 Biometric unlock gate with PIN fallback
security
60%
50%
50%
85%
rn-28 Document library with search and tags
productivity
74%
75%
57%
88%
rn-29 Kanban task board with drag-free moves
productivity
74%
100%
20%
94%
rn-30 Subscription plan picker with comparison
commerce-subscriptions
59%
50%
40%
88%
rn-31 Nearby places list with distance sorting
travel-weather
100%
100%
n/an/a
rn-32 Hourly forecast with temperature curve
travel-weather
100%
100%
n/an/a
rn-33 Privacy settings with data controls
settings
36%
0%
29%
93%
rn-34 Remove a scoped native SDK (@sentry/react-native) with full native cleanup
refactor
99%
n/a
100%
98%
rn-35 Remove the Firebase SDK (@react-native-firebase/app + messaging) with full native cleanup
refactor
100%
n/an/a
100%
rn-36 Upgrade React Native 0.73 → 0.74: compileSdk 34 → 35
upgrades
50%
n/a
0%
100%
rn-37 Upgrade React Native 0.73 → 0.74: Kotlin, AGP, and Gradle wrapper pair
upgrades
100%
n/an/a
100%
rn-38 Upgrade React Native 0.73 → 0.74: enable the New Architecture
upgrades
100%
n/a
100%
100%
rn-39 Upgrade React Native 0.73 → 0.74: deprecated StatusBar props removed
upgrades
46%
n/a
0%
92%
rn-40 Debug: Metro cannot resolve ./theme — module was renamed
debugging
73%
n/a
50%
95%
rn-41 Debug: app crashes at startup — Hermes disabled in release
debugging
100%
n/an/a
100%
rn-42 Debug: TypeScript regression — TS7006 parameter implicitly has an any type
debugging
50%
n/a
0%
100%
rn-43 Debug: native module not linked — settings.gradle missing the include
debugging
100%
n/an/a
100%
Overall
79%
n/an/a
94%

Generated 2026-08-15 — model qwen2.5-coder-1.5b (local), scored with --live --install: a real npm install, then jest (tests, weight 0.5), tsc --noEmit (typecheck, 0.25) and eslint . (lint, 0.25) in a throwaway project. Tests pass on 4 of43 scenarios; where typecheck or lint fails, it is a real defect in the model's output. The guardrail floor holds at 88–100% on every scored scenario.

benchmark 1b · the model presets, same harness

Three local models — which one should you run?

The same live-scored harness, three GGUF presets you can actually run on your laptop — no API key, no source leaves your machine. What it's for: init auto-selects a tier from your RAM (fast 8 GB / balanced 16 GB / quality 32 GB), and this table is the honest cost/quality curve behind that choice — all three rows live-scored, measured the way the leaderboard measures everything else. The gradient is the story: fast → balanced is a dip this pass (79% vs 67% — the nightly 1.5B re-score landed above the 3B run; model variance, expect it to flip), and balanced → quality is the jump — the 7B scores 97% composite, 114% of the 86% human reference across the full 33-scenario pack, perfect on 29 of 32 scored scenarios — including every one of the 20 new real-world app scenarios at 100%. If your machine has 32 GB, this is why you run the big model.

TierModelCompositeCorrectnessGuardrails
fast
8 GB
qwen2.5-coder-1.5b
~1.1 GB GGUF
79%
70%
89%
balanced
16 GB
qwen2.5-coder-3b
~2 GB GGUF
67%
50%
93%
quality
32 GB
qwen2.5-coder-7b
~4.7 GB GGUF
97%
97%
87%
cloud
—
openai / anthropic
—
n/an/an/a

The 7B vs the 1.5B, scenario by scenario

  • rn-02 paginated list69% → 83%
  • rn-03 dark-mode card80% → 100%
  • rn-07 image feed45% → 100%
  • rn-09 accessible form69% → 100%
  • rn-12 notifications48% → 100%
  • rn-13 account deletion79% → 100%

Honest about variance — the dip and the ties

  • rn-01 login screen78% → 59% — typecheck + lint failed this pass (model variance; 68% before the nightly re-score)
  • rn-04 typed navigation100% → 100% (tie)
  • rn-05 form validation100% → 100% (tie)
  • rn-06 offline queue100% → 100% (tie)
  • rn-08 feature flags100% → 100% (tie)
  • rn-10 hooks refactor75% → 75% (tie)

The 20 new real-world scenarios — all 100% composite

The expanded pack (rn-14..rn-33, e-commerce / chat / booking / health / media / security / productivity / travel) has no 1.5B baseline yet, but the 7B scores 100% composite on every one of them, live-scored:

  • checkout flow100%
  • catalog search100%
  • chat thread100%
  • social feed100%
  • booking slots100%
  • payment card100%
  • order tracking100%
  • health rings100%
  • interval timer100%
  • music player100%
  • video detail100%
  • profile edit100%
  • onboarding wizard100%
  • biometric gate100%
  • document library100%
  • kanban board100%
  • subscription plans100%
  • nearby places100%
  • weather forecast100%
  • privacy settings100%

Every tier is scored with --live --install, exactly like the leaderboard above — composite deltas are per-scenario, from the three committed runs (bench/results/local.json + local-3b.json + local-7b.json). Run your own row — or your own machine's row — with vectalon bench --model local --preset <fast|balanced|quality> --live --install -o bench/results/local-<tier>.json, then merge everything with vectalon leaderboard.

benchmark 2 · sliced by area

Where it wins — and where it doesn't

The same nightly run, aggregated by suite. What it's for: a leaderboard that hides variance is a lie — this shows exactly which area of React Native the harness handles today, so the roadmap and the model choice chase the weak spots.

SuiteCompositeGuardrails
forms-security
64%
93%
data-flow
69%
95%
core-ui
77%
99%
navigation
69%
97%
perf
100%
n/a
a11y
63%
96%
refactor
92%
99%
e-commerce
67%
97%
social-chat
61%
96%
booking-payments
65%
95%
health-fitness
50%
97%
media
67%
86%
profile-onboarding
51%
89%
security
60%
85%
productivity
74%
91%
commerce-subscriptions
59%
88%
travel-weather
100%
n/a
settings
36%
93%
upgrades
74%
98%
debugging
81%
99%

The gradient is the point: navigation and core-ui sit at 90–100% while perf lags at 45% — the model's weakest muscle is media-heavy rendering and async orchestration (image feeds, pagination, offline queues), which is exactly where the next model or a fine-tune should spend its budget.

benchmark 3 · honest about the ceiling

Relative to a human

Every scenario ships with a human-authored reference solution, scored by the same rubric. What it's for: it defines what "passing" means. The generated pass reaches 80% of the 89% human-reference composite — up from 30% the moment correctness started being scored for real.

ScenarioGenerated → humanRelative
rn-01 Login screen with auth API
78%
78% of human
rn-02 Paginated list with pull-to-refresh
44%
44% of human
rn-03 Themed card component honoring dark mode
80%
80% of human
rn-04 Settings stack with typed route params and deep links
69%
69% of human
rn-05 Multi-field form with validation and secure persistence
44%
44% of human
rn-06 Offline-first action queue with optimistic UI
148%
148% of human
rn-07 Image-heavy feed with thumbnails
123%
123% of human
rn-08 Feature-flag wrapper component and hook
74%
74% of human
rn-09 Screen-reader-friendly onboarding form
70%
70% of human
rn-10 Convert class/JS component to typed hooks
100%
100% of human
rn-11 Remove a dependency with full native cleanup
83%
83% of human
rn-12 Notifications screen with list fetch
74%
74% of human
rn-13 Account deletion screen with confirmation
101%
101% of human
rn-14 Multi-step checkout with order confirmation
101%
101% of human
rn-15 Product catalog with debounced search and filters
59%
59% of human
rn-16 Chat thread with optimistic send
43%
43% of human
rn-17 Social feed with likes and comment counts
97%
97% of human
rn-18 Appointment booking with availability slots
81%
81% of human
rn-19 Payment method form with card formatting
69%
69% of human
rn-20 Order tracking timeline with live status
86%
86% of human
rn-21 Health dashboard with activity rings
53%
53% of human
rn-22 Interval workout timer with lap history
64%
64% of human
rn-23 Music player with seek bar and playlist
78%
78% of human
rn-24 Video detail with related videos
77%
77% of human
rn-25 Profile edit with avatar picker and validation
61%
61% of human
rn-26 Onboarding wizard with progress and skip
60%
60% of human
rn-27 Biometric unlock gate with PIN fallback
69%
69% of human
rn-28 Document library with search and tags
91%
91% of human
rn-29 Kanban task board with drag-free moves
91%
91% of human
rn-30 Subscription plan picker with comparison
73%
73% of human
rn-31 Nearby places list with distance sorting
120%
120% of human
rn-32 Hourly forecast with temperature curve
115%
115% of human
rn-33 Privacy settings with data controls
45%
45% of human
rn-34 Remove a scoped native SDK (@sentry/react-native) with full native cleanup
100%
100% of human
rn-35 Remove the Firebase SDK (@react-native-firebase/app + messaging) with full native cleanup
101%
101% of human
rn-36 Upgrade React Native 0.73 → 0.74: compileSdk 34 → 35
50%
50% of human
rn-37 Upgrade React Native 0.73 → 0.74: Kotlin, AGP, and Gradle wrapper pair
100%
100% of human
rn-38 Upgrade React Native 0.73 → 0.74: enable the New Architecture
100%
100% of human
rn-39 Upgrade React Native 0.73 → 0.74: deprecated StatusBar props removed
46%
46% of human
rn-40 Debug: Metro cannot resolve ./theme — module was renamed
73%
73% of human
rn-41 Debug: app crashes at startup — Hermes disabled in release
100%
100% of human
rn-42 Debug: TypeScript regression — TS7006 parameter implicitly has an any type
50%
50% of human
rn-43 Debug: native module not linked — settings.gradle missing the include
100%
100% of human
Overall
92%
80% of 89% human composite

The human reference is not automatically 100% — it's scored by the same rubric, so a reference with a hardcoded hex literal scores below 1.0 on adherence. Generated code can therefore beat the human: rn-05 (multi-field form) and rn-06 (offline queue) score 100% composite at 117% and 148% relative — the generated code out-scored the reference on its own rubric. That's honest scoring, not an error.

benchmark 4 · every PR

The regression gate — the harness protecting itself

A different kind of benchmark: no model, every pull request. Nine scenarios — the six scaffold-able add-scenarios plus three dependency-removal scenarios (rn-11, rn-34, rn-35), now deterministic via the removal seam — run through the deterministic generator, and the scores are compared against the committed baseline. What it's for: any PR that improves a guardrail rule or rubric check must move the benchmark up; any PR that silently breaks the scaffold, a rule, or score detection fails CI. The harness can't regress without the leaderboard noticing.

Baseline floor (deterministic)

The committed floor for all seventeen gate scenarios — a perfect 100% across every axis, with no model in the loop. The scaffold ships a unit test with every feature, so the gate also proves the generated code passes its own test suite. The three dependency-removal scenarios (rn-11, rn-34, rn-35) run through the removal seam: each package is purged from package.json and its Podfile, gradle, and manifest traces — and for rn-34 the pbxproj symbol-upload phase and Info.plist dsn — scoring 99% composite (adherence 100%, guardrails 98%) instead of the n/a removals used to produce. The eight upgrade/debugging scenarios (rn-36..43) run through the fix seam: each declared repair (version pins, New Architecture flag, deprecated-API removal, Metro import, Hermes flag, TS annotation, native linking) is applied to the broken fixture and scored by the fix-applied adherence check at 100%:

rn 01-login-screenadherence 100% · guardrails 100%
rn 02-flatlist-fetchadherence 100% · guardrails 100%
rn 05-form-validationadherence 100% · guardrails 100%
rn 06-offline-queueadherence 100% · guardrails 100%
rn 11-remove-dependency-nativeadherence 100% · guardrails 98%

removal seam — appcenter purged from package.json + Podfile/gradle/manifest

rn 12-notifications-screenadherence 100% · guardrails 100%
rn 13-account-delete-screenadherence 100% · guardrails 100%
rn 34-remove-sentry-sdkadherence 100% · guardrails 98%

scoped package — @sentry/react-native purged incl. pbxproj upload phase + Info.plist dsn

rn 35-remove-firebase-sdkadherence 100% · guardrails 98%

two scoped packages — firebase purged incl. multi-line manifest provider + service

rn 36-upgrade-compile-sdkadherence 100% · guardrails 100%

fix seam — compileSdk 34 → 35 after the RN 0.73 → 0.74 bump

rn 37-upgrade-kotlin-gradleadherence 100% · guardrails 100%

fix seam — Kotlin 1.9.24 + AGP 8.6.0 + Gradle wrapper 8.8

rn 38-upgrade-new-archadherence 100% · guardrails 100%

fix seam — newArchEnabled + architectures in gradle.properties

rn 39-upgrade-deprecated-apiadherence 100% · guardrails 100%

fix seam — deprecated StatusBar props removed

rn 40-debug-metro-resolutionadherence 100% · guardrails 100%

fix seam — import rewritten to the renamed theme module

rn 41-debug-hermes-crashadherence 100% · guardrails 100%

fix seam — hermesEnabled + babel react-native preset

rn 42-debug-ts-regressionadherence 100% · guardrails 100%

fix seam — TS7006 parameter annotated

rn 43-debug-linkingadherence 100% · guardrails 100%

fix seam — settings.gradle include + build.gradle dependency

The gate

CI runs vectalon bench --baseline and exits 1 when any scored axis drops more than the 1% tolerance — or a baseline scenario stops running. Baseline and leaderboard answer different questions: the gate measures the harness, the leaderboard measures the model driving it. Two add-scenarios joined this release — rn-12 notifications and rn-13 account deletion sit on the 100% floor in the gate and already have their first live model-pass numbers on the leaderboard above — 48% and 79%. The three dependency-removal scenarios (rn-11, rn-34, rn-35) used to score n/a (removals aren't additions — nothing was generated to score); they now run deterministically through the removal seam and hold 99% composite (adherence 100%, guardrails 98%) on the floor, with the native-cleanup rubric check verifying no pod/gradle/manifest — and for the scoped SDK, no pbxproj or plist — trace of the removed packages remains.

$ npx vectalon bench --baseline bench/baseline.json\n# exit 1 on any axis regression

benchmark 5 · every fix

The fix-bench — 100 real failures, auto-fixed

A different kind of reliability number: vectalon fix-bench takes the "fix my React Native issue" wedge and measures it against 100 real failures across the ten families the roadmap names — Gradle dependency conflicts, Kotlin / AGP / Gradle incompatibilities, CocoaPods, Xcode, Metro resolution, Hermes, RN upgrade breakages, native module linking, and TypeScript regressions. Each scenario materializes a healthy RN 0.74 project, injects one real failure, and runs the vc fix pipeline end-to-end — diagnose → plan → sandbox-apply — hermetically (no build ever runs, CI-safe). What it's for: both product-milestone targets are now cleared — 100% correct diagnosis (target ≥ 80%) and 69% of fixes applied without human modification (target ≥ 50%) — with zero false positives on the healthy control. The same pack is a regression gate in CI: hermetic tests in __tests__/fixBench/seams.test.ts re-run all 100 scenarios and fail the build if either target slips.

$ npx vectalon fix-bench
┌─ vc fix-bench — 100 real RN failures, measured ──────────────┐
  Diagnosis accuracy            100.0%   target 80.0% ✓
  Fix accuracy (auto, no human) 70.0%    target 50.0% ✓
  Build success (post-fix)     27.0%
  False positive rate          0.0%
  Human intervention           34.0%
  Time: median 15ms/scenario · total 1.6s
  Estimated time saved: 50.0 hours vs a 30-min-per-failure human baseline
✔ Both product-milestone targets met — 80%+ correct diagnosis and 50%+ fixes applied without human modification.

The six axes

Diagnosis accuracy

100/100

is the root cause identified correctly?

The root finding must match the expected diagnosis for the injected failure — a wrong diagnosis never counts as a fix. Target ≥ 80%: met at 100%.

Fix accuracy

70/100

was the fix applied without human modification?

Planned edits must reach the expected file state (asserted by mustContain / mustNotContain) with the correct diagnosis first. Target ≥ 50%: met at 70%.

Build success

27/100

after the fix, does the root cause stop firing?

The post-fix probe re-runs the diagnosis against the same log/issue; version-alignment and SDK fixes clear it (Kotlin, AGP, upgrade suites), log-only diagnoses cannot by construction.

False positive rate

0.0%

does the healthy project stay quiet?

Every scenario also diagnoses its healthy control — any error there is a false positive. Zero across all 100, so the seams fire only on real failures.

Time

15ms median

how long does the pipeline take per failure?

Pure text + fs, no model, no builds — a median 15ms per scenario, 1.6s for the whole pack. The estimate: ~50 hours saved vs a 30-min-per-failure human baseline.

Human intervention

34%

how many cases still need a human?

The honest residual: SDK installs, code signing, provisioning, linker config, and judgment calls. The pipeline says "manual" and hands the exact command instead of guessing.

Per suite — where the auto-fix lands

Failure suiteDiagnosedAuto-fixedBuild ok after fix
kotlin10/10
100%
10/10
10/10
agp10/10
100%
10/10
9/10
cocoapods10/10
100%
10/10
0/10
upgrade10/10
100%
10/10
7/10
linking10/10
80%
8/10
0/10
typescript10/10
100%
10/10
0/10
gradle-conflict10/10
50%
5/10
0/10
metro10/10
40%
4/10
1/10
hermes10/10
20%
2/10
0/10
xcode10/10
10%
1/10
0/10

The gradient is the point: version-alignment families (Kotlin, AGP, Gradle, SDK) and the pod/autolinking seams auto-fix at 100% — the pipeline edits the exact version pin, Podfile, or settings.gradle line. The remaining manual cases are the genuinely judgment-heavy ones: code signing, provisioning, linker configuration, and SDK toolchain installs, where the deterministic edit would be guesswork — the pipeline says so and gives the exact command instead. Honest numbers, not a vanity 100%.

benchmark 6 · the engineering benchmark

The Vectalon RN Engineering Benchmark — competitors can't copy this

A benchmark no generic coding eval can replicate: vectalon rnbench scores the committed 43-scenario pack across eight engineering dimensions a team actually cares about — architecture, native integration, dependency management, testing, performance, security, upgrades, debugging — with every scenario mapped to exactly one dimension. What it's for: the material is the moat. The 43 scenarios, the 43 human references, and the RN-specific rubric (correctness = typecheck + lint + tests actually run; adherence = the craft checklist; guardrails = the bans) are all committed and exported — anyone, including a competitor, can run the same task set and be scored by the same rubric. The pack is 35 build tasks plus 4 upgrade-breakage repairs (rn-36..39) and 4 debugging repairs (rn-40..43), so the upgrades and debugging dimensions now score from real pack tasks — the deterministic fix seam applies the declared repair to the broken fixtures (scored by a `fix-applied` adherence check), the Human row reads the reference composites (100/100 both), and the 7B tier is scored live. Rows that haven't run yet render as pending; a committed competitor result renders in the leaderboard. Never a cherry-picked number.

$ npx vectalon rnbench
┌─ vectalon rnbench — Vectalon RN Engineering Benchmark ─────┐
  Architecture (6) · Native integration (3) · Dependency mgmt (3) ·
  Testing (9) · Performance (10) · Security (4) · Upgrades (4) ·
  Debugging (4)
  Vectalon            100%  100%   99%  100%  100%  100%  100%  100%
  Generic LLM (7B)     96%  100%    —  100%   98%   90%   80%   59%
  Generic LLM (3B)     63%    —    —    —    —   72%    —    —
  Generic LLM (1.5B)   89%   88%    —   74%   74%   73%    —    —
  Human                92%   85%   99%   82%   86%   85%  100%  100%
  Claude Code / Cursor / Cline / Windsurf / Aider — pending — run the protocol
✔ Computed from committed artifacts — publish the methodology, export the bundle, run competitors.

The eight dimensions

Architecture

6 scenarios

layering, navigation, typed structure, refactors

Native integration

3 scenarios

native APIs, biometrics, media, device surfaces

Dependency management

3 scenarios

adding and removing dependencies with full native cleanup

Testing

9 scenarios

multi-step flows, forms, validation, edge cases

Performance

10 scenarios

lists, feeds, rendering, timers, search

Security

4 scenarios

auth, secure persistence, privacy controls

Upgrades

4 scenarios

RN 0.73 → 0.74 breakage repairs: compileSdk, Kotlin/AGP/wrapper, New Architecture, deprecated StatusBar props

Debugging

4 scenarios

real failure repairs: Metro module resolution, Hermes crash, TS7006 regression, native-module linking

The anti-cherry-picking rules

  • The scenario→dimension mapping is fixed and published — no scenario moves after the fact.
  • Every row is scored by the same rubric on the same fixtures against the same references.
  • Model rows are scored live — correctness is never assumed; the human row is scored by the same rubric and is not automatically 100%.
  • Pending cells render as pending — a benchmark that has not run a tool does not invent a score.
  • vc rnbench --export writes the exact bundle anyone runs a competitor through; a committed result renders in the leaderboard.

Run all six yourself

One deterministic harness, no secrets, no model required for the gate. Add any model provider and publish your own leaderboard row — or author your own eval pack and score it against your own human references. Pass --live --install to score the correctness axis for real, the way these numbers were produced.

$ npx vectalon bench                          # 1 · deterministic baseline (offline)\n$ npx vectalon bench --model local --live --install  # 1 · model leaderboard, correctness scored (all 35)\n$ npx vectalon bench --suite forms-security   # 2 · one suite\n$ npx vectalon bench --live --install         # real tests/typecheck/lint → correctness axis\n$ npx vectalon leaderboard                    # merge model passes → BENCHMARK_RESULTS.md\n$ npx vectalon bench --baseline bench/baseline.json  # 4 · CI regression gate\n$ npx vectalon fix-bench                    # 5 · 100 real failures, diagnosed + auto-fixed\n$ npx vectalon rnbench                      # 6 · the RN engineering benchmark, 8 dimensions\n$ npx vectalon rnbench --export ./bundle  # 6 · run a competitor through the same protocol\n$ npx vectalon bench --scenarios ./my-evals --references ./my-refs  # your own eval pack