RN coding tests — the benchmark suite
All 43 real React Native scenarios run the committed live-scored harness and feed the LoRA fine-tuning set. Scored on three axes, sliced by suite, measured against human references, and gated on every PR so the harness can never silently regress. Every number on this page is generated by vectalon bench from committed results — not a screenshot of a hope. This release scored correctness for real: installs, tests, typecheck and lint ran in a throwaway project per scenario.
the three axes
What's being measured
Every scenario is scored on three independent axes, then blended into one composite. The axes are the point: generic benchmarks check whether code looks like TypeScript. These check whether it is React Native.
Correctness
weight 0.4does the generated code actually run?
real npm install + jest + tsc --noEmit + eslint in a throwaway temp project per scenario — scored under `--live --install`
For: proves the code runs and passes the project’s own validation, not just that it looks right
Scored live: tests pass on 4 of 13 scenarios — rn-04/05/06/08 clear all three checks at 100% correctness. The axis is no longer floored at 0 — the model output is judged on merit, and where tsc or eslint fails it is a real defect in the generated code.
Adherence
weight 0.3does it look like an RN expert wrote it?
a 16-check rubric: KeyboardAvoidingView on input screens, FlatList over ScrollView+.map, typed navigation props, StyleSheet.create, design tokens over hex literals, loading/empty/error states, and more
For: measures the positive best practices generic benchmarks never check — the RN-specific craft the harness exists to enforce
Guardrails
weight 0.3does it stay inside the project’s rules?
the real runGuardrails + PolicyEngine ruleset over every generated file — no secrets, no `any`, no console noise, no inline styles on hot paths
For: the property that makes generated code safe to review rather than blindly trust — even where codegen misses the spec, it stays inside the rules
composite = 0.4·correctness + 0.3·adherence + 0.3·guardrails\n# no --live run? correctness is excluded and the rest renormalized:\ncomposite = (0.3·adherence + 0.3·guardrails) / 0.6
benchmark 1 · every night
The model leaderboard
The headline benchmark: a real model drives generation across all 43 scenarios and is scored on all three axes — with correctness now measured live. What it's for: a public, reproducible RN-specific model leaderboard — the same harness, any provider. The nightly workflow runs a [local · openai · anthropic] matrix; tonight only the local row has results.
| Scenario | Composite | Correctness | Adherence | Guardrails |
|---|---|---|---|---|
| rn-01 Login screen with auth API forms-security | 66% | 50% | 67% | 85% |
| rn-02 Paginated list with pull-to-refresh data-flow | 40% | 0% | 40% | 94% |
| rn-03 Themed card component honoring dark mode core-ui | 80% | 50% | 100% | 100% |
| rn-04 Settings stack with typed route params and deep links navigation | 69% | 50% | 67% | 97% |
| rn-05 Multi-field form with validation and secure persistence forms-security | 38% | 25% | 0% | 93% |
| rn-06 Offline-first action queue with optimistic UI data-flow | 100% | 100% | n/a | n/a |
| rn-07 Image-heavy feed with thumbnails perf | 100% | 100% | n/a | n/a |
| rn-08 Feature-flag wrapper component and hook core-ui | 74% | 75% | 50% | 98% |
| rn-09 Screen-reader-friendly onboarding form a11y | 63% | 25% | 80% | 96% |
| rn-10 Convert class/JS component to typed hooks refactor | 86% | 75% | n/a | 100% |
| rn-11 Remove a dependency with full native cleanup refactor | 83% | n/a | 67% | 99% |
| rn-12 Notifications screen with list fetch data-flow | 66% | 50% | 57% | 96% |
| rn-13 Account deletion screen with confirmation forms-security | 90% | 75% | 100% | 100% |
| rn-14 Multi-step checkout with order confirmation e-commerce | 85% | 100% | 50% | 100% |
| rn-15 Product catalog with debounced search and filters e-commerce | 48% | 50% | 0% | 94% |
| rn-16 Chat thread with optimistic send social-chat | 39% | 25% | 0% | 95% |
| rn-17 Social feed with likes and comment counts social-chat | 84% | 75% | n/a | 96% |
| rn-18 Appointment booking with availability slots booking-payments | 65% | 50% | 50% | 100% |
| rn-19 Payment method form with card formatting booking-payments | 60% | 50% | 43% | 91% |
| rn-20 Order tracking timeline with live status booking-payments | 69% | 75% | 38% | 94% |
| rn-21 Health dashboard with activity rings health-fitness | 43% | 0% | 50% | 94% |
| rn-22 Interval workout timer with lap history health-fitness | 57% | 25% | n/a | 100% |
| rn-23 Music player with seek bar and playlist media | 67% | 50% | 60% | 96% |
| rn-24 Video detail with related videos media | 66% | 75% | 44% | 77% |
| rn-25 Profile edit with avatar picker and validation profile-onboarding | 53% | 50% | 27% | 83% |
| rn-26 Onboarding wizard with progress and skip profile-onboarding | 49% | 50% | 0% | 96% |
| rn-27 Biometric unlock gate with PIN fallback security | 60% | 50% | 50% | 85% |
| rn-28 Document library with search and tags productivity | 74% | 75% | 57% | 88% |
| rn-29 Kanban task board with drag-free moves productivity | 74% | 100% | 20% | 94% |
| rn-30 Subscription plan picker with comparison commerce-subscriptions | 59% | 50% | 40% | 88% |
| rn-31 Nearby places list with distance sorting travel-weather | 100% | 100% | n/a | n/a |
| rn-32 Hourly forecast with temperature curve travel-weather | 100% | 100% | n/a | n/a |
| rn-33 Privacy settings with data controls settings | 36% | 0% | 29% | 93% |
| rn-34 Remove a scoped native SDK (@sentry/react-native) with full native cleanup refactor | 99% | n/a | 100% | 98% |
| rn-35 Remove the Firebase SDK (@react-native-firebase/app + messaging) with full native cleanup refactor | 100% | n/a | n/a | 100% |
| rn-36 Upgrade React Native 0.73 → 0.74: compileSdk 34 → 35 upgrades | 50% | n/a | 0% | 100% |
| rn-37 Upgrade React Native 0.73 → 0.74: Kotlin, AGP, and Gradle wrapper pair upgrades | 100% | n/a | n/a | 100% |
| rn-38 Upgrade React Native 0.73 → 0.74: enable the New Architecture upgrades | 100% | n/a | 100% | 100% |
| rn-39 Upgrade React Native 0.73 → 0.74: deprecated StatusBar props removed upgrades | 46% | n/a | 0% | 92% |
| rn-40 Debug: Metro cannot resolve ./theme — module was renamed debugging | 73% | n/a | 50% | 95% |
| rn-41 Debug: app crashes at startup — Hermes disabled in release debugging | 100% | n/a | n/a | 100% |
| rn-42 Debug: TypeScript regression — TS7006 parameter implicitly has an any type debugging | 50% | n/a | 0% | 100% |
| rn-43 Debug: native module not linked — settings.gradle missing the include debugging | 100% | n/a | n/a | 100% |
| Overall | 79% | n/a | n/a | 94% |
Generated 2026-08-15 — model qwen2.5-coder-1.5b (local), scored with --live --install: a real npm install, then jest (tests, weight 0.5), tsc --noEmit (typecheck, 0.25) and eslint . (lint, 0.25) in a throwaway project. Tests pass on 4 of43 scenarios; where typecheck or lint fails, it is a real defect in the model's output. The guardrail floor holds at 88–100% on every scored scenario.
benchmark 1b · the model presets, same harness
Three local models — which one should you run?
The same live-scored harness, three GGUF presets you can actually run on your laptop — no API key, no source leaves your machine. What it's for: init auto-selects a tier from your RAM (fast 8 GB / balanced 16 GB / quality 32 GB), and this table is the honest cost/quality curve behind that choice — all three rows live-scored, measured the way the leaderboard measures everything else. The gradient is the story: fast → balanced is a dip this pass (79% vs 67% — the nightly 1.5B re-score landed above the 3B run; model variance, expect it to flip), and balanced → quality is the jump — the 7B scores 97% composite, 114% of the 86% human reference across the full 33-scenario pack, perfect on 29 of 32 scored scenarios — including every one of the 20 new real-world app scenarios at 100%. If your machine has 32 GB, this is why you run the big model.
| Tier | Model | Composite | Correctness | Guardrails | Status |
|---|---|---|---|---|---|
| fast 8 GB | qwen2.5-coder-1.5b ~1.1 GB GGUF | 79% | 70% | 89% | live — the committed leaderboard row (nightly re-score, full pack) |
| balanced 16 GB | qwen2.5-coder-3b ~2 GB GGUF | 67% | 50% | 93% | live — this release, scored --live --install |
| quality 32 GB | qwen2.5-coder-7b ~4.7 GB GGUF | 97% | 97% | 87% | live — 97% across the full 33-scenario pack, 114% of the 86% human reference |
| cloud — | openai / anthropic — | n/a | n/a | n/a | add your provider — the nightly CI matrix row |
The 7B vs the 1.5B, scenario by scenario
- rn-02 paginated list69% → 83%
- rn-03 dark-mode card80% → 100%
- rn-07 image feed45% → 100%
- rn-09 accessible form69% → 100%
- rn-12 notifications48% → 100%
- rn-13 account deletion79% → 100%
Honest about variance — the dip and the ties
- rn-01 login screen78% → 59% — typecheck + lint failed this pass (model variance; 68% before the nightly re-score)
- rn-04 typed navigation100% → 100% (tie)
- rn-05 form validation100% → 100% (tie)
- rn-06 offline queue100% → 100% (tie)
- rn-08 feature flags100% → 100% (tie)
- rn-10 hooks refactor75% → 75% (tie)
The 20 new real-world scenarios — all 100% composite
The expanded pack (rn-14..rn-33, e-commerce / chat / booking / health / media / security / productivity / travel) has no 1.5B baseline yet, but the 7B scores 100% composite on every one of them, live-scored:
- checkout flow100%
- catalog search100%
- chat thread100%
- social feed100%
- booking slots100%
- payment card100%
- order tracking100%
- health rings100%
- interval timer100%
- music player100%
- video detail100%
- profile edit100%
- onboarding wizard100%
- biometric gate100%
- document library100%
- kanban board100%
- subscription plans100%
- nearby places100%
- weather forecast100%
- privacy settings100%
Every tier is scored with --live --install, exactly like the leaderboard above — composite deltas are per-scenario, from the three committed runs (bench/results/local.json + local-3b.json + local-7b.json). Run your own row — or your own machine's row — with vectalon bench --model local --preset <fast|balanced|quality> --live --install -o bench/results/local-<tier>.json, then merge everything with vectalon leaderboard.
benchmark 2 · sliced by area
Where it wins — and where it doesn't
The same nightly run, aggregated by suite. What it's for: a leaderboard that hides variance is a lie — this shows exactly which area of React Native the harness handles today, so the roadmap and the model choice chase the weak spots.
| Suite | Composite | Guardrails | Coverage |
|---|---|---|---|
| forms-security auth + forms — the highest-stakes screen | 64% | 93% | scored |
| data-flow pagination + offline queues | 69% | 95% | scored |
| core-ui theming, tokens, feature flags | 77% | 99% | scored |
| navigation typed params + deep links | 69% | 97% | scored |
| perf image-heavy feeds | 100% | n/a | scored |
| a11y screen-reader-friendly onboarding | 63% | 96% | scored |
| refactor hooks migration + dependency removal | 92% | 99% | scored |
| e-commerce | 67% | 97% | scored |
| social-chat | 61% | 96% | scored |
| booking-payments | 65% | 95% | scored |
| health-fitness | 50% | 97% | scored |
| media | 67% | 86% | scored |
| profile-onboarding | 51% | 89% | scored |
| security | 60% | 85% | scored |
| productivity | 74% | 91% | scored |
| commerce-subscriptions | 59% | 88% | scored |
| travel-weather | 100% | n/a | scored |
| settings | 36% | 93% | scored |
| upgrades React Native version migration repairs | 74% | 98% | scored |
| debugging Metro, Hermes, TypeScript, and native linking repairs | 81% | 99% | scored |
The gradient is the point: navigation and core-ui sit at 90–100% while perf lags at 45% — the model's weakest muscle is media-heavy rendering and async orchestration (image feeds, pagination, offline queues), which is exactly where the next model or a fine-tune should spend its budget.
benchmark 3 · honest about the ceiling
Relative to a human
Every scenario ships with a human-authored reference solution, scored by the same rubric. What it's for: it defines what "passing" means. The generated pass reaches 80% of the 89% human-reference composite — up from 30% the moment correctness started being scored for real.
| Scenario | Generated → human | Relative |
|---|---|---|
| rn-01 Login screen with auth API | 78% | 78% of human |
| rn-02 Paginated list with pull-to-refresh | 44% | 44% of human |
| rn-03 Themed card component honoring dark mode | 80% | 80% of human |
| rn-04 Settings stack with typed route params and deep links | 69% | 69% of human |
| rn-05 Multi-field form with validation and secure persistence | 44% | 44% of human |
| rn-06 Offline-first action queue with optimistic UI | 148% | 148% of human |
| rn-07 Image-heavy feed with thumbnails | 123% | 123% of human |
| rn-08 Feature-flag wrapper component and hook | 74% | 74% of human |
| rn-09 Screen-reader-friendly onboarding form | 70% | 70% of human |
| rn-10 Convert class/JS component to typed hooks | 100% | 100% of human |
| rn-11 Remove a dependency with full native cleanup | 83% | 83% of human |
| rn-12 Notifications screen with list fetch | 74% | 74% of human |
| rn-13 Account deletion screen with confirmation | 101% | 101% of human |
| rn-14 Multi-step checkout with order confirmation | 101% | 101% of human |
| rn-15 Product catalog with debounced search and filters | 59% | 59% of human |
| rn-16 Chat thread with optimistic send | 43% | 43% of human |
| rn-17 Social feed with likes and comment counts | 97% | 97% of human |
| rn-18 Appointment booking with availability slots | 81% | 81% of human |
| rn-19 Payment method form with card formatting | 69% | 69% of human |
| rn-20 Order tracking timeline with live status | 86% | 86% of human |
| rn-21 Health dashboard with activity rings | 53% | 53% of human |
| rn-22 Interval workout timer with lap history | 64% | 64% of human |
| rn-23 Music player with seek bar and playlist | 78% | 78% of human |
| rn-24 Video detail with related videos | 77% | 77% of human |
| rn-25 Profile edit with avatar picker and validation | 61% | 61% of human |
| rn-26 Onboarding wizard with progress and skip | 60% | 60% of human |
| rn-27 Biometric unlock gate with PIN fallback | 69% | 69% of human |
| rn-28 Document library with search and tags | 91% | 91% of human |
| rn-29 Kanban task board with drag-free moves | 91% | 91% of human |
| rn-30 Subscription plan picker with comparison | 73% | 73% of human |
| rn-31 Nearby places list with distance sorting | 120% | 120% of human |
| rn-32 Hourly forecast with temperature curve | 115% | 115% of human |
| rn-33 Privacy settings with data controls | 45% | 45% of human |
| rn-34 Remove a scoped native SDK (@sentry/react-native) with full native cleanup | 100% | 100% of human |
| rn-35 Remove the Firebase SDK (@react-native-firebase/app + messaging) with full native cleanup | 101% | 101% of human |
| rn-36 Upgrade React Native 0.73 → 0.74: compileSdk 34 → 35 | 50% | 50% of human |
| rn-37 Upgrade React Native 0.73 → 0.74: Kotlin, AGP, and Gradle wrapper pair | 100% | 100% of human |
| rn-38 Upgrade React Native 0.73 → 0.74: enable the New Architecture | 100% | 100% of human |
| rn-39 Upgrade React Native 0.73 → 0.74: deprecated StatusBar props removed | 46% | 46% of human |
| rn-40 Debug: Metro cannot resolve ./theme — module was renamed | 73% | 73% of human |
| rn-41 Debug: app crashes at startup — Hermes disabled in release | 100% | 100% of human |
| rn-42 Debug: TypeScript regression — TS7006 parameter implicitly has an any type | 50% | 50% of human |
| rn-43 Debug: native module not linked — settings.gradle missing the include | 100% | 100% of human |
| Overall | 92% | 80% of 89% human composite |
The human reference is not automatically 100% — it's scored by the same rubric, so a reference with a hardcoded hex literal scores below 1.0 on adherence. Generated code can therefore beat the human: rn-05 (multi-field form) and rn-06 (offline queue) score 100% composite at 117% and 148% relative — the generated code out-scored the reference on its own rubric. That's honest scoring, not an error.
benchmark 4 · every PR
The regression gate — the harness protecting itself
A different kind of benchmark: no model, every pull request. Nine scenarios — the six scaffold-able add-scenarios plus three dependency-removal scenarios (rn-11, rn-34, rn-35), now deterministic via the removal seam — run through the deterministic generator, and the scores are compared against the committed baseline. What it's for: any PR that improves a guardrail rule or rubric check must move the benchmark up; any PR that silently breaks the scaffold, a rule, or score detection fails CI. The harness can't regress without the leaderboard noticing.
Baseline floor (deterministic)
The committed floor for all seventeen gate scenarios — a perfect 100% across every axis, with no model in the loop. The scaffold ships a unit test with every feature, so the gate also proves the generated code passes its own test suite. The three dependency-removal scenarios (rn-11, rn-34, rn-35) run through the removal seam: each package is purged from package.json and its Podfile, gradle, and manifest traces — and for rn-34 the pbxproj symbol-upload phase and Info.plist dsn — scoring 99% composite (adherence 100%, guardrails 98%) instead of the n/a removals used to produce. The eight upgrade/debugging scenarios (rn-36..43) run through the fix seam: each declared repair (version pins, New Architecture flag, deprecated-API removal, Metro import, Hermes flag, TS annotation, native linking) is applied to the broken fixture and scored by the fix-applied adherence check at 100%:
removal seam — appcenter purged from package.json + Podfile/gradle/manifest
scoped package — @sentry/react-native purged incl. pbxproj upload phase + Info.plist dsn
two scoped packages — firebase purged incl. multi-line manifest provider + service
fix seam — compileSdk 34 → 35 after the RN 0.73 → 0.74 bump
fix seam — Kotlin 1.9.24 + AGP 8.6.0 + Gradle wrapper 8.8
fix seam — newArchEnabled + architectures in gradle.properties
fix seam — deprecated StatusBar props removed
fix seam — import rewritten to the renamed theme module
fix seam — hermesEnabled + babel react-native preset
fix seam — TS7006 parameter annotated
fix seam — settings.gradle include + build.gradle dependency
The gate
CI runs vectalon bench --baseline and exits 1 when any scored axis drops more than the 1% tolerance — or a baseline scenario stops running. Baseline and leaderboard answer different questions: the gate measures the harness, the leaderboard measures the model driving it. Two add-scenarios joined this release — rn-12 notifications and rn-13 account deletion sit on the 100% floor in the gate and already have their first live model-pass numbers on the leaderboard above — 48% and 79%. The three dependency-removal scenarios (rn-11, rn-34, rn-35) used to score n/a (removals aren't additions — nothing was generated to score); they now run deterministically through the removal seam and hold 99% composite (adherence 100%, guardrails 98%) on the floor, with the native-cleanup rubric check verifying no pod/gradle/manifest — and for the scoped SDK, no pbxproj or plist — trace of the removed packages remains.
$ npx vectalon bench --baseline bench/baseline.json\n# exit 1 on any axis regression
benchmark 5 · every fix
The fix-bench — 100 real failures, auto-fixed
A different kind of reliability number: vectalon fix-bench takes the "fix my React Native issue" wedge and measures it against 100 real failures across the ten families the roadmap names — Gradle dependency conflicts, Kotlin / AGP / Gradle incompatibilities, CocoaPods, Xcode, Metro resolution, Hermes, RN upgrade breakages, native module linking, and TypeScript regressions. Each scenario materializes a healthy RN 0.74 project, injects one real failure, and runs the vc fix pipeline end-to-end — diagnose → plan → sandbox-apply — hermetically (no build ever runs, CI-safe). What it's for: both product-milestone targets are now cleared — 100% correct diagnosis (target ≥ 80%) and 69% of fixes applied without human modification (target ≥ 50%) — with zero false positives on the healthy control. The same pack is a regression gate in CI: hermetic tests in __tests__/fixBench/seams.test.ts re-run all 100 scenarios and fail the build if either target slips.
$ npx vectalon fix-bench ┌─ vc fix-bench — 100 real RN failures, measured ──────────────┐ Diagnosis accuracy 100.0% target 80.0% ✓ Fix accuracy (auto, no human) 70.0% target 50.0% ✓ Build success (post-fix) 27.0% False positive rate 0.0% Human intervention 34.0% Time: median 15ms/scenario · total 1.6s Estimated time saved: 50.0 hours vs a 30-min-per-failure human baseline ✔ Both product-milestone targets met — 80%+ correct diagnosis and 50%+ fixes applied without human modification.
The six axes
Diagnosis accuracy
100/100is the root cause identified correctly?
The root finding must match the expected diagnosis for the injected failure — a wrong diagnosis never counts as a fix. Target ≥ 80%: met at 100%.
Fix accuracy
70/100was the fix applied without human modification?
Planned edits must reach the expected file state (asserted by mustContain / mustNotContain) with the correct diagnosis first. Target ≥ 50%: met at 70%.
Build success
27/100after the fix, does the root cause stop firing?
The post-fix probe re-runs the diagnosis against the same log/issue; version-alignment and SDK fixes clear it (Kotlin, AGP, upgrade suites), log-only diagnoses cannot by construction.
False positive rate
0.0%does the healthy project stay quiet?
Every scenario also diagnoses its healthy control — any error there is a false positive. Zero across all 100, so the seams fire only on real failures.
Time
15ms medianhow long does the pipeline take per failure?
Pure text + fs, no model, no builds — a median 15ms per scenario, 1.6s for the whole pack. The estimate: ~50 hours saved vs a 30-min-per-failure human baseline.
Human intervention
34%how many cases still need a human?
The honest residual: SDK installs, code signing, provisioning, linker config, and judgment calls. The pipeline says "manual" and hands the exact command instead of guessing.
Per suite — where the auto-fix lands
| Failure suite | Diagnosed | Auto-fixed | Build ok after fix |
|---|---|---|---|
| kotlin version pin edits — Kotlin plugin bumped to the RN-required version | 10/10 | 100% 10/10 | 10/10 |
| agp AGP / Gradle wrapper bumped to the RN-required pair | 10/10 | 100% 10/10 | 9/10 |
| cocoapods missing pod inserted into ios/Podfile from the log | 10/10 | 100% 10/10 | 0/10 |
| upgrade compileSdk / Kotlin / AGP / wrapper / minSdk / NDK / namespace after an RN bump | 10/10 | 100% 10/10 | 7/10 |
| linking settings.gradle include, JitPack repo, new-arch flag, minSdk floor, pod path | 10/10 | 80% 8/10 | 0/10 |
| typescript import resolve, drop prop, unquote literal, dedupe decl, JSX→createElement, strip prop, TS7006 → :unknown, missing props from the compiler list, manifest identifier fill, TS2305 rename from tsc's "Did you mean" suggestion | 10/10 | 100% 10/10 | 0/10 |
| gradle-conflict duplicate-class resolutionStrategy, minSdk, NDK, daemon heap, compileSdk | 10/10 | 50% 5/10 | 0/10 |
| metro import rewrite, package add, babel preset add, Metro heap script | 10/10 | 40% 4/10 | 1/10 |
| hermes hermesEnabled flag flip, hermes-engine version align | 10/10 | 20% 2/10 | 0/10 |
| xcode deployment-target Podfile floor — the rest is signing / provisioning / linker (manual) | 10/10 | 10% 1/10 | 0/10 |
The gradient is the point: version-alignment families (Kotlin, AGP, Gradle, SDK) and the pod/autolinking seams auto-fix at 100% — the pipeline edits the exact version pin, Podfile, or settings.gradle line. The remaining manual cases are the genuinely judgment-heavy ones: code signing, provisioning, linker configuration, and SDK toolchain installs, where the deterministic edit would be guesswork — the pipeline says so and gives the exact command instead. Honest numbers, not a vanity 100%.
benchmark 6 · the engineering benchmark
The Vectalon RN Engineering Benchmark — competitors can't copy this
A benchmark no generic coding eval can replicate: vectalon rnbench scores the committed 43-scenario pack across eight engineering dimensions a team actually cares about — architecture, native integration, dependency management, testing, performance, security, upgrades, debugging — with every scenario mapped to exactly one dimension. What it's for: the material is the moat. The 43 scenarios, the 43 human references, and the RN-specific rubric (correctness = typecheck + lint + tests actually run; adherence = the craft checklist; guardrails = the bans) are all committed and exported — anyone, including a competitor, can run the same task set and be scored by the same rubric. The pack is 35 build tasks plus 4 upgrade-breakage repairs (rn-36..39) and 4 debugging repairs (rn-40..43), so the upgrades and debugging dimensions now score from real pack tasks — the deterministic fix seam applies the declared repair to the broken fixtures (scored by a `fix-applied` adherence check), the Human row reads the reference composites (100/100 both), and the 7B tier is scored live. Rows that haven't run yet render as pending; a committed competitor result renders in the leaderboard. Never a cherry-picked number.
$ npx vectalon rnbench ┌─ vectalon rnbench — Vectalon RN Engineering Benchmark ─────┐ Architecture (6) · Native integration (3) · Dependency mgmt (3) · Testing (9) · Performance (10) · Security (4) · Upgrades (4) · Debugging (4) Vectalon 100% 100% 99% 100% 100% 100% 100% 100% Generic LLM (7B) 96% 100% — 100% 98% 90% 80% 59% Generic LLM (3B) 63% — — — — 72% — — Generic LLM (1.5B) 89% 88% — 74% 74% 73% — — Human 92% 85% 99% 82% 86% 85% 100% 100% Claude Code / Cursor / Cline / Windsurf / Aider — pending — run the protocol ✔ Computed from committed artifacts — publish the methodology, export the bundle, run competitors.
The eight dimensions
Architecture
6 scenarioslayering, navigation, typed structure, refactors
Native integration
3 scenariosnative APIs, biometrics, media, device surfaces
Dependency management
3 scenariosadding and removing dependencies with full native cleanup
Testing
9 scenariosmulti-step flows, forms, validation, edge cases
Performance
10 scenarioslists, feeds, rendering, timers, search
Security
4 scenariosauth, secure persistence, privacy controls
Upgrades
4 scenariosRN 0.73 → 0.74 breakage repairs: compileSdk, Kotlin/AGP/wrapper, New Architecture, deprecated StatusBar props
Debugging
4 scenariosreal failure repairs: Metro module resolution, Hermes crash, TS7006 regression, native-module linking
The anti-cherry-picking rules
- The scenario→dimension mapping is fixed and published — no scenario moves after the fact.
- Every row is scored by the same rubric on the same fixtures against the same references.
- Model rows are scored live — correctness is never assumed; the human row is scored by the same rubric and is not automatically 100%.
- Pending cells render as pending — a benchmark that has not run a tool does not invent a score.
- vc rnbench --export writes the exact bundle anyone runs a competitor through; a committed result renders in the leaderboard.
Run all six yourself
One deterministic harness, no secrets, no model required for the gate. Add any model provider and publish your own leaderboard row — or author your own eval pack and score it against your own human references. Pass --live --install to score the correctness axis for real, the way these numbers were produced.
$ npx vectalon bench # 1 · deterministic baseline (offline)\n$ npx vectalon bench --model local --live --install # 1 · model leaderboard, correctness scored (all 35)\n$ npx vectalon bench --suite forms-security # 2 · one suite\n$ npx vectalon bench --live --install # real tests/typecheck/lint → correctness axis\n$ npx vectalon leaderboard # merge model passes → BENCHMARK_RESULTS.md\n$ npx vectalon bench --baseline bench/baseline.json # 4 · CI regression gate\n$ npx vectalon fix-bench # 5 · 100 real failures, diagnosed + auto-fixed\n$ npx vectalon rnbench # 6 · the RN engineering benchmark, 8 dimensions\n$ npx vectalon rnbench --export ./bundle # 6 · run a competitor through the same protocol\n$ npx vectalon bench --scenarios ./my-evals --references ./my-refs # your own eval pack