Skip to content

Field benchmarks (labs)

Transparency first

These numbers come from reproducible lab harnesses, not marketing claims. Each row links to the raw JSON artifact. We distinguish autonomous discovery (generic crawl + DEPTH-16) from lab-specific confirmation probes (known challenge solutions). When a target sanitizes output, we report OBSERVED, not CONFIRMED — we do not inflate execution.

Last updated: 2026-07-24 (UTC) · Known labs: data/bench_*.json via python scripts/regen_benchmarks_docs.py · Novel targets: GEN-01b harness (scripts/bench_novel_target.py) — metrics published after a real run (GEN-01d).

Summary table (known labs)

Lab / gate Metric (honest) Mode Gate Reproduce Raw JSON
Google Firing Range Reflected 15 browser-confirmed / 40 endpoints seen · DOM 3 confirmed (ceiling 1 stable on live lab) Autonomous discovery + DEPTH-16 dialog PASS python scripts/bench_firing_range.py --measure-only assets/benchmarks/firing_range.json
XSS Game 6/6 levels browser-confirmed (dialog) Lab-specific probes — not autonomous discovery PASS python scripts/bench_xss_game.py --measure-only assets/benchmarks/xss_game.json
webscantest.com Coverage 68% authenticated Docker-local (127.0.0.1); public host returns 403 to scanners PASS python scripts/bench_webscantest_auth.py --measure-only assets/benchmarks/webscantest_auth.json
OWASP Juice Shop (DEPTH-08) ? OBSERVED · 0 CONFIRMED Store→render pipeline; target sanitizes stored XSS FAIL python scripts/bench_admin_sinks.py assets/benchmarks/admin_sinks.json
Anti-FP corpus (BENCH-REAL) 0 FP · precision@confirmed 1.0 (29/29 cases) Fixture corpus (no live targets) PASS python scripts/bench_real.py --fixtures-only assets/benchmarks/real_report.json
Parabank (BENCH-02) parabank_recall probe PASS (public + local Docker) Targeted recall probe, not full-stream gate JSON PASS python scripts/benchmark_recall.py --include-labs BENCH-02 probe

What each number means

Firing Range — autonomous XSS depth

  • Reflected 15/40: DEPTH-16 Playwright dialog on distinct reflected paths from the seed probe surface (not the whole lab catalog).
  • DOM 3 confirmed, ceiling 1: live toxicdom matrix; innerHtml sink is intermittent. We do not claim DOM ≥10 on the public Firing Range — that requires a dedicated DOM-rich lab (Ola 3).

XSS Game — lab-specific confirmation (not discovery)

  • 6/6 counts only levels with browser dialog (CONF-MODEL-01).
  • Probes are preloaded per challenge (xss_game_lab.py) — this demonstrates cross-context confirmation (reflected, stored, DOM hash, attribute breakout, javascript: URI, gadget loader), not that the generic crawler discovers all six levels alone. Full pentest stream hit wall-clock timeout (timeout after 420.0s).

webscantest — authenticated coverage

  • 68% from scan_coverage.coverage_pct with verified session (auth=True on post-login paths).
  • Measured against Docker-local Rapid7 image; scanning the public webscantest.com hostname is blocked (403).

admin_sinks (DEPTH-08) — honest OBSERVED

  • Pipeline exercised: discover sinks → store payload → re-render → browser verify.
  • This Juice Shop build strips/escapes stored payloads → 0 CONFIRMED, ? OBSERVED. Gate PASS means the pipeline is honest, not that we faked execution.

Anti-FP (BENCH-REAL)

  • 29 negative fixtures (SPA catch-alls, JSON reflections, timing jitter, etc.) must produce zero CRITICAL/HIGH false positives.
  • precision_at_confirmed: 1.0 on the anti-FP corpus (no confirmed findings expected).

Parabank (BENCH-02)

  • Regression via parabank_recall (SQLi login bypass / session recall) in scripts/benchmark_labs.py.
  • No dedicated data/bench_parabank.json yet — PASS is from the BENCH-02 multi-lab probe harness.

Novel targets (GEN-01) — what a new app gets

Separate from known-lab gates

The rows above measure regression on labs we tune against (seed packs, lab-specific probes, or Docker-local harnesses). They do not prove depth on an arbitrary customer app. GEN-01 measures the surface-driven motor in clean mode — no seed pack, no VEX_RECALL_BENCHMARK, no lab profile.

Target Status Metric (honest) Reproduce
OWASP NodeGoat Published coverage 100% · CONFIRMED+ 1 · OBSERVED 1 · recall_core 3/3 · recall_extended 8/9 See below

How we measure novel targets

  • Harness: python scripts/bench_novel_target.py --target <url> [--auth user:pass]
  • Clean motor: VEX_DISCOVER_API_SKIP_SEEDS=1; recall-benchmark env vars unset.
  • Ground-truth: data/truth/nodegoat_truth.yaml — recall split into:
  • recall_core — only dast_detectable: yes (honest DAST denominator)
  • recall_extendedyes + partial (auth/heuristic-dependent)
  • out_of_scope_dastno entries listed separately (never inflates recall)
  • Anti-contamination: flags Meridian strings, lab_profile, recall_benchmark in meta/env.
  • Exit code 3 if target unreachable — no fabricated metrics.
git clone https://github.com/OWASP/NodeGoat.git /tmp/NodeGoat
cd /tmp/NodeGoat && docker compose up -d --build
python scripts/bench_novel_target.py \
  --target http://127.0.0.1:4000 \
  --auth user1:User1_123 \
  --json data/bench_novel_target.json
# Then: python scripts/regen_benchmarks_docs.py

NodeGoat — published run (2026-07-24)

  • Mode: autonomous surface-driven · VEX_DISCOVER_API_SKIP_SEEDS=1 · contamination CLEAN
  • Scan: 14 findings · coverage 100% · elapsed 680.2s (local Docker · Chromium OK)
  • Confidence: CONFIRMED+ 1 · OBSERVED 1 · HIGH+ 3
  • Recall core (dast_detectable: yes only): 3/3 — honest DAST denominator
  • Recall extended (yes + partial): 8/9 — includes two table-stakes-to-moderate generic checks labeled honestly: NG-A5 (Security Misconfiguration) via a missing-security-header finding (CWE-16), and NG-A2 (weak password policy) via a generic weak-password probe on the discovered registration form (CWE-521, OBSERVED — server accepts trivial passwords). The one remaining partial, NG-A7 (function-level access control), needs role comparison / 2-identity; interpret with care
  • Out of scope DAST (dast_detectable: no): 2 entries — not counted as motor failure
  • Caveats: playwright_crawl OK (13 pages, 6 DOM forms) · 6 injection point(s) on bus · discover_api: 0 API endpoint(s) · sources={} · recall_core 3/3 on NodeGoat truth (dast_detectable: yes) — surface-driven motor (GEN-02): SSJI eval timing, profile write XSS, open redirect · no seed pack · truth matching is keyword/CWE heuristic
  • Raw JSON: assets/benchmarks/novel_target.json

Finding adjudication (manual review)

  • XSS /profile (firstName): TP write-reflection — authenticated POST reflects unsanitized markup in the HTML response (admin_sinks write probe). HIGH confidence, not CONFIRMED (no browser dialog on render). Matches truth NG-A3-XSS-STORED (dast_detectable: yes).
  • SSJI /contributions: TP timing-confirmed — server-side eval busy-loop delay detected on contribution fields. CONFIRMED (matches truth NG-A1-SSJI).
  • XSS /signup (userName): TP reflejado — input echoed without encoding on validation error (NodeGoat session.js; ESAPI disabled in source). HIGH confidence (reflection detected), not CONFIRMED (no browser dialog; value= attribute context). Distinct from truth NG-A3-XSS-STORED (POST /profile). Tracked as NG-A3b-XSS-REFLECTED-SIGNUP (dast_detectable: partial).
  • Security misconfiguration (headers): TP — response missing hardening headers (e.g. X-Frame-Options). MEDIUM, generic (any app without these headers triggers it), tagged CWE-16. Matches truth NG-A5-MISCONFIG (dast_detectable: partial — "headers yes, internal config no"). Table-stakes finding; credited to extended recall but not equivalent to the injection-class hits.
  • Weak password policy (signup): TP — the discovered registration form accepts trivial industry-common passwords (123456, password, …) with no complexity block. OBSERVED, CWE-521, generic (probes any registration form found on the surface — no hardcoded path). Matches truth NG-A2-BROKEN-AUTH (dast_detectable: partial). Real product check, not a lab number; still labeled OBSERVED (acceptance detected, not full account-takeover chain).

GEN-01a — anti-FP on lab-flavored probes (internal, affects all scans)

Probes that historically smelled like Meridian tuning are now evidence-gated:

  • Password reset abuseCONFIRMED / password_reset_forged only after isolated re-login succeeds.
  • JWT confusion — forged sub from session JWT / observed email, not hardcoded lab identities.

This is not a separate public gate row; it protects novel-target honesty.

Enterprise scorecard (known labs)

Gate Status Notes
G1 Meridian recall 🟡 22/23 Stretch 23/23
G2 Firing Range 🟢 15/40 refl + DOM 3 (ceiling 1)
G3 XSS Game 🟢 6/6 lab-specific probes
G4 webscantest auth 🟢 68% Docker-local
G5 Honest reporting 🟢 RPT-09/10
G6 HIGH+ evidence 🟢 85.0% on lab corpus (scripts/bench_g6_evidence.py)
G7 jsonplaceholder anti-FP 🟢
G8 pytest 🟢 ≥2982
G9 Public benchmarks 🟢 This page
G10 Parabank BENCH-02 🟢 parabank_recall PASS
G11 OWASP taxonomy 🟢
G12 XSS browser-executed 🟢 DEPTH-16 + D4
G13 SAFETY-01 process 🟢 Documented in release checklist

Score: 12 / 13 green (92% — G6 measured 07-18). Yellow: G1 (Meridian recall stretch). Red: none.

What we do not claim

  • DOM XSS ≥10 on the live Firing Range — public lab ceiling is 1 stable sink(s); three dialogs were observed but only one is gated as stable.
  • Autonomous 6/6 on XSS Game — our 6/6 uses lab-specific solvers; autonomous metric is Firing Range reflected recall.
  • CONFIRMED stored XSS on sanitizing apps — Juice Shop admin_sinks correctly stays OBSERVED.
  • webscantest coverage on the public hostname — blocked; Docker-local only.
  • 100% recall on every OWASP class in every customer app — labs are regression targets, not guarantees.
  • Novel-target depth is measured separately (GEN-01) — see NodeGoat section below; not comparable to tuned lab gates.

Regenerate monthly

# Re-run benches (labs must be reachable / Docker up as noted), then:
./venv/bin/python scripts/bench_firing_range.py --json data/bench_firing_range.json
./venv/bin/python scripts/bench_xss_game.py --json data/bench_xss_game.json
./venv/bin/python scripts/bench_webscantest_auth.py --json data/bench_webscantest_auth.json
./venv/bin/python scripts/bench_admin_sinks.py --json data/bench_admin_sinks.json
./venv/bin/python scripts/bench_real.py --fixtures-only --json data/bench_real_report.json
# Novel target (GEN-01b) — optional; requires Docker NodeGoat up:
./venv/bin/python scripts/bench_novel_target.py \
  --target http://127.0.0.1:4000 --auth user1:User1_123 \
  --json data/bench_novel_target.json
./venv/bin/python scripts/regen_benchmarks_docs.py
cd docs-site && mkdocs build --strict