Vendune
docs/testing.mdView on GitHub ↗

Testing, coverage and maintenance contracts

The prototype has automated behavior checks, actual PostgreSQL/Qdrant integration (plus legacy conversion checks), original Shopware comparisons and partial Lean contracts. It is not fully tested at 100%, and neither coverage nor the Lean subset proves the entire system bug-free.

Current acceptance (9 October 2026)

Complete change and end-to-end record joins the source owners, all registered suites, actual browser/private purchase checks and remaining audit work. The completed merged-Core CI at 598bbec measures Rust 90.55% lines / 87.94% functions / 87.24% regions, frontend 63.23% lines / 54.78% branches / 51.14% functions / 61.94% statements and Python tooling/reference scope 81.34% lines / 57.80% branches. Exact source/CI artifact summary.

Dated measurements below retain their original source. The current suite has 187 Rust unit tests and 388 frontend tests; the private Experience application and original Storyfront container have separate evidence and no claimed percentage.

Source architecture

  • Rust: 414 source modules (merged cognitive release, 9 October 2026), each with a responsibility header, at most 320 lines; main.rs at most 120. Existing domain folders remain independent of extension app implementations.
  • Frontend: separate admin/, storefront/, platform/ and shared/ ownership. Views/controllers and styles are limited to 400 lines; locale data has a documented 700-line allowance. Runtime cycles, unresolved local imports, crossing application boundaries, missing folder contracts and undocumented source files fail CI.
  • Bundled Email/GA4/Gmail/Slack services are Rust modules in src/connectors/; independent app examples and browser SDKs remain under extensions/. Python test tooling lives under scripts/ and the previous connector is an archived comparison oracle in reference/. The generated inventory covers runtime/tooling sources and is checked for drift.
  • Studio workspaces load lazily. Root application routing is isolated in frontend/src/application/; error boundaries keep workspace failures inside the current view and provide reload recovery for rejected cached module imports. The removed CommerceManager had no call site and duplicated old operational UI. Order state management remains in admin/orders/OrderWorkflow.tsx; product review moderation lives in admin/catalog/ReviewModeration.tsx.

frontend/tests/architecture.mjs also inserts deliberately broken synthetic imports, a cycle, an oversized file and absent documentation into temporary directories; it verifies that the architecture checker actually rejects each defect.

Repeatable checks

From the repository root:

cargo fmt --check
cargo clippy --locked --all-targets -- -D warnings
cargo test --locked
cargo build --locked --bins
npm --prefix frontend ci
npm --prefix frontend run format:check
npm --prefix frontend run build
npm --prefix frontend run architecture
npm --prefix frontend run localization
npm --prefix frontend run test:coverage
python3 scripts/formal.py
python3 scripts/formal/mutations.py
QDRANT_URL=http://127.0.0.1:16333 python3 scripts/verify_integration.py

Local integration creates and removes its own uniquely named synthetic database; existing shops remain outside that run. The default database container is vendune-postgres-1; override it with --container. It uses DATABASE_URL from the process or the local .env. CI alone passes --existing-database for its already-disposable database. Failures remain failures, and child processes stop before database cleanup. SIGINT flushes optional Rust coverage profiles.

scripts/testing/suites.json is the single registry for 48 HTTP suites, 17 local provider/connector/integration suites, four browser contracts and three verification-tool commands (72 registered suites). The local server uses an offline model URL; live model checks are separate, opt-in checks. Credentials, payments, mail and Slack are exercised against loopback protocol fixtures. Passing these does not demonstrate a real provider account.

The currencies suite exercises channel restrictions, Store API/MCP/UCP agreement, fixed and converted prices, shipping/coupon conversion, stale checkout reviews, immutable orders and grouped revenue. payment_providers tests create/capture/refund with EUR/USD, zero-decimal JPY and three-decimal KWD against a local protocol fixture. production_foundations denies cross-tenant currency-job SQL access using non-owner roles, then restarts a worker after its first 100-product batch and verifies completion, exact midpoint rounding and exclusion of products created after the job snapshot. See multi-currency commerce for the configuration and proof boundaries.

The original-PHP comparisons and reflected automation catalog checks are:

python3 scripts/differential.py
python3 scripts/context_differential.py
python3 scripts/delivery_differential.py
python3 scripts/rule_differential.py
python3 scripts/automation_registry.py
python3 scripts/automation_differential.py
python3 scripts/money_boundary_differential.py

The automation suite also executes scripts/playground.py as documented: a new owner-owned shop, repeat setup after an edited rule, both checkout branches and actual invoice records. tooling_tests.py has 12 tests, including refusal of remote credential destinations and unsafe state-file symlinks. See the manual tour for the corresponding browser steps.

Production boundary regressions

production_foundations runs two non-owner, non-bypass PostgreSQL replicas and repeats native checkout, CRM and adversarial tenant checks. It rejects unsafe TRUNCATE grants, omitted/forged tenant predicates, stale operator quotas and corrupted stock snapshots; it tests commit-only cache eviction and daily AI usage shared across live/staging. The 32-client scalability fixture explicitly configures 32 tenant permits. Production defaults and saturation policies remain configurable. The architecture guide explains exactly what these checks and the Lean subset establish.

Component regressions

frontend/tests/unit/ uses real components/hooks/transports, synthetic domain fixtures, jsdom and Testing Library. Fetches fail by default unless explicitly supplied by a test. Cases include:

  • merchant authentication, chat failures, language changes, uploads and separate live/staging transports;
  • navigation through permission-filtered lazy Studio workspaces and a signed-out login path;
  • catalogue filtering, cursor pages, stale responses, customer-session expiry and forbidden operation replay;
  • independent default addresses, linked customer/order/product back paths, configured translated groups, lazy history and confirmed revision-bound restoration with retained failed drafts;
  • opt-in personalization, ranking, consent withdrawal and failed signals;
  • product gallery, variants, normalized purchase quantities, tier prices, SEO, reviews and public source-backed product questions;
  • delayed product answers after navigation and product-specific moderation with rejected writes;
  • connected flow branches/deletion, frozen execution traces, incomplete JSON drafts and source condition round-trips in four languages;
  • SMTP/Resend/SendGrid fields, saved secrets, read-only settings, translated templates and explicit dry-run test calls.

Coverage collection

Frontend coverage includes every src/**/*.ts and src/**/*.tsx, including modules never imported by a test. HTML, JSON, LCOV and summaries go to frontend/coverage/. Styles and independent guest app bundles are outside that statement metric and explicitly listed as separate scopes.

For Rust install cargo-llvm-cov 0.9.1 and llvm-tools-preview. Use the same instrumented binaries for units, original-PHP comparisons and the actual HTTP integration run:

export CARGO_LLVM_COV_TARGET_DIR="$PWD/target"
cargo llvm-cov clean --workspace
eval "$(cargo llvm-cov show-env --sh)"
cargo test --locked
cargo build --locked --bins
# Run the integration and comparison commands above, then:
mkdir -p artifacts/coverage
cargo llvm-cov report --json --output-path artifacts/coverage/rust.json
cargo llvm-cov report --html --output-dir artifacts/coverage/rust

Mutation variants write profiles outside the real-core report directory; deliberately broken code cannot inflate the production coverage metric.

The Rust report measures lines, functions and regions, not branch coverage. The JSON preserves every module in this instrumented scope. Do not call covered regions exhaustive branch coverage.

Python instrumentation is a development dependency, not a production service dependency. It collects subprocesses and includes untouched namespace modules:

python3 -m pip install -r scripts/testing/requirements.txt
COVERAGE_PROCESS_START="$PWD/scripts/testing/.coveragerc" PYTHONPATH="$PWD/scripts/testing" python3 scripts/verify_integration.py
python3 -m coverage combine --rcfile=scripts/testing/.coveragerc
python3 -m coverage json --rcfile=scripts/testing/.coveragerc
python3 -m coverage html --rcfile=scripts/testing/.coveragerc
python3 scripts/testing/coverage_report.py

The report covers Python app/service sources and verification tools, so its percentage must not be presented as coverage of the Python services alone. Before a fresh standalone Python measurement, remove only the prior synthetic artifacts/.coverage* data, to avoid mixing historical executions.

CI uploads complete reports and lists zero-hit modules and unmeasured scopes. coverage-policy.json locks the measured regression floors; Studio transport, server-health and workspace-boundary modules require 100% lines, branches, functions and statements. Adding a source file counts in the denominator. Raising those floors is an explicit reviewed change; do not silently lower them to obtain a green check.

Explicit 100% audit and open work

npm --prefix frontend run test:coverage:full
python3 scripts/testing/coverage_report.py --require-full

These commands currently fail, as they should. The second also rejects a full system claim while browser SDK/example-app JavaScript and Wasm instruction coverage remain unmeasured. CSS requires visual behavior checks; live OAuth credentials, external inference and production deployment need their own verification.

The initial measured snapshot is in quality-baseline.json. The largest remaining frontend gaps include order-detail edits/documents, customer/address edge cases, full operator workflows, app iframe errors and complex rule/flow interactions. Coverage HTML shows individual missed statements and branches, not only totals. Future changes must add behavior regressions in these owning modules and increase coverage; a passing threshold is not completion of the 100% objective.

Lean source review locks and mutation checks continue in CI unchanged in scope; see formal verification for precisely what is proved.

The account composition regression exercises both login and registration through CustomerAccount → CustomerSignIn → shopApi. The simulated transport rejects protected profile/address/order reads unless the form persisted the exact tenant session key. This reproduces the former React-reserved key prop bug; backend address/customer isolation is separately exercised over real HTTP.

Studio session regressions reproduce a product 401 after a previously loaded overview, verify that all private controller state is cleared, and reconnect with a renewed token. Shared merchant calls and focus revalidation exercise the same path. Negative cases cover 403/5xx, customer/provider/operator errors, invalid login attempts, old-token replies after reauthentication, coalesced reads, hidden tabs, overlapping checks and timer/listener cleanup. Browser verification uses the existing personal test account; no expiry bypass or new privileges are added.

International configuration regression scope (2026-10-05)

New PostgreSQL suites cover complete bundled geography, custom countries/regions, US destination taxes in actual product/cart/order paths, saved-rule conditions, non-English main-language inheritance and stale/foreign configuration rejection. Translation fixtures cover Ollama/OpenAI/Claude wire formats, durable full-catalogue drafts, cold restart, stale apply, repeated apply and permission/tenant denial. They perform no paid model calls. React tests exercise keyboard/group country selection, missing-field inheritance and a shared settings draft across navigation. UI screenshots supplement these tests; neither replaces whole-source coverage.

Recorded international-commerce verification

CI run 37285967720 verified the international-commerce tree merged as 852f3d1 on 5 October 2026: 67 Rust unit tests, 90 frontend tests and the registered 25 HTTP / 3 provider / 4 browser-contract suites. The complete measured report, including untouched modules and unmeasured scopes, is checked into quality-baseline.json.

Scope Lines Other measured metrics
Rust 81.00% Functions 78.41%; regions 77.84%
Frontend 42.28% Branches 36.34%; functions 32.49%
Python 74.11% Branches 52.78%

These percentages retain the separate source scopes above and do not establish 100% coverage or whole-system correctness. Regression floors were preserved.

Company profile HTTP/database checks run in the suite registry as company_settings; the scope, decoder/ACL, issuer snapshot, translation and selective release cases are detailed in company settings. They use synthetic tenants and local image bytes, with no external payment/model/OAuth traffic.

Channel settings and media workspaces (2026-10-05)

See the settings/media guide for the current single-language editor, field-level checkout overrides, dependency-safe method removal, gallery and optional private image jobs. These are native prototype extensions; they do not establish additional full Shopware API/DAL parity, current tax law, paid-provider quality or whole-system formal certification.

App library regression scope (2026-10-05)

Eight component regressions exercise translated search, categories/status/discovery, installation followed by the real detail page, cancellation and failed activation, viewer restrictions, registered native interfaces, locale inheritance and independent artwork failure fallbacks. The isolated PostgreSQL app suite checks metadata round trips, tenant separation, immutable same-version digests, unsafe artwork refusal and the actual shop main language. Strict model-schema regression keeps optional presentation compatible with both provider request formats.

The app-library clean frontend run passed 137 tests: 50.16% statements, 43.76% branches, 41.35% functions and 50.98% lines across all included source. These measurements are not whole-system 100% coverage. Actual browser checks supplement component tests with desktop and 375px layouts. The affected apps, app surfaces, App Studio and developer-document HTTP suites pass using local fixtures, without paid calls.

Connected knowledge workspace (2026-10-05)

The current clean frontend run passes 151 tests across 26 test files, including ten knowledge regressions: whole-shop census, a single inherited-language editor, explicit empty values, non-English main language, guarded publication, viewers, retrieval scope, failed edits, actual product navigation and recommendation review. Across all included source the run measures 52.16% statements, 45.88% branches, 43.10% functions and 52.99% lines. This is a measured frontend scope, not a whole-system coverage update or a 100% claim.

84 Rust unit tests, strict lint/format, the extraction/Lean gates and negative mutations pass. The affected isolated PostgreSQL/AGE suites (knowledge_workspace, developer_documents, apps, staging, connected_apps, automation) exercise source hashes, public/private/archive fences, enabled-language retrieval, >50-source pagination, optimistic revisions, graph reassignment, granular HTTP/MCP access and selective release. Local provider fixtures prove private document passages reach the actual merchant model prompt and disappear after archive. No paid provider calls are used.

The knowledge suite also follows source ingestion through a product-scoped event rule into a durable completed flow without an order. Four UI regressions retain the stable trigger IDs while translating their labels EN/DE/FR/ES.

Browser checks cover inherited source creation in App Studio Lab, private customer exclusion versus merchant retrieval, the real product preview, desktop and 375px layouts. See the knowledge guide for source ownership, HTTP/MCP contracts, screenshots and the fixed-model/causality boundaries.

Adversarial tenant isolation

The registered tenant_isolation HTTP suite uses two synthetic workspaces, real merchant/customer sessions, direct MCP/UCP calls, foreign IDs, forged scope headers, same-ID app records and revoked one-workspace integration keys. It re-reads victim state after denied mutations. The database helper tests all 22 new composite foreign keys with own/foreign controls, scans tenant foreign keys for missing scoped counterparts, and tests forced app RLS under a restricted role including context reset on the same connection. It runs automatically via the existing CI integration registry, with zero provider calls.

Exact scope and remaining core-RLS boundary. Grouped checks are behavioral evidence, not 100% endpoint/authorization coverage.

Checkout verification (6 October 2026)

The checkout change passes 226 frontend tests across 34 files (including twelve checkout regressions) and 96 Rust unit tests. The preceding 75bd24a whole-source frontend coverage snapshot measured lines 60.78%, statements 59.70%, branches 53.50% and functions 49.32%. These figures are frontend-only and do not replace the older full-stack baseline report.

Real isolated SQL/HTTP suites verify review admission (five checks), native payment wire behavior (ten checks) and tenant isolation (38 checks). The browser saved a synthetic order after explicit review; 390-pixel emulation had no checkout content overflow. No paid model/PSP calls were used. The public Studio's origin fix was verified separately with allowed-origin 200, foreign-origin 403 and a fresh merchant-owned test shop with products, eight discoverable apps and browser MCP.

Platform control-plane regressions (6 October 2026)

platform.py runs platform_control.py against disposable PostgreSQL shops and a local synthetic model service. Sixteen combined checks include independent operator grants, two-shop/fresh-process encrypted provider inheritance, stale revisions, pause/background deferral, reversible trash, Host-based routing and persisted measured-call latency. No external model or payment call is made.

shop-scope.test.ts covers public subdomain upgrades and protected/local URL exceptions. platform-control.test.tsx covers independent entry links, write-only key editing and typed-ID/reason confirmation before trash. The complete frontend unit run passed 38 files / 261 tests; Rust passed 100 tests. Affected provider, transaction/payment, app/flow/staging, translation/image/document/search and tenant-isolation integration suites also run through their real local paths. These checks are not live provider capability or a full-system security proof.

managed_search.py allocates its own loopback port, so an open local Studio fixture cannot accidentally satisfy its readiness check. Qdrant-backed suites require the actual configured local Qdrant address.

Core hardening regressions (8 October 2026)

The existing strict-runtime suite now checks persistent login backoff across two processes, poison-event savepoint isolation and recovery, referenced/unsettled retention, and database-wide admission saturation/expiry. transaction_pooler repeats real checkout, customer registration/login, CRM, app/staging and RLS paths through one actual PgBouncer backend; listeners stay on a direct runtime connection. This exposed and removed nested pooled borrows under checkout/login locks. CI installs PgBouncer 1.21+ and uses the same suite registry. Native execution may select TEST_PSQL and TEST_PGBOUNCER; all databases and pooler processes are isolated fixtures and cleaned up. Unknown route rights and trusted Principal extraction also have unit/static controls. Implementation and boundaries.

These checks establish those behaviors, not 100% whole-system coverage, public rollout status, a speed-up percentage or protection against a compromised server.

App-platform v2 verification

Registered PostgreSQL suites exercise app_consent, app_callbacks, app_preview, app_relations, app_schema, app_jobs, app_components, app_assets, app_secrets, app_events, app_export and app_distribution. They observe installed versions and actual records/DDL, denied foreign mutations, lease fences, provider signatures, replay and export bytes. Existing apps, services, staging, connectors and automation suites verify compatibility.

Frontend tests exercise native code-behind, action allowlists, actor draft restore, revision conflicts, iframe bundle rejection, controls and package consent. These are regression evidence, not a claim of 100% app/host coverage. Full semantic document tests additionally require the real Qdrant service; a missing retrieval service must be reported as an unverified environment dependency, never replaced by a mocked successful semantic result. See the security matrix.