Testing, coverage and maintenance contracts
The prototype has automated behavior checks, actual PostgreSQL/Qdrant integration (plus legacy conversion checks), original Shopware comparisons and partial Lean contracts. It is not fully tested at 100%, and neither coverage nor the Lean subset proves the entire system bug-free.
Current acceptance (9 October 2026)
Complete change and end-to-end record joins the
source owners, all registered suites, actual browser/private purchase checks and
remaining audit work. The completed merged-Core CI at 598bbec measures Rust
90.55% lines / 87.94% functions / 87.24% regions, frontend 63.23% lines /
54.78% branches / 51.14% functions / 61.94% statements and Python tooling/reference
scope 81.34% lines / 57.80% branches. Exact source/CI artifact summary.
Dated measurements below retain their original source. The current suite has 187 Rust unit tests and 388 frontend tests; the private Experience application and original Storyfront container have separate evidence and no claimed percentage.
Source architecture
- Rust: 414 source modules (merged cognitive release, 9 October 2026), each with a responsibility header, at most 320 lines;
main.rsat most 120. Existing domain folders remain independent of extension app implementations. - Frontend: separate
admin/,storefront/,platform/andshared/ownership. Views/controllers and styles are limited to 400 lines; locale data has a documented 700-line allowance. Runtime cycles, unresolved local imports, crossing application boundaries, missing folder contracts and undocumented source files fail CI. - Bundled Email/GA4/Gmail/Slack services are Rust modules in
src/connectors/; independent app examples and browser SDKs remain underextensions/. Python test tooling lives underscripts/and the previous connector is an archived comparison oracle inreference/. The generated inventory covers runtime/tooling sources and is checked for drift. - Studio workspaces load lazily. Root application routing is isolated in
frontend/src/application/; error boundaries keep workspace failures inside the current view and provide reload recovery for rejected cached module imports. The removedCommerceManagerhad no call site and duplicated old operational UI. Order state management remains inadmin/orders/OrderWorkflow.tsx; product review moderation lives inadmin/catalog/ReviewModeration.tsx.
frontend/tests/architecture.mjs also inserts deliberately broken synthetic imports,
a cycle, an oversized file and absent documentation into temporary directories;
it verifies that the architecture checker actually rejects each defect.
Repeatable checks
From the repository root:
cargo fmt --check
cargo clippy --locked --all-targets -- -D warnings
cargo test --locked
cargo build --locked --bins
npm --prefix frontend ci
npm --prefix frontend run format:check
npm --prefix frontend run build
npm --prefix frontend run architecture
npm --prefix frontend run localization
npm --prefix frontend run test:coverage
python3 scripts/formal.py
python3 scripts/formal/mutations.py
QDRANT_URL=http://127.0.0.1:16333 python3 scripts/verify_integration.py
Local integration creates and removes its own uniquely named synthetic database;
existing shops remain outside that run. The default database container is
vendune-postgres-1; override it with --container. It uses DATABASE_URL
from the process or the local .env. CI alone passes --existing-database for its
already-disposable database. Failures remain failures, and child processes stop
before database cleanup. SIGINT flushes optional Rust coverage profiles.
scripts/testing/suites.json is the single registry for 48 HTTP suites, 17 local provider/connector/integration suites, four browser contracts and three verification-tool commands (72 registered suites). The local server uses an offline model URL; live model checks are separate, opt-in checks. Credentials, payments, mail and Slack are exercised against loopback protocol fixtures. Passing these does not demonstrate a real provider account.
The currencies suite exercises channel restrictions, Store API/MCP/UCP agreement,
fixed and converted prices, shipping/coupon conversion, stale checkout reviews,
immutable orders and grouped revenue. payment_providers tests create/capture/refund
with EUR/USD, zero-decimal JPY and three-decimal KWD against a local protocol fixture.
production_foundations denies cross-tenant currency-job SQL access using non-owner
roles, then restarts a worker after its first 100-product batch and verifies completion,
exact midpoint rounding and exclusion of products created after the job snapshot.
See multi-currency commerce for the configuration and proof boundaries.
The original-PHP comparisons and reflected automation catalog checks are:
python3 scripts/differential.py
python3 scripts/context_differential.py
python3 scripts/delivery_differential.py
python3 scripts/rule_differential.py
python3 scripts/automation_registry.py
python3 scripts/automation_differential.py
python3 scripts/money_boundary_differential.py
The automation suite also executes scripts/playground.py as documented: a new
owner-owned shop, repeat setup after an edited rule, both checkout branches and
actual invoice records. tooling_tests.py has 12 tests, including refusal of
remote credential destinations and unsafe state-file symlinks. See
the manual tour for the corresponding browser steps.
Production boundary regressions
production_foundations runs two non-owner, non-bypass PostgreSQL replicas and
repeats native checkout, CRM and adversarial tenant checks. It rejects unsafe
TRUNCATE grants, omitted/forged tenant predicates, stale operator quotas and
corrupted stock snapshots; it tests commit-only cache eviction and daily AI usage
shared across live/staging. The 32-client scalability fixture explicitly configures
32 tenant permits. Production defaults and saturation policies remain configurable.
The architecture guide explains exactly what these
checks and the Lean subset establish.
Component regressions
frontend/tests/unit/ uses real components/hooks/transports, synthetic domain
fixtures, jsdom and Testing Library. Fetches fail by default unless explicitly
supplied by a test. Cases include:
- merchant authentication, chat failures, language changes, uploads and separate live/staging transports;
- navigation through permission-filtered lazy Studio workspaces and a signed-out login path;
- catalogue filtering, cursor pages, stale responses, customer-session expiry and forbidden operation replay;
- independent default addresses, linked customer/order/product back paths, configured translated groups, lazy history and confirmed revision-bound restoration with retained failed drafts;
- opt-in personalization, ranking, consent withdrawal and failed signals;
- product gallery, variants, normalized purchase quantities, tier prices, SEO, reviews and public source-backed product questions;
- delayed product answers after navigation and product-specific moderation with rejected writes;
- connected flow branches/deletion, frozen execution traces, incomplete JSON drafts and source condition round-trips in four languages;
- SMTP/Resend/SendGrid fields, saved secrets, read-only settings, translated templates and explicit dry-run test calls.
Coverage collection
Frontend coverage includes every src/**/*.ts and src/**/*.tsx, including
modules never imported by a test. HTML, JSON, LCOV and summaries go to
frontend/coverage/. Styles and independent guest app bundles are outside that
statement metric and explicitly listed as separate scopes.
For Rust install cargo-llvm-cov 0.9.1 and llvm-tools-preview. Use the same
instrumented binaries for units, original-PHP comparisons and the actual HTTP
integration run:
export CARGO_LLVM_COV_TARGET_DIR="$PWD/target"
cargo llvm-cov clean --workspace
eval "$(cargo llvm-cov show-env --sh)"
cargo test --locked
cargo build --locked --bins
# Run the integration and comparison commands above, then:
mkdir -p artifacts/coverage
cargo llvm-cov report --json --output-path artifacts/coverage/rust.json
cargo llvm-cov report --html --output-dir artifacts/coverage/rust
Mutation variants write profiles outside the real-core report directory; deliberately broken code cannot inflate the production coverage metric.
The Rust report measures lines, functions and regions, not branch coverage. The JSON preserves every module in this instrumented scope. Do not call covered regions exhaustive branch coverage.
Python instrumentation is a development dependency, not a production service dependency. It collects subprocesses and includes untouched namespace modules:
python3 -m pip install -r scripts/testing/requirements.txt
COVERAGE_PROCESS_START="$PWD/scripts/testing/.coveragerc" PYTHONPATH="$PWD/scripts/testing" python3 scripts/verify_integration.py
python3 -m coverage combine --rcfile=scripts/testing/.coveragerc
python3 -m coverage json --rcfile=scripts/testing/.coveragerc
python3 -m coverage html --rcfile=scripts/testing/.coveragerc
python3 scripts/testing/coverage_report.py
The report covers Python app/service sources and verification tools, so its
percentage must not be presented as coverage of the Python services alone.
Before a fresh standalone Python measurement, remove only the prior synthetic
artifacts/.coverage* data, to avoid mixing historical executions.
CI uploads complete reports and lists zero-hit modules and unmeasured scopes.
coverage-policy.json locks the measured regression floors; Studio transport, server-health and workspace-boundary
modules require 100% lines, branches, functions and statements.
Adding a source file counts in the denominator. Raising those floors is an
explicit reviewed change; do not silently lower them to obtain a green check.
Explicit 100% audit and open work
npm --prefix frontend run test:coverage:full
python3 scripts/testing/coverage_report.py --require-full
These commands currently fail, as they should. The second also rejects a full system claim while browser SDK/example-app JavaScript and Wasm instruction coverage remain unmeasured. CSS requires visual behavior checks; live OAuth credentials, external inference and production deployment need their own verification.
The initial measured snapshot is in quality-baseline.json. The largest remaining frontend gaps include order-detail edits/documents, customer/address edge cases, full operator workflows, app iframe errors and complex rule/flow interactions. Coverage HTML shows individual missed statements and branches, not only totals. Future changes must add behavior regressions in these owning modules and increase coverage; a passing threshold is not completion of the 100% objective.
Lean source review locks and mutation checks continue in CI unchanged in scope; see formal verification for precisely what is proved.
The account composition regression exercises both login and registration through
CustomerAccount → CustomerSignIn → shopApi. The simulated transport rejects
protected profile/address/order reads unless the form persisted the exact
tenant session key. This reproduces the former React-reserved key prop bug;
backend address/customer isolation is separately exercised over real HTTP.
Studio session regressions reproduce a product 401 after a previously loaded overview, verify that all private controller state is cleared, and reconnect with a renewed token. Shared merchant calls and focus revalidation exercise the same path. Negative cases cover 403/5xx, customer/provider/operator errors, invalid login attempts, old-token replies after reauthentication, coalesced reads, hidden tabs, overlapping checks and timer/listener cleanup. Browser verification uses the existing personal test account; no expiry bypass or new privileges are added.
International configuration regression scope (2026-10-05)
New PostgreSQL suites cover complete bundled geography, custom countries/regions, US destination taxes in actual product/cart/order paths, saved-rule conditions, non-English main-language inheritance and stale/foreign configuration rejection. Translation fixtures cover Ollama/OpenAI/Claude wire formats, durable full-catalogue drafts, cold restart, stale apply, repeated apply and permission/tenant denial. They perform no paid model calls. React tests exercise keyboard/group country selection, missing-field inheritance and a shared settings draft across navigation. UI screenshots supplement these tests; neither replaces whole-source coverage.
Recorded international-commerce verification
CI run 37285967720
verified the international-commerce tree merged as 852f3d1 on 5 October 2026:
67 Rust unit tests, 90 frontend tests and the registered 25 HTTP / 3 provider /
4 browser-contract suites. The complete measured report, including untouched
modules and unmeasured scopes, is checked into quality-baseline.json.
| Scope | Lines | Other measured metrics |
|---|---|---|
| Rust | 81.00% | Functions 78.41%; regions 77.84% |
| Frontend | 42.28% | Branches 36.34%; functions 32.49% |
| Python | 74.11% | Branches 52.78% |
These percentages retain the separate source scopes above and do not establish 100% coverage or whole-system correctness. Regression floors were preserved.
Company profile HTTP/database checks run in the suite registry as company_settings; the scope, decoder/ACL, issuer snapshot, translation and selective release cases are detailed in company settings. They use synthetic tenants and local image bytes, with no external payment/model/OAuth traffic.
Channel settings and media workspaces (2026-10-05)
See the settings/media guide for the current single-language editor, field-level checkout overrides, dependency-safe method removal, gallery and optional private image jobs. These are native prototype extensions; they do not establish additional full Shopware API/DAL parity, current tax law, paid-provider quality or whole-system formal certification.
App library regression scope (2026-10-05)
Eight component regressions exercise translated search, categories/status/discovery, installation followed by the real detail page, cancellation and failed activation, viewer restrictions, registered native interfaces, locale inheritance and independent artwork failure fallbacks. The isolated PostgreSQL app suite checks metadata round trips, tenant separation, immutable same-version digests, unsafe artwork refusal and the actual shop main language. Strict model-schema regression keeps optional presentation compatible with both provider request formats.
The app-library clean frontend run passed 137 tests: 50.16% statements, 43.76% branches, 41.35% functions and 50.98% lines across all included source. These measurements are not whole-system 100% coverage. Actual browser checks supplement component tests with desktop and 375px layouts. The affected apps, app surfaces, App Studio and developer-document HTTP suites pass using local fixtures, without paid calls.
Connected knowledge workspace (2026-10-05)
The current clean frontend run passes 151 tests across 26 test files, including ten knowledge regressions: whole-shop census, a single inherited-language editor, explicit empty values, non-English main language, guarded publication, viewers, retrieval scope, failed edits, actual product navigation and recommendation review. Across all included source the run measures 52.16% statements, 45.88% branches, 43.10% functions and 52.99% lines. This is a measured frontend scope, not a whole-system coverage update or a 100% claim.
84 Rust unit tests, strict lint/format, the extraction/Lean gates and negative
mutations pass. The affected isolated PostgreSQL/AGE suites (knowledge_workspace,
developer_documents, apps, staging, connected_apps, automation) exercise source hashes,
public/private/archive fences, enabled-language retrieval, >50-source pagination,
optimistic revisions, graph reassignment, granular HTTP/MCP access and selective
release. Local provider fixtures prove private document passages reach the actual
merchant model prompt and disappear after archive. No paid provider calls are used.
The knowledge suite also follows source ingestion through a product-scoped event rule into a durable completed flow without an order. Four UI regressions retain the stable trigger IDs while translating their labels EN/DE/FR/ES.
Browser checks cover inherited source creation in App Studio Lab, private customer exclusion versus merchant retrieval, the real product preview, desktop and 375px layouts. See the knowledge guide for source ownership, HTTP/MCP contracts, screenshots and the fixed-model/causality boundaries.
Adversarial tenant isolation
The registered tenant_isolation HTTP suite uses two synthetic workspaces, real
merchant/customer sessions, direct MCP/UCP calls, foreign IDs, forged scope
headers, same-ID app records and revoked one-workspace integration keys. It
re-reads victim state after denied mutations. The database helper tests all 22
new composite foreign keys with own/foreign controls, scans tenant foreign keys
for missing scoped counterparts, and tests forced app RLS under a restricted
role including context reset on the same connection. It runs automatically via
the existing CI integration registry, with zero provider calls.
Exact scope and remaining core-RLS boundary. Grouped checks are behavioral evidence, not 100% endpoint/authorization coverage.
Checkout verification (6 October 2026)
The checkout change passes 226 frontend tests across 34 files (including twelve
checkout regressions) and 96 Rust unit tests. The preceding 75bd24a whole-source
frontend coverage snapshot measured lines 60.78%, statements 59.70%, branches
53.50% and functions 49.32%. These figures are
frontend-only and do not replace the older full-stack baseline report.
Real isolated SQL/HTTP suites verify review admission (five checks), native payment wire behavior (ten checks) and tenant isolation (38 checks). The browser saved a synthetic order after explicit review; 390-pixel emulation had no checkout content overflow. No paid model/PSP calls were used. The public Studio's origin fix was verified separately with allowed-origin 200, foreign-origin 403 and a fresh merchant-owned test shop with products, eight discoverable apps and browser MCP.
Platform control-plane regressions (6 October 2026)
platform.py runs platform_control.py against disposable PostgreSQL shops and
a local synthetic model service. Sixteen combined checks include independent
operator grants, two-shop/fresh-process encrypted provider inheritance, stale
revisions, pause/background deferral, reversible trash, Host-based routing and
persisted measured-call latency. No external model or payment call is made.
shop-scope.test.ts covers public subdomain upgrades and protected/local URL
exceptions. platform-control.test.tsx covers independent entry links, write-only
key editing and typed-ID/reason confirmation before trash. The complete frontend
unit run passed 38 files / 261 tests; Rust passed 100 tests. Affected provider,
transaction/payment, app/flow/staging, translation/image/document/search and
tenant-isolation integration suites also run through their real local paths.
These checks are not live provider capability or a full-system security proof.
managed_search.py allocates its own loopback port, so an open local Studio
fixture cannot accidentally satisfy its readiness check. Qdrant-backed suites
require the actual configured local Qdrant address.
Core hardening regressions (8 October 2026)
The existing strict-runtime suite now checks persistent login backoff across two
processes, poison-event savepoint isolation and recovery, referenced/unsettled
retention, and database-wide admission saturation/expiry. transaction_pooler
repeats real checkout, customer registration/login, CRM, app/staging and RLS paths
through one actual PgBouncer backend; listeners stay on a direct runtime connection.
This exposed and removed nested pooled borrows under checkout/login locks.
CI installs PgBouncer 1.21+ and uses the same suite registry. Native execution may
select TEST_PSQL and TEST_PGBOUNCER; all databases and pooler processes are
isolated fixtures and cleaned up. Unknown route rights and trusted Principal
extraction also have unit/static controls. Implementation and boundaries.
These checks establish those behaviors, not 100% whole-system coverage, public rollout status, a speed-up percentage or protection against a compromised server.
App-platform v2 verification
Registered PostgreSQL suites exercise app_consent, app_callbacks,
app_preview, app_relations, app_schema, app_jobs, app_components,
app_assets, app_secrets, app_events, app_export and app_distribution.
They observe installed versions and actual records/DDL, denied foreign mutations,
lease fences, provider signatures, replay and export bytes. Existing apps,
services, staging, connectors and automation suites verify compatibility.
Frontend tests exercise native code-behind, action allowlists, actor draft restore, revision conflicts, iframe bundle rejection, controls and package consent. These are regression evidence, not a claim of 100% app/host coverage. Full semantic document tests additionally require the real Qdrant service; a missing retrieval service must be reported as an unverified environment dependency, never replaced by a mocked successful semantic result. See the security matrix.