Vendune
docs/managed-hosting.mdView on GitHub ↗

Managed PostgreSQL, private Qdrant and shop migration

Current implementation and test boundary

The Core now installs on ordinary PostgreSQL 17. pg_trgm remains a standard PostgreSQL extension. AGE and pgvector are not needed for fresh installs. Tenant-scoped SQL relations replace the graph adapter; canonical commerce and relationship updates remain in one PostgreSQL transaction. Qdrant v1.16.3 is a private, rebuildable search service. PostgreSQL stores source content, model identity, digest and vectors; a durable queue synchronizes Qdrant and removes deleted points. Query results are filtered by tenant/model and hydrated through current PostgreSQL rows. A failed search/index service retains queued changes and lexical search remains available. This is eventual index consistency, not an atomic cross-database transaction.

A collection is isolated by object kind and embedding model hash; points use deterministic tenant/kind/object UUIDs. A payload index identifies tenants. One SQL query hydrates the bounded product candidate set. Public document retrieval additionally checks publication status, product scope and source digest. Synthetic embeddings test these boundaries, not semantic quality.

002-managed and 011-documents-managed are new install variants. Historical migration source files/checksums remain unchanged. 033-managed-knowledge converts legacy vectors to real arrays, copies curated/observed/document relations and queues existing embeddings. Old AGE tables remain intact. Tests on the legacy image preserve product counts, relationships, all 1024 vector values, model/digest and original ledger rows; checksum drift still stops startup. Run the legacy conversion on the existing AGE-capable source server before moving its public schema/data to a fresh standard server. Do not mount an old AGE database volume under an image lacking its extension libraries. The development launcher keeps existing database containers intact with --no-recreate. This database upgrade is distinct from a complete external Shopware-shop importer.

Northflank test topology

Frankfurt: one Managed PostgreSQL addon, one Rust Core replica and one private Qdrant replica with /qdrant/storage on a 6 GB persistent volume. Core port 8787 exposes the UI and same-origin APIs over HTTPS. PostgreSQL and Qdrant ports stay private. BOOTSTRAP_MODE=migrate runs as a controlled setup job before serve is started. SEED_DEMO, ALLOW_PUBLIC_SIGNUP, ALLOW_BOOTSTRAP_AUTH are false. Personal-only hosting needs no shared merchant token. Initial operator credentials are restricted to the setup job; remove that password from cloud setup secrets after successful bootstrap. Merchant sessions use personal accounts.

SHOP_DOMAIN_SUFFIX=vendune.ai binds shop hostnames to tenant scope. Unknown shops and conflicting tenant headers are rejected. admin.vendune.ai opens the service directory (merchant/platform login and docs), admin.vendune.ai/#platform opens the operator console; SHOP.vendune.ai opens that shop, and /?shop=SHOP#merchant opens its Studio on the operator origin. Before DNS is ready the same paths work on the default Northflank HTTPS origin. Creating a shop is a database transaction, not a new container, database or certificate.

Recorded test-topology estimate (October 2026; not a live pricing quotation): Core $12, PostgreSQL $12, Qdrant $12, PostgreSQL 6 GB $0.90, Qdrant 6 GB $0.90, retained inactive legacy 20 GB volume $3 = $40.80. Example 10 GB egress adds $0.60. Jobs/builds, backup storage, AI APIs and taxes are additional. This single-instance test topology has no HA guarantee. Qdrant's attached single-writer volume requires recreate deployment and cannot be scaled by simply increasing replicas.

Growing into database islands

Start with one island and measure per-shop query latency, database CPU/IO, connection pressure, working-set size and queue lag. Add a second island when sustained load justifies it. Do not create one database per trial shop. Use a small control database directory shop_id -> island_id, epoch, state and keep identity/membership/control audit outside island data. The current Core uses one DATABASE_URL; a directory/router and automated relocation controller are design work, not implemented functionality.

Each island contains an independently sized Managed PostgreSQL and a bounded Core connection pool. Stateless HTTP replicas and durable workers can scale separately. Total pools must stay below the database's connection budget; add a verified pooler before aggressive HTTP autoscaling. An island controller admits shops based on observed capacity and promotes large shops to dedicated islands. Budget limits, minimum residency and hysteresis prevent oscillating migrations. Scaling the current $45 test budget above one Core replica requires revising the budget.

Concrete initial relocation protocol

  1. Provision the target island at the same schema version; run migrations once and verify readiness. Create a transfer record with source/target/tenant/epoch and an idempotency key. Only one transfer per tenant may run.
  2. Fence all source writes for that tenant at the database/transaction layer, including checkout, payments, webhooks, imports and every worker; drain admitted operations. New mutating requests receive a retryable maintenance response. Durably buffer inbound provider events in the control plane. An application-only host redirect is insufficient fencing.
  3. Export a consistent PostgreSQL snapshot of the tenant plus its explicitly owned dynamic app tables, binary assets, source vectors, jobs/outbox/receipts and settings. Maintain a schema-derived tenant-table manifest and foreign-key closure; fail on an unclassified tenant-bearing table. Preserve stable UUIDs. Global bigserial IDs and their references must use a collision-safe mapping or an allocated target ID range; resetting a sequence alone does not prevent collisions.
  4. Restore into a staging transaction with target writes fenced. Verify per-table row counts, content digests, foreign keys, order/invoice totals, stock, permissions, pending jobs and receipt identities. Identity/control records are references to the central directory, not copied account ownership. Explicitly move payment/account settings and encrypted app credentials using the same configured encryption authority or re-encryption; do not print secrets in exports/logs.
  5. Keep the same tenant ID and embedding model. Shared Qdrant points remain usable because index identity is independent of the PostgreSQL island; rehydrate from the target. If the index is also moved, rebuild from persisted PostgreSQL embeddings and verify queued synchronization without paying for new embeddings.
  6. Atomically compare-and-swap the directory from (source, old epoch) to (target, new epoch). Target writes require the new epoch and source writes stay fenced. Invalidate router caches; require workers to re-resolve ownership before commit. Resume target traffic and replay buffered events through preserved idempotency receipts.
  7. Observe actual checkout/order/worker results. Keep the source snapshot fenced for a retention window. Rollback before target writes is simple; after target writes it requires a reverse synchronized transfer. Never reopen a stale source as writable. Retire it only after backup/restore and transfer acceptance.

For a first real relocation, a short per-shop write pause is the smallest reliable approach. A later online copy adds snapshot + WAL change capture, catch-up and the same fenced cutover. CDC alone does not implement multi-writer conflict resolution or migrate every dynamic app table.

External Shopware imports

Keep this separate from island relocation. The existing core parity work is partial; it is not a complete importer for arbitrary Shopware plugins. First select one actual source shop and version. Use a read-only source API/export and a versioned import manifest preserving source UUID-to-target mappings. Import dependencies in order: currencies/taxes/languages, sales channels/customer groups/rules, categories/media, variant products/translations/prices, customers/addresses, orders/payments/documents, then settings/app data. Unknown plugin entities and unsupported rules must be surfaced as blocking mappings, not silently discarded.

Import into a new unpublished shop and compare real catalog/search/cart totals, historical order totals and permissions with the source. Password compatibility and PSP subscriptions require an explicit provider-specific strategy; raw passwords and active payment tokens are not portable by default. Initial cutover freezes source mutations, imports the final delta and verifies new orders/webhooks before changing the shop domain. Keep the source read-only for rollback and historical reference. A customer-facing migration wizard and automated full Shopware importer remain to be implemented and tested against an actual shop.

AI execution

Qdrant searches; it does not generate embeddings or run language models. Existing model adapters call operator-configured Ollama or external OpenAI/Anthropic endpoints. The current embedding contract is an Ollama-compatible /api/embed endpoint returning 1024 finite dimensions. No model/GPU is hosted in the $40.80 infrastructure estimate, and live model credentials have not been verified by the synthetic transport tests.

Run embeddings and batch enrichment as durable background jobs with tenant quotas, bounded concurrency, cancellation and retry limits. Pin model identity; model changes create a separate collection and trigger controlled reindexing. Merchant AI uses canonical data + approved source documents, prepares changes and executes only through existing authorization/approval transactions. Keep private app evidence outside public retrieval. Start with external inference usage billing; introduce a separate GPU inference service only after measured usage justifies its persistent cost. Do not put an unbounded inference server inside the small Core replica.

Platform provider key storage

Keep PLATFORM_SECRET_KEY (64 random hexadecimal characters) in the private runtime secret group, never in build arguments or repository files. Share the same key with every HTTP/worker replica that resolves central AI settings and retain a separate protected backup. The key encrypts provider credentials saved by operators; it grants no HTTP access. Set INFERENCE_ALLOW_LOOPBACK=false on public hosting. See the operator guide for inheritance, reversible shop lifecycle and real infrastructure metrics.