Vendune
docs/cognitive-commerce.mdView on GitHub ↗

Cognitive commerce: connected implementation and remaining work

The cognition layer reads current commerce data, retrieves relevant evidence and proposes changes. PostgreSQL remains authoritative for products, prices, stock, orders, permissions, consent, evidence and jobs. Qdrant is a rebuildable index. Model output is untrusted input; saved knowledge does not update model weights.

This implementation closes a substantial foundation slice of the intelligence audit. It does not implement the entire audit or certify the whole system. The remaining-work table below is part of the contract for subsequent development.

flowchart LR
    SOURCE[Products and document chunks] --> QUEUE[Transactional embedding intake]
    MODEL[Embedding model change] --> CURSOR[Durable bounded generation cursors]
    CURSOR --> QUEUE
    QUEUE --> WORK[Admitted Rust worker]
    WORK --> PG[(Current PostgreSQL vectors and evidence)]
    PG --> PUB[Independent fenced publication queue]
    PUB --> Q[(Private Qdrant text index)]
    PG --> SEARCH[Lexical retrieval and exact native hydration]
    Q --> SEARCH
    SEARCH --> RANK[Optional admitted reranker]
    RANK --> AGENT[Bounded read-tool rounds]
    REG[Canonical capability registry and current rights] --> AGENT
    AGENT --> PROPOSE[Revision-bound proposal]
    PROPOSE --> GUARD[Native price and autonomy predicates]
    GUARD --> APPROVE[Merchant approval or explicitly enabled budget]
    APPROVE --> NATIVE[Existing transactional commerce owner]

Retrieval and model serving

Products and document chunks enter embedding_jobs in their source transaction. The memory worker claims one tenant and at most 32 jobs, reads current content, releases its database transaction, then calls the embedding provider. Publication checks the source text and job revision/lease before writing. Provider failures have bounded retries and classified diagnostics; admission saturation does not consume the provider retry budget. A manual reindex resets failed intake jobs.

Migration 076 adds per-tenant model-generation cursors. A configured model change re-enqueues existing sources in 256-record keyset batches without a request-sized catalog scan. Cursor progress and queue intake commit together and survive a restart. Source triggers cover concurrent inserts/edits. Configure one consistent embedding model across all replicas; mixed models are not a supported rollout. A model's name identifies its semantic space: changing weights behind the same name requires a new versioned name or an explicit reindex. Retired Qdrant collections are not automatically garbage-collected.

Embedding dimensions are validated per response, 1–8192, with consistent batch geometry and finite, nonzero vectors. Qdrant collection names bind kind, model and dimension; they contain a named text vector and an indexed tenant payload. Candidates are rehydrated against current PostgreSQL tenant/model/content state. Pending source changes cannot use their old dense product vector. Publication claims are ordered by next due attempt, so old retries do not repeatedly displace new work; ready batches drain without a two-second delay between every packet. This is dense plus PostgreSQL full-text reciprocal-rank fusion, not BM25 or image search. The optional TEI reranker validates every candidate index; failures retain fused ordering. Whole records/quotes are omitted when agent context budgets are exceeded; truncated quotations are not presented as complete evidence.

Configuration Contract
EMBEDDING_MODEL Versioned model identifier, default existing local configuration
EMBEDDING_PROTOCOL=openai OpenAI-compatible /embeddings; otherwise Ollama /api/embed
EMBEDDING_BASE_URL, EMBEDDING_API_KEY Operator-configured server-only endpoint/key; base URL includes /v1 where required
QDRANT_URL, QDRANT_API_KEY Private index; pooled HTTP client, bounded responses, no redirects
QDRANT_QUANTIZATION=int8 Scalar quantization when creating new collections; existing collection configuration is retained
RERANKER_URL, RERANKER_API_KEY Optional TEI /rerank endpoint, 24 candidates, three-second ceiling
OPENAI_PROTOCOL=chat Self-hosted OpenAI-compatible chat-completions, suitable for configured vLLM/SGLang servers
INFERENCE_MODEL_PLANNER_OPENAI, INFERENCE_MODEL_EXTRACTION_OPENAI, INFERENCE_MODEL_CONCIERGE_OPENAI Task-specific model routing for explicitly configured providers; the Platform choice retains centrally inherited configuration
INFERENCE_PROMPT_CACHE=true Provider-native system-prefix caching; no cache of private commerce answers
INFERENCE_CONCURRENCY Shared per-process inference admission, default four; tenant fleet leases additionally bound model, embedding and reranker concurrency

Running a compatible adapter does not prove a server's batching throughput. Interactive embeddings and background indexing share admission. Reranker requests are admitted before provider invocation. Saturation falls back to lexical/fused retrieval instead of creating an unbounded inference queue. Interactive chat (including SSE), buyer-facing Concierge, MCP merchant planning, document extraction via HTTP/MCP and Flow AI actions consume the same atomic UTC-day tenant attempt quota. Staging shares its live tenant budget; failed provider attempts count. This is not token/spend accounting.

Migration 078 maintains per-error-code counters in the source-job transaction. Status reads inspect this small per-tenant projection, not every failed job. Error changes, reset/retry and deletion remove old contributions before adding new ones; the migration backfills the projection once. This does not change retry admission.

Lexical candidates under forced row security

Migration 079 adds two rebuildable word projections, knowledge_product_lexemes and knowledge_chunk_lexemes. Source text is tokenized once in its native write transaction. The tables store only tenant, source identity and distinct simple lexemes, with composite source foreign keys and forced tenant RLS. Source deletion cascades; name/description/chunk edits replace their tokens atomically. Price or stock-only updates do not rewrite words. Source content, publication, association, locale, rank and commerce state are still read from the native owners.

This choice addresses a PostgreSQL security/performance interaction: the native FTS operator is not leakproof, so a GIN index can work under a database owner but fail to narrow candidates under the ordinary forced-RLS runtime. B-tree equality on tenant/lexeme narrows candidates without relaxing RLS or reclassifying a PostgreSQL function. Products require every query term; documents retain their existing any-term behavior. Both paths recheck the native text predicate and current source admission before returning results. A forged/stale candidate is not sufficient to admit a product or quote. No extra service or truth ledger is introduced; the historic document GIN index is retained, not duplicated.

The migration takes a source-write lock for one backfill and trigger installation. Provision a maintenance window and storage headroom for large existing catalogs: this is not an online zero-downtime index build. Storage grows with distinct source words, and writes maintain two B-tree indexes per projection. Large/common term posting lists still require candidate ranking. This is lexical filtering, not BM25, language stemming, semantic quality or a universal sub-millisecond API. Rebuilding these projections is an operator migration task, not a public endpoint. Runtime grants follow the existing post-migration role setup.

lexical_search executes the actual product/document SQL under a non-owner NOSUPERUSER/NOBYPASSRLS role, with 5,000 products and 5,000 document chunks. It checks the actual indexed plans without disabling sequential scans, preexisting source backfill, AND/OR semantics, absent/foreign contexts, composite references, scoped edits, private/archive withdrawals, native rechecks and cascade deletion. Document hydration is fenced to lexical/semantic candidate IDs before checking current locale, publication and product associations. Parameterized lateral lookups prevent a broad document scan under forced RLS; unexecuted alternative plans are not counted as actual scans. Semantic IDs are split on their final position suffix, so document IDs containing colons remain valid. Native hash/model checks still reject stale semantic candidates. The separately measured 20,000-product selective case used five paired baseline/candidate SQL executions with identical native results. Median execution was 36.209 ms before and 0.243 ms after under the scoped runtime role. Recorded conditions and source hashes accompany this database micro-test. It does not measure HTTP, inference, mixed tenant load or the entire platform.

Agent reads, writes and progress

Planner/Concierge separate read decisions from the final response. Up to four read-decision calls allow at most three executed read rounds, three tools per round, followed by one final response call (at most five model calls). The read-decision schema cannot contain proposed changes. Malformed decisions stop without a task. Core tools are explicitly classified; current actor permissions and input schemas are checked at dispatch. Installed app tools require a current authorized read-only contract. Storefront agents can read only public catalog tools. Tools cannot execute writes. Provider output may propose the existing price/stock/layout/managed-app changes; the server binds actual revisions, validates and persists the proposal. Expanding this to all registry mutations remains open.

Proposal review retains the existing permission-filtered, bounded appContext projection (up to 32 KiB). modelAppContext records the separate 4 KiB prompt projection, including explicit whole-record omissions. Inspecting a proposal's review data does not establish that the model read those omitted records.

POST /api/agent/chat/stream carries acceptance, elapsed waiting updates and final completion/errors through the existing conversation and lease owner. It is not provider token streaming, a new chat backend or process-resumable model execution. Reloading retrieves the persisted conversation. Reviewed claims, experiment lifecycle/results and applied proposals emit native outbox events selectable in Flow Builder; the existing worker executes them. Scoped app subscriptions receive minimized IDs/states, never private quotes or actor identities. Context omits oversized records explicitly and never dumps a million-product catalog into the model.

Buyer context admission and response freshness

Concierge shares one retrieval between its localized catalog and relevant graph. It admits the initially supplied products through current sales-channel visibility and stock before model calls. Search hits, relation endpoints and approved pair observations outside that admitted catalog are omitted from both model input and the returned knowledge context. Public read tools still use their native Store API admission; the model cannot widen merchant permissions.

After inference, the advisor rehydrates only its original at-most-24 IDs, checks current locale/channel admission, compares the exact catalog projection and rebuilds its admitted native evidence neighborhood. Changed price, stock, description/translation, visibility, confirmed evidence or channel state rejects the response with a localized 409. This adds bounded native reads rather than another embedding/search request or a catalog scan. It does not freeze commerce state across inference, certify free prose, or revision-fence every additional model-selected read-tool result; those remain explicit limits.

testing/advisor_sources.py runs within the existing registered provider fixture. It checks the actual provider input and response against an isolated lamp-only channel, then pauses final inference while native product/channel/document APIs change each source. No private shop data or paid provider is used.

Typed evidence and public statements

The existing knowledge_relations ledger now has typed source/target nodes, proposed/evidenced/confirmed/rejected states, confidence, business-valid time and recorded-time history. Confidence is not a probability that an assertion is true. An omitted candidate confidence uses 0.5; supplied values must be finite JSON numbers in 0–1. Nulls, strings and out-of-range values are rejected before a claim is stored, rather than silently substituted with a score. Current node types include product, variant, material, property, intent, problem, occasion, audience, claim, return reason, supplier, policy, document and support.

knowledge.extract performs document extraction through HTTP/MCP and the native graphical Flow Builder. An explicitly configured pipeline can subscribe to knowledge.document.ingested or knowledge.document.updated and extract its event document. Leave sourceId/productId empty to use the current source and its owning product; extraction uses the source language. The existing durable worker checks the creator's current catalog/knowledge permissions before each step, shares the daily AI quota, records the result and marks interrupted effects uncertain rather than automatically repeating a provider call. No extraction flow is enabled by default. Each candidate must quote an exact current source chunk; it remains proposed. Merchant review requires current catalog permissions, explicit approval and the exact claim revision. Confirmation does not turn an arbitrary model paraphrase into a mathematical entailment. Changing/archiving/private-marking the source removes public admission. Continuous review/return/support/product extraction is not yet wired merely because those node types exist.

Document retrieval carries native product/family/shop association separately from untrusted titles. The shared Planner/Concierge/product-question instruction prioritizes native structured properties and requires conflicting source values to be reported. These instructions do not establish semantic entailment or guarantee that a model follows them.

Two-hop retrieval limits branch width and total evidence. The public compiler accepts only exact current confirmed statements backed by public documents. Batch intake/compilation handles 24 products and at most 20 statements per product, without a per-product network loop. It does not certify surrounding generated prose. The private Storyfront bridge adds these statements to the original native manifest and checks bound source statements again before publish/public delivery; its implementation is kept outside this public repository.

Signed facts bind native product data, public evidence, tenant, channel and a five-minute validity window. Configure FACT_SIGNING_SEED as 32-byte hexadecimal or use the domain-separated existing platform secret. Verify the exact decoded payloadBase64 bytes with the public JWK, not a separately reserialized payload. Stock is not reserved and estimated delivery is not guaranteed by a signature. The UCP facts route is a Vendune extension, not complete UCP certification.

Saved fact navigation

Studio Products → Categories can augment a listing with a saved fact query. Choose an intent, problem, occasion, audience, material or property node type, minimum recorded confidence and a localized literal phrase. The existing category CAS, history, staging clone/release and enabled-language editor own this data; there is no second navigation registry. Missing/null phrases inherit the shop's main language and search claims in that source language. An explicit empty phrase adds no fact matches for that locale. Manual product assignments remain independent.

The ordinary Store API/MCP category listing selects current confirmed public claims, checks their document hash/revision, product scope and validity window, then applies current channel visibility, active products and cursor pagination. Archiving or making a source private withdraws only its graph-based membership. Migration 077 adds partial type/confidence and text-trigram indexes. Browsing runs no model and stores no duplicated category-product projection. This is a bounded one-hop typed fact predicate with literal substring matching, not arbitrary Cypher, inferred semantic entailment or a full intent-planning language.

Merchant guardrails and optional autonomy

Studio Shop knowledge → Guardrails edits commerce.settings.aiPolicy with the same native settings revision and currency scales. Product corridors can require minimum/maximum prices, net unit cost/margin, maximum discount relative to the current base price, brand price lock and available stock. Tax conversion uses the native product tax basis. The margin excludes payment fees, shipping and returns; it is not contribution margin.

Autonomy is disabled by default. When explicitly enabled, it is price-only, requires current catalog/settings rights, an explicit product corridor and a daily unique-SKU quota. Default limits are ±5% and 20 products per UTC day. The original daily price and currency prevent repeatedly compounding a permitted adjustment. Native transactional apply rechecks settings/product revisions, guardrails and budgets; preview is not execution. There is no autonomous recurring merchant-goal worker in this slice.

Controlled experiments and private preferences

Studio Shop knowledge → Experiments creates immutable preregistrations for two native layouts, one channel/currency, horizon, settlement delay, minimum arm size, outcome cap and fixed CUPED coefficient. Explicit start/stop/finish operations use current settings rights and revisions. Consent-bound cart units receive a persisted 50/50 assignment that the actual native experience consumer uses.

The report reads live confirmed payment captures minus completed refunds from the existing payment ledger. Demo purchases do not reward the experiment. The conservative fixed-final-look Hoeffding interval requires a mature horizon/delay, sufficient units, matching currency, no early stop and no privacy withdrawals. A mature report is saved once and written as an evidenced graph result/outbox event. Erasure invalidates inference and the published graph result rather than hiding attrition. Repeated reports cannot manufacture additional final looks. This assumes independent cart units; repeated buyers/interference and unrecorded late returns remain limitations. No actual commerce uplift has been established by synthetic transport tests. Switchbacks, contribution-margin outcomes and automatic goal selection are not implemented.

Customer profile controls save a typed private graph only with personalization consent. AI advice uses it only with a separate explicit sharing choice and for 30 days after update. Export and deletion use the same native cart context; revocation deletes it. This is cart/browser-context memory, not authenticated cross-device customer memory. Customer preferences are never public product facts. The advisor captures a request-only fingerprint of its admitted private graph (tenant/cart identity, revision, update time and exact content). After all model rounds, it rechecks current consent/sharing and that fingerprint before returning an answer. Withdrawal, erasure, edits, unsharing or erase/recreate invalidate the in-flight response with a localized 409 asking the customer to retry. No private answer is cached, and no database transaction is held during inference. This is validation at the response boundary, not a guarantee that consent cannot change after validation or an authenticated cross-device memory implementation.

Ownership and verification

Owner Responsibility
src/cognition/indexing.rs, generations.rs; migrations 071/074/076 Source-trigger intake, model-change cursors, leased embeddings and status counters (including migration 078 diagnostic deltas)
src/knowledge/{search,vectors,vector_cache,rerank,embeddings}.rs, src/documents/search.sql; migration 079 Retrieval, current native hydration, forced-RLS word candidates, index transport/geometry and provider validation
src/cognition/{context,tools,stream}.rs; src/inference/protocol.rs Bounded context, authorized read rounds, existing-chat SSE and model protocols
src/cognition/{evidence,extraction,claim_batches,contracts,signed}.rs; migration 072 Source-bound evidence lifecycle and current public statement/signature adapters
src/cognition/{guardrails,autonomy}.rs; migration 073 Native policy, current revisions, daily price budget and transactional consumer
src/cognition/experiments/; migration 075 Native assignment, preregistration and actual payment-ledger readout
src/cognition/preferences.rs; existing privacy owner Private graph, current consent, advice admission and erasure
src/categories/{admin,graph_query}.rs, listing.sql; migration 077 Native saved fact-category predicates, source admission and existing catalog consumer
frontend/src/admin/intelligence/, storefront account and shared i18n/API Multilingual editors and native request transport

Registered verification includes users, providers, knowledge_workspace, cognitive_experiments, managed_search and tenant_isolation. Providers use local synthetic HTTP fixtures; managed search uses actual PostgreSQL/Qdrant with synthetic embeddings, 150 automatic product inserts and model-change restarts. lexical_search adds actual non-owner indexed plans and transactional lexical-source regression cases. These tests establish contracts and state effects, not semantic model quality or throughput. The existing providers suite also pauses actual concierge HTTP inference and mutates native consent/preferences concurrently. It checks all five invalidation cases, unchanged-context completion, explicit sharing, independent carts and buyer-facing daily-quota denial before any model invocation. Frontend tests cover SSE parsing, permission/revision failures, preregistration controls and consent/sharing. Rust/Lean comparison and negative mutations cover the exact extracted predicates for price admission, autonomy, claim rendering and experiment-result admission. SQL, statistical assumptions, signatures, provider protocols and browser code remain outside those proofs.

Local real-model check

An isolated local run used qwen3.6:35b and qwen3-embedding:0.6b through the ordinary HTTP, worker, retrieval, tool and proposal consumers. An English data sheet explicitly assigned to a mug said 500 ml; the native product said 350 ml. A German question retrieved the indexed sheet. Extraction saved four exact-quote candidates as proposed, without confirming them. The initial combined read/proposal schema skipped an explicitly requested read. With separate phases, the model called merchant.product.content once, then produced a price-only 17.90 EUR proposal awaiting approval and noted the capacity conflict.

This is one real inference example, not a semantic-quality benchmark. Earlier answers varied, including mistaken product association inferred from the sheet name and inconsistent handling of absent native fields. Shared instructions and association metadata reduce that ambiguity; they do not prove every answer true. The run used no paid provider and no existing customer/shop data, and applied no price change. Transport fixtures separately test denials and malformed output.

Audit completion tracker

Requested point Current implementation Still required
Currency-neutral planning, pooling, bounded indexing/status Native currencies, pooled Qdrant, automatic source intake/model rebuild, exact counters Large-scale mixed-load measurements; large-scale end-to-end latency and admission measurements
Hybrid multimodal retrieval Lexical+dense RRF, optional reranker, named text vector, optional INT8 BM25/sparse scoring and real image embeddings/search
Modern serving/tool use Configurable self-hosted protocol, routing, admission, read rounds, progress SSE Real deployment/throughput; token streaming; generic write proposals
Provenance ontology Typed nodes/time/history; on-demand and opt-in document-event extraction/review; bounded public graph retrieval Continuous product/review/return/support extraction with privacy gates (document events already use the native Flow worker)
Verified autonomy Native price corridors, margin/budget checks and exact extracted predicates Broader actions and recurring goal execution
Causal learning Native layout holdouts/CUPED/delayed net-cash readout Switchbacks, recorded-return/CM metrics, actual live evaluation
Intent navigation Saved typed current-public-fact predicates in native categories, Store API/MCP listing, localized editor and staging Multi-hop intent resolution, benchmarked semantic quality and richer query language
Claim compiler Exact current confirmed statement admission; private native Storyfront bridge Every generated/edited prose claim must bind its original native knowledge path
Agent offers Signed native facts; existing authoritative checkout Constrained bundle/quantity/delivery negotiation and reservation contracts
Digital twin Existing sandbox/quote/proposal owners retained Validated historical replay/simulator, uncertainty and native decision linkage
Retouren/reviews feedback Node types declared Actual feedback intake, size advice proposal and return-rate experiment
Private customer memory Consent-bound cart graph and explicit AI sharing/export/delete Customer-owned cross-device memory with complete privacy lifecycle
Graph-native apps Optional namespaced node/edge mappings over selected native app fields/references; current revisions and grants; one Studio/agent manifest and API/MCP/planner consumer Unstructured app extraction, product-fact source admission and global semantic traversal
Merchant goals Existing approved proposals Durable observe→hypothesis→experiment→proposal loop and progress UI
Infrastructure recommendations Existing Rust/PG17/RLS/outbox/cache/lease/CDN contracts remain PG18 upgrade, generated OpenAPI, OTEL export, analytic projection, CoW staging; measured sharding strategy

These remaining items must extend their existing owners. Do not introduce another pricing engine, approval system, truth graph, document source store or workflow queue to make a recommendation look implemented.

Public answers and concurrent source changes

Product questions re-admit the native product snapshot after inference and fence every source supplied to the model, including uncited inputs, against current PostgreSQL tenant, publication, archive, association, revision, hash, chunk text and locale. Withdrawal, re-publication, reassignment or a changed product returns a localized conflict asking the customer to retry; no answer/excerpt is returned. Provider work holds no commerce transaction. This admission check does not undo previously authorized input already sent to a provider and does not prove prose semantically correct. Delayed-provider HTTP tests change sources and products through their real merchant APIs while the question is in flight.

App-owned ontology extensions

Apps map selected native models/fields and real core/app reference edges through intelligence.ontology, edited in App Studio or by the coding agent. The existing native list projects a bounded current graph view after permission/RLS checks; merchant planner, API and MCP consume it. No duplicated fact tables or new queue are introduced. Public app list actions deliberately expose selected graph fields; private mappings retain current action rights. Contract and limits, example. App records cannot bypass source admission and merchant confirmation in the product claim compiler.