Less manual onboarding: most products publish untouched, and people review the uncertain ones
Onboard products faster, and be able to prove where every value came from
Onboarding a product means finding its weight, dimensions, datasheet, images and customs code, then writing the copy, per product, per supplier, in whatever format the supplier happened to send. Vareoprettelse removes most of that work and makes the part that remains defensible: every value it publishes can be traced back to the source it came from.
It sits between a company's PIM and its messy supplier inputs. What it hands over is not scraped data but a reviewed proposal: one value per attribute, with the source it came from, how confident the system is, and why that value won over the alternatives.
Most products pass through without anyone looking at them. The ones where the sources disagree, or where a required field is missing, are the ones that reach a person.
The Problem
Why manual onboarding caps how fast a catalogue can grow
For distributors and retailers, onboarding a product means hunting down weights, dimensions, datasheets, images, customs codes and writing copy, per product, per supplier, in inconsistent formats.
The result is slow, expensive, and unauditable. In a flat PIM, weight = 12.5 has no story: when it is wrong, nobody knows why or where it came from.
The same supplier's files get re-keyed every time, and regulated classification like CN commodity codes is error-prone by hand.
- No provenance behind published values
- Repeated re-keying of the same supplier layouts
- Error-prone, regulated customs classification
- Every product eyeballed, regardless of confidence
What It Changes
Two things a catalogue owner can act on
The commercial case does not depend on understanding the architecture. It comes down to how much human time a product costs to onboard, and whether a published value can be defended when somebody questions it.
The effort drops because a supplier's file layout is recognised the second time it arrives, because the values are gathered and reconciled automatically, and because a product only reaches a person when something about it is genuinely unclear.
The defensibility comes from keeping collected evidence and accepted product data as two separate things. A flat PIM records that the weight is 12.5 kg. This records that the weight is 12.5 kg, that two trusted sources agreed on it, which ones they were, and when. When it turns out to be wrong, the trail exists.
The sections below explain how that is built, for readers who need to assess the engineering rather than the outcome.
Why It Matters
Speed and auditability, without choosing between them
Treating product acquisition as evidence, resolution and provenance is what makes AI safe to use here rather than a liability.
Provable data quality, with the source and confidence recorded for every value
AI that shows its work and never overrides a human silently
Tenant isolation enforced by the database, not by application filtering you have to trust
A supplier's file layout is recognised on the second upload, so it is mapped once
Regulated commodity-code classification improves as corrections accumulate
Collected data ≠ accepted product data
Many candidate values are collected per attribute. That is evidence. Exactly one becomes the selected truth, chosen by configurable rules, with a confidence score and a reason you can read. Nothing is silently authoritative.
12.5 kg
Producer datasheet · Tier 1 · Producer
12.5 kg
Distributor PDF · Tier 2 · Distributor
12 kg
Retailer listing · Tier 3 · Retailer
12.5 kg
confidence 0.94
Producer value, confirmed by a second source. Retailer value disagreed and was down-weighted.
Every value carries provenance, a confidence score, and an explanation, so your PIM stops receiving mystery data, and any published number is defensible.
A durable, crash-safe pipeline from messy file to publish-ready product
Each product runs through a job graph whose entire state lives in database rows. Work fans out across many sources in parallel, a phase only advances when the previous one drains, and a crash is a non-event.
- 1
Upload
Supplier file → mapped candidates
Intake - 2
Discover
Web search + match verification
Enrich (parallel) - 3
Fetch
Fan-out across sources
Enrich (parallel) - 4
Score
Multi-source agreement
Resolve - 5
Resolve
One value per attribute
Resolve - 6
Validate
Checksums + completeness
Resolve - 7
Review
By exception only
Approve & publish - 8
Publish
Approved proposal → PIM
Approve & publish
root job
Children are spawned atomically as the parent completes and are only ever appended, so cycles are structurally impossible, so one dead source never stalls a product.
Resolve Phase
Score, resolve, and validate before anything reaches a human
When the parallel enrichment phase drains, a short linear tail turns competing evidence into one accepted value per attribute, and decides whether the product can flow through untouched.
Step 01
Score agreement
Multi-source agreement is recomputed so values confirmed across trusted sources gain confidence and outliers lose it.
Step 02
Resolve one value
A configurable strategy picks a single winner per attribute, with a reason: best confidence, priority source, or authoritative only.
Step 03
Convert & canonicalize
Units are converted and values normalized to internal codes so the proposal is consistent regardless of source formatting.
Step 04
Validate readiness
Checksums, type checks and required-field coverage decide: ready for auto-import, or routed to manual review.
Key Workflows
One pipeline, several operator-facing flows
The same evidence spine powers everything from remembered column mappings to grounded copy generation and PIM publishing.
Scenario context
An operator uploads a supplier file. The system fingerprints the header, looks up a remembered mapping for that sender, and only asks the operator to confirm low-confidence columns before creating one candidate per row.
- Map each supplier once, never again
- Raw rows preserved alongside mapped attributes
- Low-confidence columns surfaced for confirmation
Governance note
Mappings are versioned per tenant and source key, so a layout change is a new version, not a silent overwrite.
Corrections on AI-sourced values are captured as a learning signal, aggregated into governed, human-approved improvements, never silent rule rewrites.
Proof Layer
How it is built to stay reliable at volume
The parts that are expensive to retrofit, the data model, the orchestration, tenant isolation and testing, were built to production standards from the start.
Context
- Operational domain
- High-volume product onboarding from heterogeneous supplier catalogs
- Primary users
- Product data, e-commerce and PIM operations teams
Scope
- Data model
- A dedicated schema for evidence and accepted values, held apart from each other, with values kept per language and per sales channel (52 tables)
- Orchestration
- Work is broken into jobs that run in parallel across sources and survive a crash, with the resolution step waiting until all of them have finished
Constraints
- Isolation requirement
- Tenant separation enforced by the database itself, not by application-layer filtering
- AI requirement
- Grounded and governed, with no value silently authoritative
Artifacts Delivered
- Resolution engine
- Three strategies, multi-component confidence, typed value system
- Classifiers
- Retrieve-then-rerank category and CN commodity code with kNN memory
Outcome Signals
- Throughput signal
- Most products publish without human involvement; people handle the exceptions
- Quality signal
- Multi-source agreement, checksum validation, and grounding guards catch errors early
Outcome signals are anonymized measurements from a defined pilot period. Ranges are used to preserve client confidentiality. The measurement period, baseline, and scope are stated in the classification and evidence note on this page.
How to read this case
Classification and evidence method
A product we build and run in production today, not a one-off engagement.
Anonymized measurements taken over the stated period against the stated baseline. Ranges rather than single figures preserve client confidentiality.
- Case type
- Live Product
- Number basis
- Measured
- Measurement period
- Continuous production operation, counted over Q1 2026
- Baseline
- Manual product onboarding effort per item, timed on the same catalogue before the pipeline
- Scope
- Products onboarded through the evidence and acceptance pipeline for live tenants
Published · Updated
Sovereign AI
The grounded, cost-governed AI in this platform routes through the same kind of central LLM gateway. Read how one control layer governs every model call.
Explore the Sovereign AI caseOperational AI
See the same embedded, human-in-the-loop approach applied inside day-to-day operational workflows rather than product intake.
Explore the Operational AI caseexplore further
Capabilities and reading behind this work
Related capabilities
- Operational AI
AI systems that integrate with existing platforms and workflows, with control, traceability, and operational reliability.
- Systems Architecture
Designing architectural foundations that allow complex organizations to operate reliably and evolve safely.
Further reading
- AI in your editor, or AI in their platform
Two things get called the same name, and only one of them puts an AI system into the customer's estate. Which one you mean decides who has to govern it.
- How to turn a business-built AI prototype into production software
A business-built prototype proves intent and interaction. It does not prove architecture, security, data integrity or operational readiness. Treat it as an executable specification: preserve intent by default, and preserve generated code only where evidence justifies it.
- AI-assisted implementation with frontier models
Frontier models can accelerate implementation, but only when they are used inside a disciplined delivery method: clear architecture, review, testing, security, and production ownership.
Interested in how this approach could work for your organization?
Get in touch