Case Study

Product data with provenance, not scraped guesses

Vareoprettelse sits between messy supplier inputs — Excel exports, producer feeds, bespoke sheets — and a company's PIM. It does not deliver scraped product data; it delivers a resolved, validated product proposal where every value has provenance, a confidence score, and an explanation of why it was chosen.

Many candidate values are collected per attribute as evidence. Exactly one becomes the selected truth, chosen by configurable rules — or escalated to a human. Automation handles the 90%; experts review the 10% that needs judgment.

Evidence-first intakeGrounded, governed AIHuman review by exceptionMulti-tenant with database-enforced isolation
Scroll to see how it works

The Problem

Manual product onboarding is the bottleneck on catalog growth

For distributors and retailers, onboarding a product means hunting down weights, dimensions, datasheets, images, customs codes and writing copy — per product, per supplier, in inconsistent formats.

The result is slow, expensive, and unauditable. In a flat PIM, weight = 12.5 has no story: when it is wrong, nobody knows why or where it came from.

The same supplier's files get re-keyed every time, and regulated classification like CN commodity codes is error-prone by hand.

  • No provenance behind published values
  • Repeated re-keying of the same supplier layouts
  • Error-prone, regulated customs classification
  • Every product eyeballed, regardless of confidence
The Core Idea

Collected data accepted product data

Many candidate values are collected per attribute — that is evidence. Exactly one becomes the selected truth, chosen by configurable rules, with a confidence score and a reason you can read. Nothing is silently authoritative.

Collected evidence

12.5 kg

Producer datasheet · Tier 1 · Producer

0.94

12.5 kg

Distributor PDF · Tier 2 · Distributor

0.78

12 kg

Retailer listing · Tier 3 · Retailer

0.41
Resolve
Accepted value

12.5 kg

confidence 0.94

Producer value, confirmed by a second source. Retailer value disagreed and was down-weighted.

Every value carries provenance, a confidence score, and an explanation — so your PIM stops receiving mystery data, and any published number is defensible.

The Flow

A durable, crash-safe pipeline from messy file to publish-ready product

Each product runs through a job graph whose entire state lives in database rows. Work fans out across many sources in parallel, a phase only advances when the previous one drains, and a crash is a non-event.

  1. 1

    Upload

    Supplier file → mapped candidates

    Intake
  2. 2

    Discover

    Web search + match verification

    Enrich (parallel)
  3. 3

    Fetch

    Fan-out across sources

    Enrich (parallel)
  4. 4

    Score

    Multi-source agreement

    Resolve
  5. 5

    Resolve

    One value per attribute

    Resolve
  6. 6

    Validate

    Checksums + completeness

    Resolve
  7. 7

    Review

    By exception only

    Approve & publish
  8. 8

    Publish

    Approved proposal → PIM

    Approve & publish
Inside “Fetch” — a variable, append-only fan-out
Discover
root job
Producer site
Datasheet PDF
Official DB
Product images
Marketplace

Children are spawned atomically as the parent completes and are only ever appended — cycles are structurally impossible, so one dead source never stalls a product.

Resolve Phase

Score, resolve, and validate before anything reaches a human

When the parallel enrichment phase drains, a short linear tail turns competing evidence into one accepted value per attribute — and decides whether the product can flow through untouched.

Step 01

Score agreement

Multi-source agreement is recomputed so values confirmed across trusted sources gain confidence and outliers lose it.

Step 02

Resolve one value

A configurable strategy — best confidence, priority source, or authoritative only — picks a single winner per attribute, with a reason.

Step 03

Convert & canonicalize

Units are converted and values normalized to internal codes so the proposal is consistent regardless of source formatting.

Step 04

Validate readiness

Checksums, type checks and required-field coverage decide: ready for auto-import, or routed to manual review.

Key Workflows

One pipeline, several operator-facing flows

The same evidence spine powers everything from remembered column mappings to grounded copy generation and PIM publishing.

Scenario context

An operator uploads a supplier file. The system fingerprints the header, looks up a remembered mapping for that sender, and only asks the operator to confirm low-confidence columns before creating one candidate per row.

  • Map each supplier once, never again
  • Raw rows preserved alongside mapped attributes
  • Low-confidence columns surfaced for confirmation

Governance note

Mappings are versioned per tenant and source key, so a layout change is a new version — not a silent overwrite.

Corrections on AI-sourced values are captured as a learning signal — aggregated into governed, human-approved improvements, never silent rule rewrites.

Proof Layer

Architecture built for scale, isolation, and auditability

The hard, expensive-to-retrofit parts — domain model, orchestration, tenancy and testing — are built to production standards.

Context

Operational domain
High-volume product onboarding from heterogeneous supplier catalogs
Primary users
Product data, e-commerce and PIM operations teams

Scope

Evidence spine
52-model schema with per-language, per-channel values and JSONB evidence
Orchestration
Durable hub-and-spoke job DAG with transactional outbox and phase-drain barriers

Constraints

Isolation requirement
Database-enforced Row-Level Security via a two-role Postgres model
AI requirement
Grounded and governed — no value silently authoritative

Artifacts Delivered

Resolution engine
Three strategies, multi-component confidence, typed value system
Classifiers
Retrieve-then-rerank category and CN commodity code with kNN memory

Outcome Signals

Throughput signal
Most products flow through untouched; humans handle only exceptions
Quality signal
Multi-source agreement, checksum validation, and grounding guards catch errors early

Why It Matters

Fast and auditable and trusted — not a trade-off

Reframing product acquisition as evidence, resolution and provenance is exactly what makes AI safe to use here rather than a liability.

Faster onboarding — automate the 90%, review the 10% that needs judgment

Provable data quality with per-value provenance and confidence

Grounded AI that shows its work and never overrides a human silently

Database-enforced tenant isolation, not app-layer filtering you have to trust

Mapping-with-memory removes repeated re-keying per supplier

Shared, regulated commodity-code learning compounds across tenants

Explore the Sovereign AI case: Sovereign AI
Related Infrastructure Case
LLM GatewayGoverned AI

Sovereign AI

The grounded, cost-governed AI in this platform routes through the same kind of central LLM gateway — read how one control layer governs every model call.

Explore the Sovereign AI case
Explore the Operational AI case: Operational AI
Related Operational Case
In-App AIWorkflow Automation

Operational AI

See the same embedded, human-in-the-loop approach applied inside day-to-day operational workflows rather than product intake.

Explore the Operational AI case

explore further

Related capabilities

  • Operational AI

    AI systems that integrate with existing platforms and workflows, with control, traceability, and operational reliability.

  • Systems Architecture

    Designing architectural foundations that allow complex organizations to operate reliably and evolve safely.

Further reading

  • From vibe coding to final product

    A vibe-coded prototype is not production software, but it can be a powerful input. Used well, AI-assisted prototypes become a direct path from business intent to final product implementation.

  • AI-assisted implementation with frontier models

    Frontier models can accelerate implementation, but only when they are used inside a disciplined delivery method: clear architecture, review, testing, security, and production ownership.

Interested in how this approach could work for your organization?

Get in touch
core purpose. techTechnology consulting with purpose.