Redefining Technology

Retail & E-CommerceAI-Driven Disruptions & Innovations

Multimodal AI models in retail and e-commerce: from pixels to published product claims

Multimodal AI models are systems that read images, video, text, speech and behavioural signals together rather than one modality at a time. In retail and e-commerce they enrich catalogues, power visual and conversational search, and triage photo-evidenced returns — but they only pay when every claim they make is grounded against the product record.

Generated scene: an e-commerce content operations desk where product photography, video frames and catalogue attribute records are reviewed side by side
Retail & E-Commerce · AI-Driven Disruptions & Innovations

Key takeaways

  1. A multimodal model's commerce value is set by grounding, not fluency. A model that describes a product image beautifully has produced marketing copy; a model whose reading of that image can be reconciled against the GTIN-level attribute record and the supplier specification has produced a product claim you can publish.
  2. The join comes before the model. Until an image, a video, a review and a session can be resolved to one product identity, there is nothing to train against, nothing to score against and nowhere to write back to — and every vendor accuracy claim is measured on their assortment, not yours.
  3. Precision floors differ by attribute class. Colour and style can auto-publish at a high-but-reachable precision; dimensional and compatibility attributes need a source citation; composition, allergen, hazard and origin claims should never be published from imagery alone at any confidence.
  4. Most retailers stall at the bolt-on stage because there is no golden set. Without a human-verified answer set at SKU level you cannot justify a confidence threshold, so a merchandiser re-checks everything, the cost per enriched SKU never falls, and the programme is quietly cancelled at budget.
  5. Rights and disclosure are the binding constraint on generative multimodal content, not model capability. Image licensing, consent for a person's likeness, personal data arriving inside customer photos, and the EU AI Act's transparency duties decide what can ship long before accuracy does.

Abbreviations used on this page

VLM
Vision-language model — a model that reads images and text in one representation
PIM
Product information management system — the record of what a product is
DAM
Digital asset management system — where product imagery and video live
PDP
Product detail page
OMS
Order management system
GTIN
Global Trade Item Number — the GS1 product identifier
GPC
GS1 Global Product Classification — the shared category and attribute taxonomy
OCR
Optical character recognition — reading text off packaging and labels
ASR
Automatic speech recognition — transcribing voice and call audio
UGC
User-generated content — shopper photos, videos and reviews
ANN
Approximate nearest neighbour — the index behind embedding search
CVR
Conversion rate

Free · 8 questions · ~3 minutes

Score your multimodal grounding

Eight questions, one at a time, about three minutes. Answer them and we build your personalised grounding report — your stage on the ladder, your score on each of the four dimensions, and the specific blocker standing between you and the next stage — and send it to your inbox. Your result doubles as the baseline for the enrichment business case.

0 of 8 answered

Question 1 of 8Modality joins & identity

Can a product image in your DAM be resolved to a GTIN or SKU without a person opening the file?

Nothing multimodal can be trained, scored or written back until pixels and product records share an identity.

How the score maps to a stage
  • 05 — Stage 1, Single-modality. Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
  • 611 — Stage 2, Bolt-on. Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
  • 1216 — Stage 3, Grounded. Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
  • 1721 — Stage 4, Shared representation. Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
  • 2224 — Stage 5, Bounded autonomy. Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.

What multimodal AI models are — and what changes in a catalogue

A definition, the separation between what ships today and what is still research, and the path a pixel has to travel before it becomes a product claim you can publish.

Multimodal AI models are models that take more than one kind of input into a single representation — an image and a sentence, a video and a transcript, a photograph and a structured product record — and reason across them rather than about each separately. In retail and e-commerce that matters for a mundane reason: almost everything a shopper wants to know about a product is in the photograph, and almost nothing about the photograph is in the database.

The commercially useful subset is narrower than the phrase suggests. Four things are in production across the sector today: attribute extraction and catalogue enrichment from imagery and supplier documents; visual and conversational search, where a photograph or a described-not-named query is matched against the assortment; generated and AI-edited product and marketing imagery; and triage of customer-supplied photographs in returns, claims and marketplace moderation. Everything else — from fully agentic buying to models that understand a garment's drape from a still — is at various distances from a shipping system, and this page keeps that separation explicit.

The organising idea underneath all four is grounding. A multimodal model produces a proposition: this garment is ribbed; this cable fits that socket; this returned item is damaged. A retailer cannot publish a proposition — it publishes a claim, which is a proposition plus the evidence that it holds. The distance between the two is the whole engineering problem, and it is why two retailers using the same model get results a factor apart.

Value released against position on the grounding ladder

The curve is not linear. Value stays close to flat through stages 1 and 2 — where most retailers are — and inflects at stage 3, when output starts being written into the product record under a defensible gate rather than reviewed by a person. This is why programmes that measure progress in images processed report activity without result.

Commercial value released by stage

  • Stage 1 · Single-modality — 21% of operators. Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
  • Stage 2 · Bolt-on — 38% of operators. Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
  • Stage 3 · Grounded — 26% of operators. Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
  • Stage 4 · Shared representation — 12% of operators. Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
  • Stage 5 · Bounded autonomy — 3% of operators. Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with McKinsey's State of AI adoption research.

How a pixel becomes a published product claim

The path from a product photograph to a live attribute, stage by stage. The stage is set by where the arrow ends: at stages 1–2 it ends at a person retyping, at stages 3–4 it ends in the PIM behind a per-attribute gate with provenance, and at stage 5 a bounded set publishes unattended. Most retailers are in the top lane.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • At stages 1–2 a vision-language model reads a hand-picked set of images and produces tags into a spreadsheet or a staging queue. Nothing carries a GTIN or a provenance record, so a merchandiser retypes whatever they accept and the cost per enriched SKU does not move. This is where the value leaks, and it leaks quietly because the output reads well.
  • At stages 3–4 assets, copy, reviews and behaviour are joined on the product identifier. The model proposes a value and a confidence per attribute; a maintained golden set says what that confidence is worth for that attribute; a gate set per attribute class passes the confident values straight into the PIM with full provenance and queues the rest for a merchandiser to adjudicate.
  • At stage 5 a versioned policy lets an enumerated set of judgements publish or settle unattended — cosmetic attributes below a value threshold, photo-evidenced claims under a ceiling. Anything outside the bounds escalates to a person, and every automated judgement can be reconstructed months later with its model version, its source pixel and the policy version in force.
Step-by-step insights
Images in the DAM — the asset that cannot be found
The single most predictive artefact of a stalled multimodal programme is a DAM organised by campaign. Assets filed as "AW26_Shoot3_Final_v2_USE_THIS.jpg" cannot be joined to a product, which means they cannot be training data, cannot be scored, and cannot be written back from. Retailers routinely discover that the join is not one problem but three: assets without identifiers, identifiers without assets, and the far more awkward middle case of one lifestyle image carrying four sellable products and belonging, arguably, to all of them. Resolve that before any model conversation; it is weeks of unglamorous work that every later stage depends on.
The demo trap — fluency is not evidence
A vision-language model reading forty curated photographs is a persuasive artefact and a near-useless one. The curated set has clean backgrounds, single products, consistent lighting and categories the model has seen a great deal of. Your long tail has none of that. Worse, generated attribute text is grammatical and confident whether it is right or wrong, so the usual human error-detection reflex — that something reads oddly — never fires. This is the specific reason content AI needs harder evidence than forecasting AI, not softer: a wrong number looks wrong, a wrong sentence does not.
The join — identity before intelligence
Joining assets to the GTIN is what turns a pile of media into a dataset. Once an image resolves to a product, everything else follows cheaply: the reviews attached to that product become supervision, the returns reason codes become labels, the clickstream becomes a relevance signal, and the supplier specification becomes the grounding evidence. GS1 identifiers are the natural spine because they already reach outside the business to suppliers and marketplaces, which is where most of your attribute disagreements will come from.
The golden set — the artefact that unlocks everything downstream
A golden set is a few hundred to a few thousand SKUs whose attributes have been verified by a human against the physical product or the supplier specification. It is boring to build and it is the highest-leverage week in the entire programme, because it converts every subsequent argument from opinion to measurement: which model, which prompt, which threshold, which attributes can be automated and which cannot. Build it per category, refresh it as the assortment turns, and treat any model or prompt change as a release that has to clear it.
The gate — one threshold per attribute class, never one global number
A single global confidence threshold is the most common design error at stage 3 and it fails in both directions at once: too low for composition and allergen claims, too high for colour and pattern, so you are simultaneously over-exposed and under-automating. Gate per attribute class, with a floor derived from the golden set and a consequence analysis behind it. Some classes should have no threshold at all that permits automatic publication, and saying so in the policy is a stronger control than a very high number.
Provenance and the bounded policy — what an auditor actually reads
Provenance costs almost nothing at write time — model identifier, version, confidence, source asset, clearing check — and is the difference between explaining a bad batch in an afternoon and reconstructing it for a fortnight. At stage 5 the same record becomes external evidence: a marketplace asking why a listing was suppressed, or a consumer body asking how a claim was substantiated, is asking to read the policy version and the decision trail, not the model. Version the policy, pin the model, and log the judgement.

The five stages of multimodal grounding

For each stage: what it looks like inside a real catalogue team, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps retailers there, and what leaving costs.

Each stage below describes how far a multimodal model's output has travelled toward being a claim the business will stand behind. The hallmarks are observable conditions in the PIM, the DAM and the search stack; the diagnostic signals are checks you can run against your own systems this week; and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Single-modality

21% of operators sit here

Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.

Stage 1 is not a shortage of pixels. A mid-sized retailer already holds tens of thousands of product images, thousands of hours of video, years of call recordings and a decade of shopper photographs attached to reviews. What is missing is any path from those assets to a decision: the images are addressed by filename, the transcripts sit in a contact-centre archive nobody queries, and the product record is a separate universe maintained by hand.

The clearest tell is the search box. Type a description rather than a name — "black ribbed midi with a side slit", "the cable that fits the older model" — and a lexical index returns either nothing or the wrong department, because the words a shopper uses are not the words in the title and the information that would answer them is sitting in the photograph. Retailers at this stage usually know their search is weak; what they have not yet connected is that the fix is a data-joining problem before it is a model problem.

This stage is cheap to leave and expensive to sit in, because the cost is invisible. Nobody books "the enrichment we could not do" or "the long tail that never got attributes", so the deficit shows up somewhere else: agency invoices for product copy, a returns rate the category manager blames on suppliers, and a search team that keeps buying synonym packs.

In practice

The forty-SKU demo

A homewares retailer sat through a vendor demo in which a vision-language model read forty hand-picked product photographs and produced clean, plausible attributes for each. Everyone in the room was impressed. Nobody could answer the two questions that mattered: what it would do on the 12,000 SKUs in the long tail, because there was no verified answer set to compare against; and how a value would ever reach the PIM, because no image in the DAM could be resolved to a GTIN without a person opening the file.

What it looks like

  • Product imagery lives in the DAM organised by campaign and filename, not by GTIN
  • Site search is lexical: a shopper's photo, or a described-not-named query, returns nothing useful
  • Catalogue attributes are typed by merchandisers or accepted from supplier feeds unchecked
  • Any multimodal work is a vendor demo on a hand-picked set of SKUs

Diagnostic signals you can check this week

  • Ask what share of live SKUs carry every attribute their category template requires — the honest number is usually far below the reported one
  • Search your own site for a described-not-named product and count how many of the first twenty results are usable
  • Pick a product photograph at random and ask whether anything but a human can tell you which GTIN it belongs to
  • Ask who holds the rights record for that photograph. If the answer is "the agency, probably", generative work cannot start here

Anti-pattern · Buying the model before building the join

The instinct is to procure a multimodal enrichment platform, because the demo is compelling and the pricing is per SKU. It fails for a structural reason: the vendor's accuracy is measured on their evaluation set and your assortment is not in it, and even a perfect output has nowhere to land while assets and product records share no identifier. Build the asset-to-GTIN join and a small verified answer set on one category first. That work is unglamorous, takes weeks rather than quarters, and it is what turns every subsequent vendor conversation from a demo into a measurement.

What holds you here

Images, video, reviews and behaviour are not joined to a single product identity, so nothing can be trained, scored or written back.

Highest-leverage next move

Build the asset-to-GTIN join and a rights register for one category, then a small human-verified answer set for the attributes that actually matter in it.

Cost of leaving

Effort
2–4 months
Team
One data engineer, one merchandising lead, part-time
Risk
Low — nothing customer-facing changes and no product record is written
To next stage
2–4 months

If this is you, the next step is

A two-week engagement: map assets, product records and rights, then size the join.

Scope the asset-to-GTIN join

Stage 2

Bolt-on

38% of operators sit here

Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.

Stage 2 is where most retailers are, and it is the most misleading stage on the ladder because everything about it looks like progress. A model is live. It produces output daily. The output is visibly better than nothing. And the cost per enriched SKU has not moved at all, because a human still reads every line before it is trusted — which was the expensive part in the first place.

The structural cause is the absence of an answer set. Without a human-verified reference for what the attributes on a given SKU actually are, there is no way to say that the model is right 96% of the time on colour and 71% of the time on material. And without that split, no threshold can be defended: a merchandising director asked to let a machine write into the catalogue at 90% confidence will reasonably ask what 90% means in returns, and at stage 2 nobody can answer.

Time spent here is not neutral. Merchandisers learn that the model produces work for them rather than removing it, and the next proposal is heard against that memory. The pattern is close to the dashboard trap in other AI programmes, with one twist specific to content: because generated copy is pleasant to read, quality problems hide better here than in a forecast. A wrong number looks wrong; a wrong sentence reads fine.

In practice

The alt-text backlog that came back

A fashion retailer generated alt text for 60,000 product images to close an accessibility gap ahead of an audit. The copy was fluent and the backlog cleared in a fortnight. Three months later a manual sample found a recurring failure: on flat-lay and lifestyle shots the model frequently described the styling props — the chair, the plant, the second model's jacket — as if they were the product. The alt text was accessible, well written and, for a meaningful slice of the catalogue, about the wrong object. There had never been a verified reference to score it against, so nothing had flagged it.

What it looks like

  • A VLM generates tags, alt text or draft descriptions into a staging queue
  • Every output is human-reviewed, because none of it is scored
  • Nothing is written to the PIM under a rule — a merchandiser retypes what they accept
  • Quality is discussed in impressions ("it's pretty good") rather than per-attribute precision

Diagnostic signals you can check this week

  • Ask for per-attribute precision on the last thousand generated values. A single overall figure, or an anecdote, means you are here
  • Count how many generated values reached the PIM without a human retyping them. If it is zero, the model has added review work, not removed it
  • Compute cost per enriched SKU including review time, then compare it with the agency or offshore quote it was meant to replace
  • Ask which model version produced last quarter's descriptions. If nobody knows, there is no provenance and no way to explain a bad batch

Anti-pattern · Prompt-engineering past a measurement problem

When output is uneven the reflex is to rewrite the prompt, and it works often enough to be seductive: the failures you happened to look at go away. What has actually happened is that the prompt has been fitted to a handful of remembered examples, with no way to know what it broke elsewhere in the assortment. Every prompt change is a model change and needs the same evidence as one. Build the golden set first; then prompt, fine-tune or swap models freely, because you can see the effect per attribute rather than per anecdote.

What holds you here

There is no human-verified answer set, so no confidence threshold can be justified and every output must be re-checked by a person.

Highest-leverage next move

Build a golden set on one category, score the model per attribute rather than overall, and set precision floors by attribute class before writing anything.

Cost of leaving

Effort
3–6 months
Team
One ML engineer, one merchandising QA lead, a named category owner
Risk
Medium — the first automated write into the PIM needs a fallback to the supplier-supplied value
To next stage
3–6 months

If this is you, the next step is

Ten working days to a 1,000-SKU verified answer set and defensible per-attribute floors.

Build the golden set

Stage 3

Grounded

26% of operators sit here

Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.

Stage 3 is the first stage at which multimodal work changes the unit economics of the catalogue. The mechanism is narrow and specific: for each attribute the model proposes a value with a confidence, the value is checked against whatever independent evidence exists — the supplier specification sheet, the GS1 category template, the packaging text read by OCR — and only values above that attribute class's floor are written. Everything else queues. The merchandiser's day changes from typing to adjudicating, and adjudication is a far smaller job.

The discipline that makes it work is provenance. Every written attribute carries the model identifier, the version, the confidence, the source asset and the check that cleared it. That record costs almost nothing to produce at write time and is the difference between a supportable catalogue and an unexplainable one. When a category manager asks why 400 SKUs suddenly claim a fabric they do not have, provenance turns a forensic exercise into a filtered query.

The constraint that emerges at stage 3 is duplication. Search has built its own embedding, the content team has its own captioner, the support assistant has its own retrieval index, and the returns team is evaluating a third vendor. Each works. None compounds, and worse, the same product now means slightly different things in four places — which is how a shopper reads one thing on the PDP, is told another by the assistant and receives a third in the box.

In practice

The category that stopped being retyped

A homewares team started with one category — occasional furniture — and the twelve attributes that drove its return reasons: dimensions, assembly requirement, material family, weight class and finish. They verified 1,100 SKUs by hand, scored the model per attribute, and found precision ranged from the high nineties on finish and colour to barely two-thirds on material. They gated accordingly: finish and colour auto-published, dimensions published only where OCR of the supplier sheet agreed, and material always queued. Roughly seven in ten attribute values cleared the gate. The queue that remained was a morning's work a week rather than a permanent headcount.

What it looks like

  • Attributes are written to the PIM with model, version, confidence and source asset recorded
  • The confidence gate is set per attribute class — regulated attributes never auto-pass
  • Merchandisers work a queue of the low-confidence tail instead of re-checking everything
  • Value is reported as attribute completeness, null-result rate and return-reason movement, not model accuracy

Diagnostic signals you can check this week

  • Open any AI-populated attribute in the PIM and ask which model version and which source asset produced it — the answer should be one click, not an investigation
  • Check whether the gate uses one global threshold or different floors for colour and for allergen. One global threshold means the classes have not been thought about
  • Ask what happens on modality dropout: a SKU with no photograph, or a lifestyle shot with three products in it
  • Ask whether search, the PDP and the contact-centre assistant read the same attribute record. Three answers means three truths

Anti-pattern · One model per surface

Each team ships the multimodal capability it needs, and each is individually justified: search wants an embedding tuned for retrieval, content wants a captioner tuned for tone, support wants an assistant tuned for policy. The cost lands later and lands on the customer, as inconsistency — the assistant confidently describing a colourway the PDP does not list. Consolidating after four surfaces exist is a migration; agreeing one shared product representation while there are two is an afternoon. The signal to act is the second surface, not the fourth.

What holds you here

Every surface has its own model and its own embedding, so the same product means different things in search, on the PDP and in support — and nothing compounds.

Highest-leverage next move

Serve one shared product representation — a single embedding index and a single attribute record with provenance — to every surface that needs it.

Cost of leaving

Effort
6–12 months
Team
ML engineer, catalogue/PIM engineer, merchandising owner, plus a compliance partner for the regulated attribute classes
Risk
Medium — the risk shifts from accuracy to consistency across surfaces
To next stage
6–12 months

If this is you, the next step is

We map what each surface reads today and what a shared representation would have to guarantee.

Consolidate onto one product representation

Stage 4

Shared representation

12% of operators sit here

Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.

At stage 4 the marginal cost of a new multimodal surface collapses, and the conversation changes shape with it. Teams stop proposing "an AI project for returns" and start asking which decisions should read the product representation next. That is a healthier conversation because it is about decisions and metrics rather than about models, and it is the reliable signature of this stage.

The measurement discipline is what separates stage 4 from a well-engineered stage 3. Each surface has a named commerce metric and, where the surface allows it, a holdout: a category left on the previous ranking, a market left on supplier-supplied attributes, a returns queue left on manual triage. Holdouts are harder in retail than in logistics because seasonality and promotions move everything at once, which is exactly why the comparison group has to exist rather than be argued from a year-on-year chart.

The governance work also consolidates here, and this is the part most programmes underestimate. Rights and consent for imagery, disclosure of synthetic content, personal data arriving inside customer photographs, and accessibility of generated alt text are all properties of the representation layer, not of the individual surface. Handled once, they are a register and a set of checks. Handled per team, they are four incompatible answers to the same regulator's question.

In practice

Returns triage in three weeks

An electricals retailer already served one product representation to search, the PDP and its assistant. When the returns team asked for automated triage on photo-evidenced claims — is the item damaged, is it the item that was ordered, is the accessory in the box — the build was three weeks, most of it spent agreeing the metric and the holdout with operations rather than writing code. The representation, the rights handling and the review queue already existed. That ratio, specification-heavy and build-light, is the stage-4 tell.

What it looks like

  • One embedding index and one attribute record serve search, recommendations, PDP content, support and returns
  • A new surface ships in weeks because the representation already exists
  • Every surface is evaluated against a holdout in its own commerce metric
  • Disclosure, rights and personal-data controls are handled once in the representation layer, not per team

Diagnostic signals you can check this week

  • Measure elapsed time from "we want visual search in the app" to it serving traffic, for the last two surfaces
  • Change an attribute definition in one place and check whether every surface picks it up, or whether three teams have to be told
  • Ask whether any multimodal surface has a live holdout rather than a year-on-year comparison
  • Ask whether the rights and disclosure register covers generated assets on every channel, including marketplaces and paid social

Anti-pattern · Automating because the representation is good

A gate working well on colour and finish creates pressure to extend it to compatibility, composition and origin, because the machinery is identical and the volumes are tempting. The classes are not comparable. Cosmetic attributes fail into a poor facet; composition and origin claims fail into consumer-protection and labelling exposure, and a single wrong allergen inference is a different category of event from a thousand wrong colour tags. Each attribute class earns automation from its own evidence, with its own floor and its own sign-off.

What holds you here

The model still only reads and proposes; the remaining value needs it to publish, suppress or settle — which needs bounded policy and audit evidence, not more accuracy.

Highest-leverage next move

Enumerate the multimodal decisions allowed to execute unattended, each with a value ceiling, a precision floor and a reconstructable trail.

Cost of leaving

Effort
12–24 months
Team
Platform team, merchandising product owner, legal and compliance partner
Risk
Higher — a representation error propagates to every surface at once
To next stage
12–24 months

If this is you, the next step is

Which multimodal judgements may publish or settle unattended, and the evidence that makes it defensible.

Design your multimodal action policy

Stage 5

Bounded autonomy

3% of operators sit here

Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.

Stage 5 is much narrower than the word autonomy suggests. It is not a self-running catalogue; it is a short, enumerated list of judgements that may execute unattended inside stated bounds, with everything outside those bounds escalating. Colour and finish on a low-value SKU qualify. A photo-evidenced damage claim under a value ceiling qualifies. Allergen inference, origin claims and anything with a safety or labelling consequence are correctly held at stage 4 permanently, and a mature operator says so in the policy rather than leaving it to a threshold.

The engineering is largely finished by the time an operator arrives here. The hard artefact is the evidence: showing a marketplace, a regulator or a consumer body why one specific listing was suppressed nine months ago, which model version and which pixel produced the judgement, and under which version of the policy. Treat the policy with the same rigour as the model — it is the document that will be read out, and it is far more likely to be examined than the weights.

Sustaining this stage is a change-control problem rather than a technical one, and it is the stage most likely to regress. Assortments turn over seasonally, so an evaluation set stops representing the catalogue within a couple of drops. Photography style changes when the agency does. And model providers deprecate versions on their own calendar, which means a capability you validated in March can behave differently in September without a single line of your code changing.

In practice

The bounded publish set

A marketplace operator publishes AI-derived colour, pattern and silhouette attributes unattended for listings below a stated value, and auto-settles photo-evidenced damage claims below a ceiling set per category. Roughly one judgement in twelve escalates. The escalation rate is itself monitored: a rise means the assortment or the imagery has moved outside the policy's validity — a new seller cohort, a new photographic style — and it triggers a re-validation against a refreshed golden set before it triggers a complaint.

What it looks like

  • Cosmetic and dimensional attributes below a value threshold publish unattended
  • Photo-evidenced returns under a value ceiling settle automatically; everything else routes to a person
  • Listings whose imagery contradicts their claims are suppressed pending review, with a logged reason and an appeal path
  • The multimodal policy is versioned and reviewed like code, and model versions are pinned rather than floating

Diagnostic signals you can check this week

  • Check whether the policy names attribute classes and value ceilings explicitly, with a version history and an owner
  • Ask when the kill switch — revert every gated attribute to the supplier-supplied value — was last exercised deliberately
  • Ask whether escalation and appeal rates are watched as leading indicators rather than reported after an incident
  • Pick one suppressed listing from nine months ago and ask for the reason, the model version and the policy version

Anti-pattern · Letting the vendor's release calendar set your policy

A hosted model version is deprecated, the pipeline silently rolls forward to its successor, and behaviour changes mid-season on a catalogue nobody re-scored. It is the most common regression at this stage and it is entirely avoidable: pin versions, subscribe to deprecation notices as an operational feed, and treat every model change as a release that must clear the golden set before it touches the gate. A model swap is a deployment, not a configuration tweak.

What holds you here

Sustaining autonomy is a governance and change-control problem: model versions are deprecated on the provider's calendar, assortments turn over seasonally and disclosure duties move.

Highest-leverage next move

Version the multimodal decision policy, pin model versions, and treat a deprecation as a planned operational event with a re-validation gate.

Cost of leaving

Effort
Continuous
Team
Platform team plus a standing content-governance forum with legal and merchandising
Risk
Concentrated — low frequency, high consequence, consumer-law and platform-policy in nature

If this is you, the next step is

We stress-test the policy, the provenance trail and the revert against a real scenario.

Audit an automated publish path

Where retailers actually sit on the ladder

The distribution across the five stages, and why the bolt-on stage is both the mode and the trap.

Most retailers are at the bolt-on stage. The distribution is heavily weighted toward work where a multimodal model is live and producing output that a human still reads before anything is trusted — which is to say, live but not yet load-bearing. A minority have a defensible gate into the product record, and a very small group let any multimodal judgement stand without a person behind it.

Distribution of retailers across the five grounding stages

Illustrative distribution — a model-derived synthesis, not a survey. Stage 2 is the mode and the plateau: the drop from bolt-on to grounded is the largest single transition loss on the ladder, and it is an evidence problem rather than a model problem.

Share of retailers

  • 21% — 1 · Single-modality
  • 38% — 2 · Bolt-on (the plateau)
  • 26% — 3 · Grounded
  • 12% — 4 · Shared representation
  • 3% — 5 · Bounded autonomy

Source: Illustrative distribution, synthesised from McKinsey, NRF and Baymard research

The shape is not specific to multimodal work — cross-industry research has consistently found a wide gap between organisations experimenting with AI and organisations reporting material bottom-line impact, which is the pattern McKinsey's State of AI (opens in a new tab) has tracked over successive years. What is specific to retail is the shape of the blocker. In logistics the usual gap is a write-back path into an execution system; in a catalogue the write path is comparatively easy and the missing piece is evidence: nobody can say what the model's confidence is worth on this assortment, so nobody will let it write.

That figure is worth sitting with, because it reframes catalogue enrichment from a content-operations cost line into a margin lever. When a shopper returns an item as "not as described", the description is usually not false so much as absent — the attribute that would have set the expectation was never populated, because populating it by hand for a long-tail SKU has never been worth anyone's time. This is the specific gap a grounded multimodal pipeline closes, and it is why the business case belongs to the category manager rather than to the innovation team. Baymard Institute's e-commerce search research (opens in a new tab) makes the same point from the discovery side: shoppers routinely fail to find products that exist in the catalogue, because the attributes they search on were never recorded.

Where multimodal models land in a retail operation

Eight decisions worth wiring, the modalities each one actually needs, the system of record it writes to, the KPI it moves and the stage at which it earns its keep.

Multimodal value in retail concentrates in eight decisions, and each has a natural home on the ladder. A decision is a good first candidate when three things hold: the system of record is one you control, the evidence to ground the model already exists somewhere in the business, and the KPI it moves is one a category manager already reports. The map below is how we scope first and second multimodal use cases with retailers.

DecisionModalities in playSystem of recordKPI it movesEarns its keep
Catalogue enrichment and attribute extractionProduct image + supplier spec sheet (OCR) + existing copyPIM / DAMAttribute completeness, cost per enriched SKUStage 3
Visual search and camera entry pointsShopper photograph + catalogue imagery + clickstreamSearch indexNull-result rate, search-entry CVRStage 3
Conversational and described-not-named searchQuery text + attribute record + review textSearch index / merchandising rulesZero-result queries, add-to-basket from searchStage 3–4
PDP content and accessible alt textProduct image + attributes + brand tone rulesCMS / PIMAccessibility conformance, organic entrancesStage 3
Generated and AI-edited product imagerySource photography + generated variants + rights recordDAMContent cost per SKU, disclosure coverageStage 3–4
Returns and claims triage on photo evidenceCustomer photograph + order record + policy textOMS / returns platformProcessing cost per return, claims fraud lossStage 4
Marketplace and third-party listing integrityListing imagery + title and description + category taxonomyMarketplace / seller platformPolicy-violating listings caught, notice turnaroundStage 4
Shelf, pick and pack verificationCamera stream + planogram + inventory recordStore systems / WMSOn-shelf availability, mispick rateStage 4
The retail multimodal decision map. "Earns its keep" is the grounding stage at which the decision typically becomes worth operating — running a stage-4 decision on a stage-2 evidence base is the fragile quadrant described later on this page.

Catalogue enrichment is where most retailers should begin, for reasons that have nothing to do with it being the most exciting. The system of record is yours, the grounding evidence — supplier specifications, existing verified SKUs, returns reason codes — is already in the building, and the KPI moves inside a quarter. Visual search is the more visible project and the harder first one, because its quality is bounded by exactly the attribute coverage that enrichment produces: a camera entry point on a thin catalogue returns thin results, and the shopper blames the feature.

The compliance column is not a separate domain here — it runs through every row. Product identity and classification lean on GS1 standards (opens in a new tab), and the GTIN is the join key that makes attribute disagreements with suppliers and marketplaces tractable at all; GS1 Digital Link (opens in a new tab) extends the same identifier into the web resources a model can retrieve. Shopper-facing assistants and synthetic content attract transparency duties under the EU AI Act (opens in a new tab), marketplace moderation decisions attract notice-and-action and statement-of-reasons duties under the Digital Services Act (opens in a new tab), and any imagery containing a person pulls in data-protection obligations set out in the ICO's guidance on AI and data protection (opens in a new tab). None of these prohibits multimodal work. All of them assume you can explain a decision afterwards, which is a grounding property.

Why the bolt-on stage stalls: the grounding problem

Three structural patterns account for most of the plateau, none of them a modelling problem — plus the precision floors that decide which attributes may ever be automated.

The bolt-on stage stalls because the pilot was scoped to prove that the model can read an image, and reading the image was never the constraint. Modern vision-language models read product photography well. The constraint is that a retailer cannot act on a proposition it cannot check, and at stage 2 there is nothing to check against — no verified answer set, no supplier specification joined to the SKU, no record of which model version said what.

  • There is no verified answer set, so no threshold can be defended

    Confidence scores are meaningless until they are calibrated against ground truth on your own assortment. Without that, "publish above 0.9" is a number somebody chose in a meeting, and the merchandising director who declines to sign it off is being rigorous rather than obstructive. The golden set is the cheapest artefact on this page and the one whose absence stops everything.

  • The pilot was scoped to a category and the business case needs the tail

    Enrichment pilots are run on a category with good imagery and clean supplier data, because that is where a pilot succeeds. But the value of enrichment is concentrated precisely where the data is worst — the long tail nobody has had time to populate by hand. A pilot on the best-documented 5% of the catalogue proves the model works and proves nothing about the economics.

  • Content quality is discussed in taste, not in precision

    Because generated copy is a matter of tone as well as fact, review meetings drift toward whether the descriptions sound right. Tone is real and it is a brand decision. It is also a different conversation from whether the fibre composition is correct, and merging the two is how factual error survives review. Separate them: tone is signed off once, per template; facts are measured per attribute, continuously.

The remedy is a per-attribute view of consequence. Not every wrong value costs the same, and treating them uniformly is what produces both over-exposure and under-automation at once. The table below is the framework we use with merchandising and compliance teams to decide, class by class, what may ever be automated and what must remain a human or supplier-supplied claim regardless of how confident a model becomes.

Attribute classExamplesWhat a wrong value costsFloor before any auto-publishWho signs the class off
Cosmetic / descriptiveColour family, pattern, silhouette, neckline, finishA poor facet and a slightly worse search resultHigh but reachable — derive from the golden set, then sample continuouslyMerchandising, with sampled QA
DimensionalLength, capacity, drop, weight class, pack quantity"Not as described" returns and support contactsHigher, and only where OCR of the supplier sheet agrees with the imageCategory manager
CompatibilityFits model X, socket type, thread size, cartridge seriesReturns plus a support contact plus a negative reviewHighest, and only with a citation to the supplier specificationTechnical merchant or supplier quality
Composition / materialFibre content, food ingredients, packaging materialLabelling and consumer-law exposure; sustainability claimsNever from imagery alone — document-grounded only, model may proposeSupplier data plus compliance
Regulated / safetyAllergens, hazard pictograms, age restriction, medical claimsLegal exposure, recall, delistingNever auto-published; the model's only role is flagging disagreementCompliance and legal
Provenance claimsCountry of origin, "organic", "recycled content", certificationsGreenwashing and consumer-protection exposureNever inferred — the claim comes from the certificate, not the pixelLegal and sustainability
Precision floors by attribute class. The floors are working defaults to argue from, not universal constants — derive your own from your golden set and your category's consequence profile. The two "never" rows are policy positions, not thresholds: no confidence score makes an inferred allergen or origin claim publishable.

Building the golden set — ten working days

  1. Pick the category by consequence, not by convenience

    Choose the category where missing attributes are already costing you — highest "not as described" return rate, or highest zero-result search volume. The point is not to make the model look good; it is to produce a number a category manager will act on.

  2. Choose the twelve attributes that actually matter

    Read the returns reason codes and the last quarter's search queries for that category. Twelve attributes is usually enough to cover the reasons shoppers return and fail to find, and a set small enough that human verification finishes inside a fortnight.

  3. Verify 800–1,500 SKUs by hand against evidence, not memory

    Verification means against the supplier specification, the physical sample or the packaging — not against the existing PIM value, which is frequently the thing being wrong. Sample across the tail deliberately: the well-photographed hero SKUs are not the population you need to measure.

  4. Score per attribute and publish the confusion, not just the score

    A single accuracy figure hides everything useful. Report precision and recall for each attribute, and keep the disagreements — the systematic ones (styling props read as products, lifestyle shots, multipacks) tell you what to fix in the pipeline rather than in the model.

  5. Set the floors with merchandising and compliance in the room

    Floors are a business decision informed by measurement, not a modelling output. Agree them per attribute class, write down the two or three classes that will never auto-publish, and version the document — it becomes the multimodal decision policy later.

  6. Schedule the refresh against the assortment cycle

    A golden set built in spring stops representing an autumn assortment. Tie its refresh to the buying calendar rather than to a fixed interval, and re-run it whenever the model version, the prompt or the photography style changes.

The AI system's outputs must be sufficiently accurate for the purpose, and organisations must be able to explain the reasoning behind decisions to the people affected by them.

What multimodal deployment looks like in public

Three publicly reported programmes, read against the grounding ladder. None is an Atomic Loops engagement — each links to the operator's own published material.

The most instructive public examples are not the ones with the largest models but the ones where the operator built the surrounding apparatus at the same time — the catalogue, the identity, the rights record. In each case below the differentiator is what sits behind the modality, and the lesson maps onto a specific rung of the ladder.

Three programmes read against the grounding ladder

Outcomes as reported by the operators themselves; verify figures against the linked source before reusing them, and note that dated press releases for two of these programmes no longer resolve, so the stable corporate newsroom is linked instead. Card images are generated industry scenes from our asset library, not operator photography, and imply no endorsement.

Generated scene: a shopper using a phone camera as a search entry point in front of a product displayAmazonGlobal marketplace and retailer · catalogue at internet scale35
Challenge
Shoppers arrive with a photograph or a vague description far more often than with a product name, and a lexical index answers neither. At marketplace scale the problem compounds: much of the catalogue is seller-supplied, so imagery and attributes vary in quality across hundreds of millions of listings.
Approach
Amazon has publicly described two complementary multimodal entry points in its shopping app: Lens, which matches a photograph taken by the shopper against the catalogue, and Rufus, a generative shopping assistant that answers product questions grounded in catalogue data, listings and community content rather than from the model's own recall.
Reported outcome
Amazon reports both features as generally available to customers in its shopping app, and describes Rufus explicitly as answering from product information, listings and customer reviews — that is, as a grounded assistant rather than a free-standing chatbot.
What it shows about the curveThe modality is the entry point; the catalogue is the product. Visual and conversational search are only as good as the attribute coverage and review corpus behind them, which is why enrichment usually has to precede the camera icon rather than follow it.

About Amazon — Amazon Lens (opens in a new tab)

Generated scene: an omnichannel retail operations team reviewing product search results alongside catalogue recordsWalmartGlobal omnichannel retailer · stores plus marketplace34
Challenge
Search and product content across an enormous, largely supplier-fed assortment, spanning grocery, general merchandise and a third-party marketplace — categories whose attribute vocabularies have almost nothing in common with each other.
Approach
Walmart has publicly described building retail-specific language models trained on its own catalogue and customer data, and applying generative AI to search and shopping assistance across its properties, rather than relying solely on general-purpose models with no view of its assortment.
Reported outcome
Walmart publishes ongoing reporting on its AI-powered search and assistant work through its corporate technology channel; treat the specifics there as the authority rather than any secondary summary, including this one.
What it shows about the curveThe defensible asset is the representation — catalogue, behaviour, identity — not the model. The stage-4 signature is exactly this: one representation good enough that new surfaces read from it instead of building their own.

Walmart — technology (opens in a new tab)

Generated scene: a fashion content studio where photographed and generated garment imagery are compared against usage rights recordsH&M GroupGlobal fashion retailer · high assortment turnover23
Challenge
Producing enough on-model imagery for a fast-turning assortment across many markets and channels, where photography cost and lead time constrain how much of the catalogue can be shown well at all.
Approach
H&M Group publicly reported exploring AI-generated digital twins of models — likenesses created with the models' agreement, with the models retaining rights over the use of their digital twin and being remunerated for it, and the resulting imagery labelled as AI-generated where used.
Reported outcome
The programme was reported through H&M Group's own newsroom, which remains the authority on its current scope; the durable point is the shape of the controls rather than any volume figure.
What it shows about the curveFor generative multimodal content, the deployment work is the governance work. Consent, rights and disclosure are what decide whether an asset can ship — capability stopped being the constraint some time ago.

H&M Group newsroom (opens in a new tab)

Read together, the three make one argument. Amazon's assistant is described as answering from listings and reviews — grounding, stated as a product property. Walmart's investment is in a retail-specific representation rather than in a bigger general model. H&M Group's reported programme is notable less for the imagery than for the consent, rights and labelling apparatus around it, which is the part a European retailer has to build regardless of which model it picks. See About Amazon's own description of Rufus (opens in a new tab) for the grounding claim in the operator's words.

The four dimensions that set your stage

Multimodal maturity is not one number. Four dimensions gate each other, and the lowest is the real stage — which in retail is more often rights than model quality.

Multimodal maturity is not a single number. A retail operation is scored on four dimensions — modality joins and identity, grounding and evaluation, write path and surfaces, and rights, transparency and safeguards — and the lowest of the four is the real stage, because each one gates the others. An excellent model served from unjoined assets writes nothing; a beautifully gated pipeline with no rights register cannot publish a generated image in a regulated market.

  • Modality joins and identity

    Whether an image, a video, a review, a call transcript and a session can all be resolved to one product and one customer. The binding question is whether a machine can do it without a person opening a file. Until it can, the other three dimensions have nothing to operate on, and GS1 identifiers are the usual spine because they already reach suppliers and marketplaces — see GS1's standards (opens in a new tab).

  • Grounding and evaluation

    Whether a proposed value can be checked against independent evidence, and whether you know per attribute how often the model is right on your assortment. This dimension is where the stage-2 plateau lives, and it is almost always the cheapest to fix — a golden set is measured in person-days, not quarters.

  • Write path and surfaces

    Where output lands, and how many representations your customer-facing surfaces read from. One shared representation is the difference between a capability and four capabilities that disagree. The tell is what happens when someone changes an attribute definition: one change and everything follows, or three tickets and a reconciliation meeting.

  • Rights, transparency and safeguards

    Whether you can say, per asset, who holds the rights, whether a person consented to their likeness, whether the asset is disclosed as AI-generated where required, and how personal data arriving inside imagery is handled. Retailers consistently score lowest here and consistently discover it last, because nothing breaks until a market, a marketplace or a regulator asks — and the risk-management practices in the NIST AI framework (opens in a new tab) map onto this dimension almost directly.

Diagnosing the real constraint

Plot your grounding evidence against your write path. The quadrant names the next investment — and three of the four answers are not "get a better model".

Publishing what nobody checked

  • Output reaches the catalogue with no per-attribute evidence
  • The most dangerous quadrant — errors are fluent and durable
  • Fix: build the golden set before widening the gate an inch

Compounding

  • Evidence and write path both in place
  • Constraint moves to consistency across surfaces
  • Fix: consolidate onto one shared representation

Demo shelf

  • Neither foundation in place
  • Common at stage 1 and early stage 2
  • Fix: join assets to the GTIN on one category

Well measured and unused

  • Good per-attribute evidence, output goes nowhere
  • Highest-leverage position on the matrix
  • Fix: build the gated PIM write, not another evaluation
Write path — top: Into the product record under a gate, bottom: Staging queue or spreadsheet
Grounding & evaluation — left: No verified answer set, right: Per-attribute evidence, maintained

The two lower-left-to-upper-right diagonals are the interesting ones. "Well measured and unused" is frustrating but cheap to resolve: the hard artefact already exists and what remains is integration work with a known shape. "Publishing what nobody checked" feels like progress and is the position from which retailers get hurt, because generated attribute errors are grammatical, plausible and propagate to every surface that reads the record — including the marketplaces and comparison feeds you do not control.

The reference architecture, layer by layer

What has to exist for each stage of multimodal grounding — defined by what each layer must guarantee, not by which product provides it.

A grounded multimodal capability needs five layers, and the order in which you build them decides whether the programme compounds or stalls. The architecture below is deliberately vendor-neutral: nothing in it names a product, and every layer is specified by the guarantee it owes the layer above.

Layers required by grounding stage

Each layer is annotated with the stage that first requires it. A programme aiming at stage 3 without the identity and evaluation layers is running a stage-2 bolt-on with more infrastructure.

  1. Sources and assets

    Stage 1+

    • DAMProduct photography, video, generated variants
    • PIM and supplier feedsSpecifications, category templates, certificates
    • Reviews, UGC and transcriptsShopper photography, review text, contact-centre audio
    • Behavioural eventsSearch queries, clickstream, returns reason codes
  2. Identity and rights

    Stage 2+

    • GTIN / SKU resolutionOne product identity across every source
    • Asset-to-product joinIncluding the multi-product lifestyle shot case
    • Rights and consent registerLicence, term, territory, likeness consent, per asset
  3. Representation

    Stage 3+

    • Multimodal embeddings + ANN indexOne index, served to every surface
    • Attribute candidates with confidencePer attribute, never one score per SKU
    • Golden evaluation setHuman-verified, refreshed on the buying calendar
  4. Decision and write

    Stage 3+

    • Confidence gate by attribute classFloors derived from the golden set, classes that never pass
    • Write with provenanceModel, version, confidence, source asset, clearing check
    • Human review queueThe tail only — adjudication, not re-typing
    • Fallback to supplier valueThe previous value, one switch away
  5. Governance and transparency

    Stage 4+

    • Synthetic-content disclosurePer asset and per channel, including marketplaces
    • Personal-data controls for imageryLawful basis, retention, redaction, deletion path
    • Model version pinning and change logDeprecations handled as planned releases
    • Decision audit trailReconstructable months later, per judgement

Pipeline described

  1. Sources and assets (stage 1+) — DAM: Product photography, video, generated variants; PIM and supplier feeds: Specifications, category templates, certificates; Reviews, UGC and transcripts: Shopper photography, review text, contact-centre audio; Behavioural events: Search queries, clickstream, returns reason codes
  2. Identity and rights (stage 2+) — GTIN / SKU resolution: One product identity across every source; Asset-to-product join: Including the multi-product lifestyle shot case; Rights and consent register: Licence, term, territory, likeness consent, per asset
  3. Representation (stage 3+) — Multimodal embeddings + ANN index: One index, served to every surface; Attribute candidates with confidence: Per attribute, never one score per SKU; Golden evaluation set: Human-verified, refreshed on the buying calendar
  4. Decision and write (stage 3+) — Confidence gate by attribute class: Floors derived from the golden set, classes that never pass; Write with provenance: Model, version, confidence, source asset, clearing check; Human review queue: The tail only — adjudication, not re-typing; Fallback to supplier value: The previous value, one switch away
  5. Governance and transparency (stage 4+) — Synthetic-content disclosure: Per asset and per channel, including marketplaces; Personal-data controls for imagery: Lawful basis, retention, redaction, deletion path; Model version pinning and change log: Deprecations handled as planned releases; Decision audit trail: Reconstructable months later, per judgement
Step-by-step insights
Sources and assets — the transcripts nobody counts
Most retailers inventory their imagery and forget three sources that are already multimodal supervision: shopper photographs attached to reviews, contact-centre audio, and returns reason codes with free-text notes. Each is a labelled dataset about the gap between what the catalogue said and what arrived. Speech recognition over a quarter of contact-centre audio, joined to order records, routinely surfaces the specific attributes that generate contacts — which is a far better prioritisation input than asking merchandisers which attributes matter.
Identity and rights — build the register while the assets are new
The rights register is trivial to maintain going forward and near-impossible to reconstruct backwards. For each asset: who created it, under what licence, for which territories and how long, whether an identifiable person appears and what they consented to, and whether the asset is wholly or partly generated. Retailers that skip this discover the cost when they want to use existing photography as training or conditioning input for generated variants, and cannot establish whether they are allowed to.
Representation — one index, and resist the second
The pressure to build a second embedding index is constant and always locally reasonable: search wants retrieval-tuned vectors, recommendations want behavioural ones, support wants document chunks. Serve one representation with surface-specific heads or re-rankers on top rather than duplicating the base. The test is a question you should be able to answer instantly: if a product's imagery is replaced tomorrow, how many indexes have to be rebuilt before every surface agrees again?
Decision and write — the fallback is what unlocks the approval
The component most often skipped is the revert: the supplier-supplied value, restorable in one action per attribute or in bulk. It reads as engineering pessimism and functions as the political key to the write approval, because a merchandising director will accept a machine writing into the catalogue if the previous state is one switch away. A gated-write proposal without a drilled revert sits in a change queue; one with it ships.
Governance and transparency — a register, not a review board
Disclosure, personal data and model versioning are best implemented as records the pipeline emits rather than as meetings. Every generated asset carries its disclosure state and channel permissions; every image containing a person carries a lawful basis and a retention date; every model version has a pinned identifier and a change log entry. Built this way, a regulator's question is a query against the register. Built as a governance forum, it is a fresh project every time somebody asks, and the answer arrives weeks after it was useful.

The layer retailers most often defer is identity and rights, because it produces nothing visible. It is also the layer whose absence caps everything above it: without the asset-to-product join there is no training set, no evaluation, no write-back and no way to answer whether a given image may be used to condition a generated variant. Build it first even though it demos badly.

A 90-day plan: cutting "not as described" returns in one category

The bolt-on to grounded transition made concrete on one retail problem — missing dimensional attributes driving returns in a single category. Contains no model development.

Moving one stage takes about 90 days when it is scoped to a single category and a single failure, and multiple years when it is scoped to "the catalogue". To make that concrete, the plan below runs the transition on a specific, common retail problem: items returned as "not as described" because the dimensional and material attributes that would have set the shopper's expectation were never populated. The model already exists at most stage-2 retailers, so the quarter contains no model development at all — it is joins, evidence, a gate and an attribution.

Bolt-on to grounded on one category, in one quarter

One category, twelve attributes, one named owner. If a phase needs more than its window, narrow the scope — fewer attributes, fewer SKUs — rather than extending the plan.

  1. Days 1–15

    Pick the category and baseline the loss

    Choose the category with the highest "not as described" return rate. Pull twelve months of return reason codes, current attribute completeness against the category template, and zero-result search queries for that category. Build the asset-to-GTIN join for its SKUs. Name the category manager as owner — returns and margin are their numbers, not the AI team's.

    A baseline return rate, completeness figure and a working asset join

  2. Days 16–40

    Build the golden set and score per attribute

    Human-verify 800–1,500 SKUs across the tail — not the hero products — on the twelve attributes that drive the return reasons, against supplier specifications and physical samples rather than the existing PIM values. Score whichever model you already use per attribute, keep the systematic disagreements, and agree the precision floors with merchandising and compliance in the room.

    Per-attribute precision and signed-off floors by attribute class

  3. Days 41–70

    Ship the gated write with provenance and a revert

    Write above-floor values into the PIM carrying model, version, confidence and source asset. Queue the tail to merchandisers as adjudication rather than re-typing. Leave the supplier-supplied value one switch away, and exercise the bulk revert once deliberately on a quiet day. Record disclosure and rights state for any generated imagery in scope.

    Attributes flowing under a gate, with a drilled revert

  4. Days 71–90

    Attribute the result against a holdout

    Hold out a comparable category — similar price band, similar return profile — on supplier-supplied attributes. Report the difference in "not as described" return rate, attribute completeness, zero-result search share and cost per enriched SKU. Report model precision separately and last; it is the input, not the outcome.

    A returns and cost delta the finance team will accept

The order matters

  1. Join before model

    Every day spent evaluating models before assets resolve to products is a day spent on a decision you cannot act on. The join is unglamorous, finishes in weeks, and turns every later vendor conversation into a measurement rather than a demo.

  2. Evidence before threshold

    A confidence threshold with no golden set behind it is a number someone chose. Merchandising directors are right to refuse it, and the fastest route through that refusal is measurement, not persuasion.

  3. Gate before autonomy

    Keep the merchandiser adjudicating the tail through the first quarter even where the precision would permit more. The adjudication log — which proposals were accepted, corrected and how — is the dataset that sets safe bounds later, and there is no way to acquire it retrospectively.

  4. One category before the catalogue

    The shared representation is worth building when the second and third categories are asking for the same joins. Building it before the first has an attributed number encodes guesses as architecture.

Instrumenting multimodal work: formula, source, cadence

Where each metric actually comes from — the formula, the system that produces it, and the stage at which it starts measuring something real. All telemetry, no self-report.

A multimodal metric you cannot name a source system for is an opinion, and this is the area of retail AI most prone to opinion because the output is prose. Every metric below reduces to counts and comparisons that the PIM, the DAM, the search platform or the model-serving layer already records. The table is the build sheet: formula, source, cadence, and the grounding stage at which the metric first becomes honest.

MetricFormula / readSourceCadenceHonest from
Attribute completenessPopulated required attributes ÷ required attributes, per category templatePIMWeeklyStage 1
Per-attribute precisionCorrect proposals ÷ proposals made, scored against the golden setGolden set + serving logPer model or prompt releaseStage 2
Gate pass rateValues published above floor ÷ values proposedGate logDailyStage 3
Tail clearance timeMedian age of items in the merchandiser review queueReview queueWeeklyStage 3
Cross-modal disagreement rateProposals contradicting the supplier spec or OCR ÷ proposalsGrounding check logDailyStage 3
Cost per enriched SKU(Compute + review minutes × loaded rate) ÷ SKUs enrichedServing log + queue timingsMonthlyStage 3
Zero-result search shareSearches returning nothing ÷ searches, for the covered categorySearch platformWeeklyStage 3
"Not as described" return rateReturns on that reason code ÷ orders, versus a holdout categoryOMS / returns platformMonthlyStage 3
Alt-text coverageImages with non-empty, product-accurate alt text ÷ imagesCMS / PIMMonthlyStage 3
Disclosure coverageGenerated assets carrying a disclosure state ÷ generated assets publishedDAM / rights registerWeeklyStage 4
Model version ageDays since the serving version last cleared the golden setRelease logWeeklyStage 4
Unattended publish rateJudgements executed within policy bounds ÷ judgements in automated scopeDecision logWeeklyStage 5
Instrumentation build sheet for multimodal work in a retail estate. "Honest from" is the grounding stage at which the metric first measures something real rather than something available.

Two disciplines make the sheet trustworthy. First, model metrics and commerce metrics are always reported as a pair: a precision figure with no movement in returns or zero-result share is an input nobody has cashed, and a returns improvement with no precision figure cannot be defended when someone asks whether the season did it. Second, every commerce metric is read against a holdout category, because in retail promotions, weather and assortment changes will otherwise take the credit or the blame.

Grounded-stage readiness checklist

If you cannot tick all seven, you are still at the bolt-on stage regardless of how well the model reads your imagery. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Tick honestly — the blank list is data too

Most stage-2 retailers can genuinely tick one or two of these, not zero. If none apply yet, do not start with tooling: pick the category with the worst "not as described" return rate and run the 90-day plan above. Every item on this list falls out of doing that once, on one category.

Failure modes specific to multimodal retail work

Five regressions account for almost all of it, and none is caught by a model accuracy dashboard.

Multimodal grounding is not monotonic. Retailers regress, usually without noticing, because the conditions that made a gate defensible quietly stopped holding — a new photographic style, a new seller cohort, a model version rolled forward by a provider. Five failure modes account for almost all of it, and the common property is that none produces an obvious error message.

Likelihood: highImpact: high

Confident description of the wrong object

On lifestyle, flat-lay and multipack imagery the model describes the styling props, the second garment or the display packaging as if it were the product. The output is grammatical and plausible, so the usual human error reflex never fires, and the wrong attribute propagates to search, the PDP and every syndicated feed.

PreventionScore lifestyle and multi-product imagery as its own slice of the golden set, and route it to a stricter gate than clean pack shots.

Likelihood: highImpact: medium

The golden set ages out with the assortment

An evaluation set built on spring stock stops representing an autumn catalogue within two drops. Precision measured against it keeps reporting the old number while live quality falls, which is the most reassuring way to be wrong.

PreventionTie golden-set refresh to the buying calendar rather than a fixed interval, and re-score whenever photography style or supplier mix changes.

Likelihood: mediumImpact: high

A model version is deprecated on the provider's calendar

A hosted version is retired, the pipeline rolls forward to its successor, and behaviour changes mid-season on a catalogue nobody re-scored. Nothing in your code changed, so nothing in your change log explains it.

PreventionPin versions, treat deprecation notices as an operational feed, and require every model change to clear the golden set before it reaches the gate.

Likelihood: mediumImpact: high

Personal data walks in through the imagery

Customer return photographs, try-on captures and shopper UGC arrive containing faces, homes and bodies. Used for verification or matching, that processing can become biometric, and the pipeline built for product photography has none of the retention, access or deletion controls that implies.

PreventionClassify image sources at ingest, redact or crop where possible, and complete an impact assessment before any customer-supplied imagery reaches a model.

Likelihood: mediumImpact: medium

Generated imagery ships without a disclosure state

An asset created or edited by a model is published into a channel with different disclosure duties from the one it was made for — a marketplace, a paid social placement, a market with stricter labelling rules — because the disclosure state travelled in a spreadsheet rather than with the asset.

PreventionCarry disclosure and channel permissions as asset metadata in the DAM, so the publishing step can refuse rather than rely on a human remembering.

The regulatory backdrop is worth stating plainly rather than dramatically. The EU AI Act (opens in a new tab) sets transparency obligations for systems that interact with people and for synthetic content; the Digital Services Act (opens in a new tab) governs how marketplaces handle notices and explain moderation decisions; the European Data Protection Board (opens in a new tab) and national regulators such as the ICO address the personal data inside imagery; ISO's AI management standards (opens in a new tab) and the NIST AI Risk Management Framework (opens in a new tab) give you the management-system shape to hang it on. Sector context on European e-commerce practice is published by Ecommerce Europe (opens in a new tab). Confirm the current text and timetable of each with its publisher before acting — this page is not legal advice, and the applicable dates have moved more than once.

Accessibility deserves its own line because multimodal work touches it directly and usually improves it. Generated alt text is the single fastest route to closing a long-standing gap on a large catalogue, provided it is grounded: alt text describing the wrong object is worse than none, because assistive technology presents it with the same authority as a human-written description. Score alt text as an attribute class with its own floor, and read the W3C's WCAG materials (opens in a new tab) for what the text actually has to accomplish rather than treating coverage as the goal.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Multimodal model
A model that takes more than one kind of input — image, video, text, audio, structured data — into a single representation and reasons across them, rather than processing each modality separately and joining the outputs afterwards.
Vision-language model (VLM)
The subclass of multimodal model that reads images and text together. In retail it is the workhorse behind attribute extraction, alt-text generation, visual search and imagery-versus-claim comparison.
Grounding
Tying a model's proposition to independent evidence — a supplier specification, packaging text read by OCR, a category template, a verified SKU — so that a proposition becomes a claim the business can publish and defend.
Golden set
A few hundred to a few thousand SKUs whose attributes have been verified by a human against the supplier specification or the physical product. The reference against which per-attribute precision is measured and confidence floors are set.
Confidence gate
The rule that decides whether a proposed attribute value is written automatically, queued for a merchandiser, or refused outright. Set per attribute class, never as one global threshold across the catalogue.
Attribute provenance
The record written alongside every AI-populated attribute: which model, which version, what confidence, which source asset and which grounding check cleared it. The difference between explaining a bad batch and reconstructing it.
Cross-modal retrieval
Matching across modalities — a photograph against catalogue imagery, a described-not-named query against product attributes — using a shared embedding space rather than keyword overlap.
Modality dropout
The condition where an expected input is missing or unusable: a SKU with no photograph, a lifestyle shot with several products, an unreadable supplier scan. Handling it explicitly is what stops a pipeline degrading silently.
Cross-modal disagreement
A conflict between what the imagery implies and what a document or existing record states. Valuable as a signal in its own right: a rising disagreement rate usually means the assortment, the photography or the model has moved.
Null-result rate
The share of site searches returning nothing. The most direct commerce read on attribute coverage, because a shopper searching on an attribute nobody recorded gets an empty page for a product you stock.
Synthetic content disclosure
The record and the label indicating that an asset was generated or materially edited by a model, carried per asset and per publishing channel so that a market with stricter duties cannot be served an undisclosed asset by accident.
Digital twin (of a person)
An AI-generated likeness of a real person, used in place of new photography. In retail its governing questions are consent, remuneration, control over use, and disclosure — not image quality.

Frequently asked questions

The questions retail and e-commerce teams ask most often when deciding what multimodal models can safely do in their catalogue.

What is a multimodal AI model in e-commerce?

It is a model that reads more than one kind of input at once — typically product imagery together with text, and sometimes video, speech or behavioural signals — and reasons across them in a single representation. In retail that enables four things in production today: attribute extraction from imagery and supplier documents, visual and conversational search, generated or AI-edited product imagery, and triage of customer-supplied photographs in returns and moderation. What it does not do on its own is make a claim you can publish; that requires grounding the model's reading against the product record.

What is the difference between a multimodal model and a vision-language model?

A vision-language model is one type of multimodal model — the one that handles images and text. Multimodal is the wider family, which also covers audio, video and structured or behavioural inputs. In retail the distinction matters when scoping: a VLM covers catalogue enrichment, alt text and imagery-versus-claim checks, but returns triage often needs the order record too, and contact-centre work needs speech. Ask which modalities a decision genuinely requires before buying a capability that spans all of them, because each additional modality adds joins, storage and data-protection obligations.

Do we need our own multimodal model, or is a general-purpose one enough?

For most retailers a general-purpose model plus your own grounding layer is the right starting point, and it stays right longer than vendors suggest. The differentiating asset is almost never the model — it is the asset-to-GTIN join, the golden set, the supplier specifications and the behavioural signals, none of which a provider has. Retailers who train or heavily adapt their own models tend to do so at marketplace scale, where category vocabularies are enormous and seller-supplied content is the dominant input. Build the grounding layer first: it makes swapping models cheap, which is the real hedge.

How big does a golden set need to be?

Between 800 and 1,500 SKUs per category is enough to set defensible floors on around a dozen attributes, and getting there takes roughly ten working days. Size matters less than sampling: the set must reflect the long tail rather than the well-photographed hero products, because the tail is where enrichment earns its money and where the model is weakest. Verify against supplier specifications or physical samples rather than the existing catalogue values, which are frequently the thing being measured. Refresh on the buying calendar, not a fixed interval.

Which product attributes should never be auto-published from an image?

Composition, allergen and hazard information, age restrictions, medical claims, country of origin and certification claims such as organic or recycled content. These fail into labelling, consumer-protection and safety exposure rather than into a poor search facet, and no confidence score changes that calculus — the claim's authority comes from a document, not a pixel. The model still has a valuable role for these classes: flagging where the imagery appears to disagree with the recorded value, which is a genuinely useful signal and costs nothing to act on cautiously.

Does visual search actually improve conversion?

The honest answer is that it depends almost entirely on what sits behind it. A camera entry point on a catalogue with thin attributes returns thin results, and shoppers attribute the failure to the feature. Where attribute coverage is good, visual and described-not-named search mainly recovers demand that was previously lost to zero-result pages, so the metric to watch first is null-result share and search-entry conversion rather than a site-wide conversion figure. Sequence enrichment before the camera icon; the reverse order is the most common way to make visual search look like a failure.

How does the EU AI Act apply to AI-generated product imagery and shopper assistants?

The Act sets transparency obligations for AI systems that interact directly with people and for content that is artificially generated or manipulated, which is why a shopper-facing assistant and a generated product image both attract disclosure considerations. Most retail multimodal uses are not high-risk under the Act, but transparency is not conditional on risk classification. The practical implication is architectural rather than legal: carry disclosure state with the asset and the conversation, per channel, so the publishing step can enforce it. Confirm current obligations and dates against the Commission's own materials, as the timetable has been amended.

Is a customer's return photograph personal data?

Usually yes, and it often carries more than the retailer intended: faces, other people, home interiors and documents in the frame. Where imagery is used to identify or verify a person the processing may also become biometric, which raises the bar considerably. Treat customer-supplied imagery as a distinct source class from product photography, with its own lawful basis, retention limit, access control and deletion path, and complete an impact assessment before it reaches a model. Redaction or cropping at ingest removes most of the exposure at very little cost to the triage task.

How do we stop a multimodal model hallucinating product attributes?

You do not stop it; you catch it. The three controls that work are grounding checks against an independent source such as the supplier specification or packaging text read by OCR, per-attribute confidence floors derived from a golden set, and a queue for anything below the floor. Add a fourth for retail specifically: treat lifestyle, flat-lay and multipack imagery as its own slice with a stricter gate, because describing the styling prop instead of the product is the dominant failure and it produces confident, well-formed, wrong text.

What does multimodal enrichment actually do to returns?

It attacks the "not as described" reason code, which is a product-content failure before it is a logistics one. The mechanism is unglamorous: the attribute that would have set the shopper's expectation — a dimension, a material, a compatibility note — was never populated, because populating it by hand for a long-tail SKU was never worth anyone's time. Measure it properly by holding out a comparable category on supplier-supplied attributes and comparing that reason code, rather than reading a year-on-year chart in which promotions and weather have already voted.

What happens when the model version we depend on is deprecated?

Behaviour changes without a line of your code changing, which is why this is the most common regression at the automated stage. Pin the serving version explicitly rather than tracking a floating alias, subscribe to the provider's deprecation notices as an operational feed with a named owner, and require any successor version to clear the golden set before it reaches the gate. Budget for it: a model migration is a release with re-validation, not a configuration change, and scheduling one mid-peak is how a stable pipeline produces a bad season.

Who should own multimodal work — merchandising or technology?

Both, with distinct accountabilities, and the split is what makes the programme survive a budget cycle. Merchandising owns the outcome — attribute completeness, the return reason code, the search null rate — and signs off the precision floors per attribute class, because those are consequence decisions rather than modelling ones. Technology owns the pipeline's production behaviour: joins, evaluation, gate, provenance, model versioning and revert. Programmes with only a technology owner stall because nobody's targets improve when the model is used; programmes with only a merchandising owner stall because nobody is paged when it drifts.

About the author

Atomic Loops Engineering

Applied AI practice

Atomic Loops builds production AI systems for retail and e-commerce operators — catalogue enrichment, search and ranking, content generation and returns decisioning — integrated into the PIM, OMS and search stack rather than delivered as a demo console.

  • · Multimodal catalogue enrichment shipped into live PIM estates
  • · Golden-set design and per-attribute evaluation with merchandising teams
  • · Rights, disclosure and personal-data controls built into the asset pipeline
  • · 18 cited sources on this page

Sources

  1. McKinsey & CompanyThe state of AI (opens in a new tab)
  2. Baymard InstituteE-commerce search research (opens in a new tab)
  3. Baymard InstituteCart abandonment rate statistics (opens in a new tab)
  4. National Retail FederationResearch and insights (opens in a new tab)
  5. GS1GS1 standards (opens in a new tab)
  6. GS1GS1 Digital Link (opens in a new tab)
  7. Information Commissioner's OfficeGuidance on AI and data protection (opens in a new tab)
  8. EDPBEuropean Data Protection Board (opens in a new tab)
  9. European CommissionRegulatory framework for AI (AI Act) (opens in a new tab)
  10. European CommissionDigital Services Act package (opens in a new tab)
  11. NISTAI Risk Management Framework (opens in a new tab)
  12. ISOArtificial intelligence standards (opens in a new tab)
  13. W3CWeb Content Accessibility Guidelines (opens in a new tab)
  14. AmazonAmazon Lens visual search (opens in a new tab)
  15. AmazonAmazon Rufus shopping assistant (opens in a new tab)
  16. WalmartTechnology (opens in a new tab)
  17. H&M GroupNewsroom (opens in a new tab)
  18. Ecommerce EuropeEuropean e-commerce association (opens in a new tab)

Find out what your imagery is actually worth in the catalogue

We take one of your categories, build a small verified answer set, score whichever model you use per attribute, and hand back the precision floors and the cost-per-SKU maths. You keep the golden set and the numbers whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.