Retail & E-CommerceAI-Driven Disruptions & Innovations
Multimodal AI models in retail and e-commerce: from pixels to published product claims
Multimodal AI models are systems that read images, video, text, speech and behavioural signals together rather than one modality at a time. In retail and e-commerce they enrich catalogues, power visual and conversational search, and triage photo-evidenced returns — but they only pay when every claim they make is grounded against the product record.

Key takeaways
- A multimodal model's commerce value is set by grounding, not fluency. A model that describes a product image beautifully has produced marketing copy; a model whose reading of that image can be reconciled against the GTIN-level attribute record and the supplier specification has produced a product claim you can publish.
- The join comes before the model. Until an image, a video, a review and a session can be resolved to one product identity, there is nothing to train against, nothing to score against and nowhere to write back to — and every vendor accuracy claim is measured on their assortment, not yours.
- Precision floors differ by attribute class. Colour and style can auto-publish at a high-but-reachable precision; dimensional and compatibility attributes need a source citation; composition, allergen, hazard and origin claims should never be published from imagery alone at any confidence.
- Most retailers stall at the bolt-on stage because there is no golden set. Without a human-verified answer set at SKU level you cannot justify a confidence threshold, so a merchandiser re-checks everything, the cost per enriched SKU never falls, and the programme is quietly cancelled at budget.
- Rights and disclosure are the binding constraint on generative multimodal content, not model capability. Image licensing, consent for a person's likeness, personal data arriving inside customer photos, and the EU AI Act's transparency duties decide what can ship long before accuracy does.
Abbreviations used on this page
- VLM
- Vision-language model — a model that reads images and text in one representation
- PIM
- Product information management system — the record of what a product is
- DAM
- Digital asset management system — where product imagery and video live
- PDP
- Product detail page
- OMS
- Order management system
- GTIN
- Global Trade Item Number — the GS1 product identifier
- GPC
- GS1 Global Product Classification — the shared category and attribute taxonomy
- OCR
- Optical character recognition — reading text off packaging and labels
- ASR
- Automatic speech recognition — transcribing voice and call audio
- UGC
- User-generated content — shopper photos, videos and reviews
- ANN
- Approximate nearest neighbour — the index behind embedding search
- CVR
- Conversion rate
Free · 8 questions · ~3 minutes
Score your multimodal grounding
Eight questions, one at a time, about three minutes. Answer them and we build your personalised grounding report — your stage on the ladder, your score on each of the four dimensions, and the specific blocker standing between you and the next stage — and send it to your inbox. Your result doubles as the baseline for the enrichment business case.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised grounding report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, how you compare with retailers of similar assortment size and channel mix, and the 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Single-modality
Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
Your next moveBuild the asset-to-GTIN join and a rights register for one category, then a small human-verified answer set for the attributes that actually matter in it.
Stage 2 · Bolt-on
Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
Your next moveBuild a golden set on one category, score the model per attribute rather than overall, and set precision floors by attribute class before writing anything.
Stage 3 · Grounded
Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
Your next moveServe one shared product representation — a single embedding index and a single attribute record with provenance — to every surface that needs it.
Stage 4 · Shared representation
Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
Your next moveEnumerate the multimodal decisions allowed to execute unattended, each with a value ceiling, a precision floor and a reconstructable trail.
Stage 5 · Bounded autonomy
Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.
Your next moveVersion the multimodal decision policy, pin model versions, and treat a deprecation as a planned operational event with a re-validation gate.
0 / 24
Modality joins & identity
— / 6
Grounding & evaluation
— / 6
Write path & surfaces
— / 6
Rights, transparency & safeguards
— / 6
Your score maps to a stage on the grounding ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps you, and in retail it is very often rights and safeguards rather than model quality. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the grounding ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps you, and in retail it is very often rights and safeguards rather than model quality.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this run against your actual catalogue?
We take one category, build a small verified answer set, score whichever model you are using per attribute, and hand back the precision floors and the cost-per-SKU maths. You keep the golden set and the numbers whether or not we build anything.
How the score maps to a stage
- 0–5 — Stage 1, Single-modality. Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
- 6–11 — Stage 2, Bolt-on. Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
- 12–16 — Stage 3, Grounded. Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
- 17–21 — Stage 4, Shared representation. Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
- 22–24 — Stage 5, Bounded autonomy. Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.
What multimodal AI models are — and what changes in a catalogue
A definition, the separation between what ships today and what is still research, and the path a pixel has to travel before it becomes a product claim you can publish.
Multimodal AI models are models that take more than one kind of input into a single representation — an image and a sentence, a video and a transcript, a photograph and a structured product record — and reason across them rather than about each separately. In retail and e-commerce that matters for a mundane reason: almost everything a shopper wants to know about a product is in the photograph, and almost nothing about the photograph is in the database.
The commercially useful subset is narrower than the phrase suggests. Four things are in production across the sector today: attribute extraction and catalogue enrichment from imagery and supplier documents; visual and conversational search, where a photograph or a described-not-named query is matched against the assortment; generated and AI-edited product and marketing imagery; and triage of customer-supplied photographs in returns, claims and marketplace moderation. Everything else — from fully agentic buying to models that understand a garment's drape from a still — is at various distances from a shipping system, and this page keeps that separation explicit.
The organising idea underneath all four is grounding. A multimodal model produces a proposition: this garment is ribbed; this cable fits that socket; this returned item is damaged. A retailer cannot publish a proposition — it publishes a claim, which is a proposition plus the evidence that it holds. The distance between the two is the whole engineering problem, and it is why two retailers using the same model get results a factor apart.
Value released against position on the grounding ladder
The curve is not linear. Value stays close to flat through stages 1 and 2 — where most retailers are — and inflects at stage 3, when output starts being written into the product record under a defensible gate rather than reviewed by a person. This is why programmes that measure progress in images processed report activity without result.
Commercial value released by stage
- Stage 1 · Single-modality — 21% of operators. Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
- Stage 2 · Bolt-on — 38% of operators. Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
- Stage 3 · Grounded — 26% of operators. Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
- Stage 4 · Shared representation — 12% of operators. Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
- Stage 5 · Bounded autonomy — 3% of operators. Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with McKinsey's State of AI adoption research.
How a pixel becomes a published product claim
The path from a product photograph to a live attribute, stage by stage. The stage is set by where the arrow ends: at stages 1–2 it ends at a person retyping, at stages 3–4 it ends in the PIM behind a per-attribute gate with provenance, and at stage 5 a bounded set publishes unattended. Most retailers are in the top lane.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- At stages 1–2 a vision-language model reads a hand-picked set of images and produces tags into a spreadsheet or a staging queue. Nothing carries a GTIN or a provenance record, so a merchandiser retypes whatever they accept and the cost per enriched SKU does not move. This is where the value leaks, and it leaks quietly because the output reads well.
- At stages 3–4 assets, copy, reviews and behaviour are joined on the product identifier. The model proposes a value and a confidence per attribute; a maintained golden set says what that confidence is worth for that attribute; a gate set per attribute class passes the confident values straight into the PIM with full provenance and queues the rest for a merchandiser to adjudicate.
- At stage 5 a versioned policy lets an enumerated set of judgements publish or settle unattended — cosmetic attributes below a value threshold, photo-evidenced claims under a ceiling. Anything outside the bounds escalates to a person, and every automated judgement can be reconstructed months later with its model version, its source pixel and the policy version in force.
Step-by-step insights
- Images in the DAM — the asset that cannot be found
- The single most predictive artefact of a stalled multimodal programme is a DAM organised by campaign. Assets filed as "AW26_Shoot3_Final_v2_USE_THIS.jpg" cannot be joined to a product, which means they cannot be training data, cannot be scored, and cannot be written back from. Retailers routinely discover that the join is not one problem but three: assets without identifiers, identifiers without assets, and the far more awkward middle case of one lifestyle image carrying four sellable products and belonging, arguably, to all of them. Resolve that before any model conversation; it is weeks of unglamorous work that every later stage depends on.
- The demo trap — fluency is not evidence
- A vision-language model reading forty curated photographs is a persuasive artefact and a near-useless one. The curated set has clean backgrounds, single products, consistent lighting and categories the model has seen a great deal of. Your long tail has none of that. Worse, generated attribute text is grammatical and confident whether it is right or wrong, so the usual human error-detection reflex — that something reads oddly — never fires. This is the specific reason content AI needs harder evidence than forecasting AI, not softer: a wrong number looks wrong, a wrong sentence does not.
- The join — identity before intelligence
- Joining assets to the GTIN is what turns a pile of media into a dataset. Once an image resolves to a product, everything else follows cheaply: the reviews attached to that product become supervision, the returns reason codes become labels, the clickstream becomes a relevance signal, and the supplier specification becomes the grounding evidence. GS1 identifiers are the natural spine because they already reach outside the business to suppliers and marketplaces, which is where most of your attribute disagreements will come from.
- The golden set — the artefact that unlocks everything downstream
- A golden set is a few hundred to a few thousand SKUs whose attributes have been verified by a human against the physical product or the supplier specification. It is boring to build and it is the highest-leverage week in the entire programme, because it converts every subsequent argument from opinion to measurement: which model, which prompt, which threshold, which attributes can be automated and which cannot. Build it per category, refresh it as the assortment turns, and treat any model or prompt change as a release that has to clear it.
- The gate — one threshold per attribute class, never one global number
- A single global confidence threshold is the most common design error at stage 3 and it fails in both directions at once: too low for composition and allergen claims, too high for colour and pattern, so you are simultaneously over-exposed and under-automating. Gate per attribute class, with a floor derived from the golden set and a consequence analysis behind it. Some classes should have no threshold at all that permits automatic publication, and saying so in the policy is a stronger control than a very high number.
- Provenance and the bounded policy — what an auditor actually reads
- Provenance costs almost nothing at write time — model identifier, version, confidence, source asset, clearing check — and is the difference between explaining a bad batch in an afternoon and reconstructing it for a fortnight. At stage 5 the same record becomes external evidence: a marketplace asking why a listing was suppressed, or a consumer body asking how a claim was substantiated, is asking to read the policy version and the decision trail, not the model. Version the policy, pin the model, and log the judgement.
The five stages of multimodal grounding
For each stage: what it looks like inside a real catalogue team, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps retailers there, and what leaving costs.
Each stage below describes how far a multimodal model's output has travelled toward being a claim the business will stand behind. The hallmarks are observable conditions in the PIM, the DAM and the search stack; the diagnostic signals are checks you can run against your own systems this week; and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Single-modality
21% of operators sit here
Single-modality is the stage where images, video and speech are assets the business stores rather than signals any system reads — all commerce intelligence runs on text and numbers.
Stage 1 is not a shortage of pixels. A mid-sized retailer already holds tens of thousands of product images, thousands of hours of video, years of call recordings and a decade of shopper photographs attached to reviews. What is missing is any path from those assets to a decision: the images are addressed by filename, the transcripts sit in a contact-centre archive nobody queries, and the product record is a separate universe maintained by hand.
The clearest tell is the search box. Type a description rather than a name — "black ribbed midi with a side slit", "the cable that fits the older model" — and a lexical index returns either nothing or the wrong department, because the words a shopper uses are not the words in the title and the information that would answer them is sitting in the photograph. Retailers at this stage usually know their search is weak; what they have not yet connected is that the fix is a data-joining problem before it is a model problem.
This stage is cheap to leave and expensive to sit in, because the cost is invisible. Nobody books "the enrichment we could not do" or "the long tail that never got attributes", so the deficit shows up somewhere else: agency invoices for product copy, a returns rate the category manager blames on suppliers, and a search team that keeps buying synonym packs.
In practice
The forty-SKU demo
A homewares retailer sat through a vendor demo in which a vision-language model read forty hand-picked product photographs and produced clean, plausible attributes for each. Everyone in the room was impressed. Nobody could answer the two questions that mattered: what it would do on the 12,000 SKUs in the long tail, because there was no verified answer set to compare against; and how a value would ever reach the PIM, because no image in the DAM could be resolved to a GTIN without a person opening the file.
What it looks like
- Product imagery lives in the DAM organised by campaign and filename, not by GTIN
- Site search is lexical: a shopper's photo, or a described-not-named query, returns nothing useful
- Catalogue attributes are typed by merchandisers or accepted from supplier feeds unchecked
- Any multimodal work is a vendor demo on a hand-picked set of SKUs
Diagnostic signals you can check this week
- Ask what share of live SKUs carry every attribute their category template requires — the honest number is usually far below the reported one
- Search your own site for a described-not-named product and count how many of the first twenty results are usable
- Pick a product photograph at random and ask whether anything but a human can tell you which GTIN it belongs to
- Ask who holds the rights record for that photograph. If the answer is "the agency, probably", generative work cannot start here
Anti-pattern · Buying the model before building the join
The instinct is to procure a multimodal enrichment platform, because the demo is compelling and the pricing is per SKU. It fails for a structural reason: the vendor's accuracy is measured on their evaluation set and your assortment is not in it, and even a perfect output has nowhere to land while assets and product records share no identifier. Build the asset-to-GTIN join and a small verified answer set on one category first. That work is unglamorous, takes weeks rather than quarters, and it is what turns every subsequent vendor conversation from a demo into a measurement.
What holds you here
Images, video, reviews and behaviour are not joined to a single product identity, so nothing can be trained, scored or written back.
Highest-leverage next move
Build the asset-to-GTIN join and a rights register for one category, then a small human-verified answer set for the attributes that actually matter in it.
Cost of leaving
- Effort
- 2–4 months
- Team
- One data engineer, one merchandising lead, part-time
- Risk
- Low — nothing customer-facing changes and no product record is written
- To next stage
- 2–4 months
If this is you, the next step is
A two-week engagement: map assets, product records and rights, then size the join.
Stage 2
Bolt-on
38% of operators sit here
Bolt-on is the stage where a vision-language model reads product imagery in one bounded task, but its output stops at a person, a spreadsheet or a staging queue — never the product record.
Stage 2 is where most retailers are, and it is the most misleading stage on the ladder because everything about it looks like progress. A model is live. It produces output daily. The output is visibly better than nothing. And the cost per enriched SKU has not moved at all, because a human still reads every line before it is trusted — which was the expensive part in the first place.
The structural cause is the absence of an answer set. Without a human-verified reference for what the attributes on a given SKU actually are, there is no way to say that the model is right 96% of the time on colour and 71% of the time on material. And without that split, no threshold can be defended: a merchandising director asked to let a machine write into the catalogue at 90% confidence will reasonably ask what 90% means in returns, and at stage 2 nobody can answer.
Time spent here is not neutral. Merchandisers learn that the model produces work for them rather than removing it, and the next proposal is heard against that memory. The pattern is close to the dashboard trap in other AI programmes, with one twist specific to content: because generated copy is pleasant to read, quality problems hide better here than in a forecast. A wrong number looks wrong; a wrong sentence reads fine.
In practice
The alt-text backlog that came back
A fashion retailer generated alt text for 60,000 product images to close an accessibility gap ahead of an audit. The copy was fluent and the backlog cleared in a fortnight. Three months later a manual sample found a recurring failure: on flat-lay and lifestyle shots the model frequently described the styling props — the chair, the plant, the second model's jacket — as if they were the product. The alt text was accessible, well written and, for a meaningful slice of the catalogue, about the wrong object. There had never been a verified reference to score it against, so nothing had flagged it.
What it looks like
- A VLM generates tags, alt text or draft descriptions into a staging queue
- Every output is human-reviewed, because none of it is scored
- Nothing is written to the PIM under a rule — a merchandiser retypes what they accept
- Quality is discussed in impressions ("it's pretty good") rather than per-attribute precision
Diagnostic signals you can check this week
- Ask for per-attribute precision on the last thousand generated values. A single overall figure, or an anecdote, means you are here
- Count how many generated values reached the PIM without a human retyping them. If it is zero, the model has added review work, not removed it
- Compute cost per enriched SKU including review time, then compare it with the agency or offshore quote it was meant to replace
- Ask which model version produced last quarter's descriptions. If nobody knows, there is no provenance and no way to explain a bad batch
Anti-pattern · Prompt-engineering past a measurement problem
When output is uneven the reflex is to rewrite the prompt, and it works often enough to be seductive: the failures you happened to look at go away. What has actually happened is that the prompt has been fitted to a handful of remembered examples, with no way to know what it broke elsewhere in the assortment. Every prompt change is a model change and needs the same evidence as one. Build the golden set first; then prompt, fine-tune or swap models freely, because you can see the effect per attribute rather than per anecdote.
What holds you here
There is no human-verified answer set, so no confidence threshold can be justified and every output must be re-checked by a person.
Highest-leverage next move
Build a golden set on one category, score the model per attribute rather than overall, and set precision floors by attribute class before writing anything.
Cost of leaving
- Effort
- 3–6 months
- Team
- One ML engineer, one merchandising QA lead, a named category owner
- Risk
- Medium — the first automated write into the PIM needs a fallback to the supplier-supplied value
- To next stage
- 3–6 months
If this is you, the next step is
Ten working days to a 1,000-SKU verified answer set and defensible per-attribute floors.
Stage 3
Grounded
26% of operators sit here
Grounded is the stage where multimodal output is written into the product record under a per-attribute confidence gate, with provenance, a review queue for the low-confidence tail and measurement in commerce units.
Stage 3 is the first stage at which multimodal work changes the unit economics of the catalogue. The mechanism is narrow and specific: for each attribute the model proposes a value with a confidence, the value is checked against whatever independent evidence exists — the supplier specification sheet, the GS1 category template, the packaging text read by OCR — and only values above that attribute class's floor are written. Everything else queues. The merchandiser's day changes from typing to adjudicating, and adjudication is a far smaller job.
The discipline that makes it work is provenance. Every written attribute carries the model identifier, the version, the confidence, the source asset and the check that cleared it. That record costs almost nothing to produce at write time and is the difference between a supportable catalogue and an unexplainable one. When a category manager asks why 400 SKUs suddenly claim a fabric they do not have, provenance turns a forensic exercise into a filtered query.
The constraint that emerges at stage 3 is duplication. Search has built its own embedding, the content team has its own captioner, the support assistant has its own retrieval index, and the returns team is evaluating a third vendor. Each works. None compounds, and worse, the same product now means slightly different things in four places — which is how a shopper reads one thing on the PDP, is told another by the assistant and receives a third in the box.
In practice
The category that stopped being retyped
A homewares team started with one category — occasional furniture — and the twelve attributes that drove its return reasons: dimensions, assembly requirement, material family, weight class and finish. They verified 1,100 SKUs by hand, scored the model per attribute, and found precision ranged from the high nineties on finish and colour to barely two-thirds on material. They gated accordingly: finish and colour auto-published, dimensions published only where OCR of the supplier sheet agreed, and material always queued. Roughly seven in ten attribute values cleared the gate. The queue that remained was a morning's work a week rather than a permanent headcount.
What it looks like
- Attributes are written to the PIM with model, version, confidence and source asset recorded
- The confidence gate is set per attribute class — regulated attributes never auto-pass
- Merchandisers work a queue of the low-confidence tail instead of re-checking everything
- Value is reported as attribute completeness, null-result rate and return-reason movement, not model accuracy
Diagnostic signals you can check this week
- Open any AI-populated attribute in the PIM and ask which model version and which source asset produced it — the answer should be one click, not an investigation
- Check whether the gate uses one global threshold or different floors for colour and for allergen. One global threshold means the classes have not been thought about
- Ask what happens on modality dropout: a SKU with no photograph, or a lifestyle shot with three products in it
- Ask whether search, the PDP and the contact-centre assistant read the same attribute record. Three answers means three truths
Anti-pattern · One model per surface
Each team ships the multimodal capability it needs, and each is individually justified: search wants an embedding tuned for retrieval, content wants a captioner tuned for tone, support wants an assistant tuned for policy. The cost lands later and lands on the customer, as inconsistency — the assistant confidently describing a colourway the PDP does not list. Consolidating after four surfaces exist is a migration; agreeing one shared product representation while there are two is an afternoon. The signal to act is the second surface, not the fourth.
What holds you here
Every surface has its own model and its own embedding, so the same product means different things in search, on the PDP and in support — and nothing compounds.
Highest-leverage next move
Serve one shared product representation — a single embedding index and a single attribute record with provenance — to every surface that needs it.
Cost of leaving
- Effort
- 6–12 months
- Team
- ML engineer, catalogue/PIM engineer, merchandising owner, plus a compliance partner for the regulated attribute classes
- Risk
- Medium — the risk shifts from accuracy to consistency across surfaces
- To next stage
- 6–12 months
If this is you, the next step is
We map what each surface reads today and what a shared representation would have to guarantee.
Stage 4
Shared representation
12% of operators sit here
Shared representation is the stage where one multimodal product record and one session representation serve every surface, so a new surface is configuration rather than a project.
At stage 4 the marginal cost of a new multimodal surface collapses, and the conversation changes shape with it. Teams stop proposing "an AI project for returns" and start asking which decisions should read the product representation next. That is a healthier conversation because it is about decisions and metrics rather than about models, and it is the reliable signature of this stage.
The measurement discipline is what separates stage 4 from a well-engineered stage 3. Each surface has a named commerce metric and, where the surface allows it, a holdout: a category left on the previous ranking, a market left on supplier-supplied attributes, a returns queue left on manual triage. Holdouts are harder in retail than in logistics because seasonality and promotions move everything at once, which is exactly why the comparison group has to exist rather than be argued from a year-on-year chart.
The governance work also consolidates here, and this is the part most programmes underestimate. Rights and consent for imagery, disclosure of synthetic content, personal data arriving inside customer photographs, and accessibility of generated alt text are all properties of the representation layer, not of the individual surface. Handled once, they are a register and a set of checks. Handled per team, they are four incompatible answers to the same regulator's question.
In practice
Returns triage in three weeks
An electricals retailer already served one product representation to search, the PDP and its assistant. When the returns team asked for automated triage on photo-evidenced claims — is the item damaged, is it the item that was ordered, is the accessory in the box — the build was three weeks, most of it spent agreeing the metric and the holdout with operations rather than writing code. The representation, the rights handling and the review queue already existed. That ratio, specification-heavy and build-light, is the stage-4 tell.
What it looks like
- One embedding index and one attribute record serve search, recommendations, PDP content, support and returns
- A new surface ships in weeks because the representation already exists
- Every surface is evaluated against a holdout in its own commerce metric
- Disclosure, rights and personal-data controls are handled once in the representation layer, not per team
Diagnostic signals you can check this week
- Measure elapsed time from "we want visual search in the app" to it serving traffic, for the last two surfaces
- Change an attribute definition in one place and check whether every surface picks it up, or whether three teams have to be told
- Ask whether any multimodal surface has a live holdout rather than a year-on-year comparison
- Ask whether the rights and disclosure register covers generated assets on every channel, including marketplaces and paid social
Anti-pattern · Automating because the representation is good
A gate working well on colour and finish creates pressure to extend it to compatibility, composition and origin, because the machinery is identical and the volumes are tempting. The classes are not comparable. Cosmetic attributes fail into a poor facet; composition and origin claims fail into consumer-protection and labelling exposure, and a single wrong allergen inference is a different category of event from a thousand wrong colour tags. Each attribute class earns automation from its own evidence, with its own floor and its own sign-off.
What holds you here
The model still only reads and proposes; the remaining value needs it to publish, suppress or settle — which needs bounded policy and audit evidence, not more accuracy.
Highest-leverage next move
Enumerate the multimodal decisions allowed to execute unattended, each with a value ceiling, a precision floor and a reconstructable trail.
Cost of leaving
- Effort
- 12–24 months
- Team
- Platform team, merchandising product owner, legal and compliance partner
- Risk
- Higher — a representation error propagates to every surface at once
- To next stage
- 12–24 months
If this is you, the next step is
Which multimodal judgements may publish or settle unattended, and the evidence that makes it defensible.
Stage 5
Bounded autonomy
3% of operators sit here
Bounded autonomy is the stage where an enumerated set of multimodal judgements publishes, suppresses or settles without human approval inside stated bounds, with people owning exceptions and policy.
Stage 5 is much narrower than the word autonomy suggests. It is not a self-running catalogue; it is a short, enumerated list of judgements that may execute unattended inside stated bounds, with everything outside those bounds escalating. Colour and finish on a low-value SKU qualify. A photo-evidenced damage claim under a value ceiling qualifies. Allergen inference, origin claims and anything with a safety or labelling consequence are correctly held at stage 4 permanently, and a mature operator says so in the policy rather than leaving it to a threshold.
The engineering is largely finished by the time an operator arrives here. The hard artefact is the evidence: showing a marketplace, a regulator or a consumer body why one specific listing was suppressed nine months ago, which model version and which pixel produced the judgement, and under which version of the policy. Treat the policy with the same rigour as the model — it is the document that will be read out, and it is far more likely to be examined than the weights.
Sustaining this stage is a change-control problem rather than a technical one, and it is the stage most likely to regress. Assortments turn over seasonally, so an evaluation set stops representing the catalogue within a couple of drops. Photography style changes when the agency does. And model providers deprecate versions on their own calendar, which means a capability you validated in March can behave differently in September without a single line of your code changing.
In practice
The bounded publish set
A marketplace operator publishes AI-derived colour, pattern and silhouette attributes unattended for listings below a stated value, and auto-settles photo-evidenced damage claims below a ceiling set per category. Roughly one judgement in twelve escalates. The escalation rate is itself monitored: a rise means the assortment or the imagery has moved outside the policy's validity — a new seller cohort, a new photographic style — and it triggers a re-validation against a refreshed golden set before it triggers a complaint.
What it looks like
- Cosmetic and dimensional attributes below a value threshold publish unattended
- Photo-evidenced returns under a value ceiling settle automatically; everything else routes to a person
- Listings whose imagery contradicts their claims are suppressed pending review, with a logged reason and an appeal path
- The multimodal policy is versioned and reviewed like code, and model versions are pinned rather than floating
Diagnostic signals you can check this week
- Check whether the policy names attribute classes and value ceilings explicitly, with a version history and an owner
- Ask when the kill switch — revert every gated attribute to the supplier-supplied value — was last exercised deliberately
- Ask whether escalation and appeal rates are watched as leading indicators rather than reported after an incident
- Pick one suppressed listing from nine months ago and ask for the reason, the model version and the policy version
Anti-pattern · Letting the vendor's release calendar set your policy
A hosted model version is deprecated, the pipeline silently rolls forward to its successor, and behaviour changes mid-season on a catalogue nobody re-scored. It is the most common regression at this stage and it is entirely avoidable: pin versions, subscribe to deprecation notices as an operational feed, and treat every model change as a release that must clear the golden set before it touches the gate. A model swap is a deployment, not a configuration tweak.
What holds you here
Sustaining autonomy is a governance and change-control problem: model versions are deprecated on the provider's calendar, assortments turn over seasonally and disclosure duties move.
Highest-leverage next move
Version the multimodal decision policy, pin model versions, and treat a deprecation as a planned operational event with a re-validation gate.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team plus a standing content-governance forum with legal and merchandising
- Risk
- Concentrated — low frequency, high consequence, consumer-law and platform-policy in nature
If this is you, the next step is
We stress-test the policy, the provenance trail and the revert against a real scenario.
Where retailers actually sit on the ladder
The distribution across the five stages, and why the bolt-on stage is both the mode and the trap.
Most retailers are at the bolt-on stage. The distribution is heavily weighted toward work where a multimodal model is live and producing output that a human still reads before anything is trusted — which is to say, live but not yet load-bearing. A minority have a defensible gate into the product record, and a very small group let any multimodal judgement stand without a person behind it.
Distribution of retailers across the five grounding stages
Illustrative distribution — a model-derived synthesis, not a survey. Stage 2 is the mode and the plateau: the drop from bolt-on to grounded is the largest single transition loss on the ladder, and it is an evidence problem rather than a model problem.
Share of retailers
- 21% — 1 · Single-modality
- 38% — 2 · Bolt-on (the plateau)
- 26% — 3 · Grounded
- 12% — 4 · Shared representation
- 3% — 5 · Bounded autonomy
Source: Illustrative distribution, synthesised from McKinsey, NRF and Baymard research
The shape is not specific to multimodal work — cross-industry research has consistently found a wide gap between organisations experimenting with AI and organisations reporting material bottom-line impact, which is the pattern McKinsey's State of AI (opens in a new tab) has tracked over successive years. What is specific to retail is the shape of the blocker. In logistics the usual gap is a write-back path into an execution system; in a catalogue the write path is comparatively easy and the missing piece is evidence: nobody can say what the model's confidence is worth on this assortment, so nobody will let it write.
That figure is worth sitting with, because it reframes catalogue enrichment from a content-operations cost line into a margin lever. When a shopper returns an item as "not as described", the description is usually not false so much as absent — the attribute that would have set the expectation was never populated, because populating it by hand for a long-tail SKU has never been worth anyone's time. This is the specific gap a grounded multimodal pipeline closes, and it is why the business case belongs to the category manager rather than to the innovation team. Baymard Institute's e-commerce search research (opens in a new tab) makes the same point from the discovery side: shoppers routinely fail to find products that exist in the catalogue, because the attributes they search on were never recorded.
Where multimodal models land in a retail operation
Eight decisions worth wiring, the modalities each one actually needs, the system of record it writes to, the KPI it moves and the stage at which it earns its keep.
Multimodal value in retail concentrates in eight decisions, and each has a natural home on the ladder. A decision is a good first candidate when three things hold: the system of record is one you control, the evidence to ground the model already exists somewhere in the business, and the KPI it moves is one a category manager already reports. The map below is how we scope first and second multimodal use cases with retailers.
| Decision | Modalities in play | System of record | KPI it moves | Earns its keep |
|---|---|---|---|---|
| Catalogue enrichment and attribute extraction | Product image + supplier spec sheet (OCR) + existing copy | PIM / DAM | Attribute completeness, cost per enriched SKU | Stage 3 |
| Visual search and camera entry points | Shopper photograph + catalogue imagery + clickstream | Search index | Null-result rate, search-entry CVR | Stage 3 |
| Conversational and described-not-named search | Query text + attribute record + review text | Search index / merchandising rules | Zero-result queries, add-to-basket from search | Stage 3–4 |
| PDP content and accessible alt text | Product image + attributes + brand tone rules | CMS / PIM | Accessibility conformance, organic entrances | Stage 3 |
| Generated and AI-edited product imagery | Source photography + generated variants + rights record | DAM | Content cost per SKU, disclosure coverage | Stage 3–4 |
| Returns and claims triage on photo evidence | Customer photograph + order record + policy text | OMS / returns platform | Processing cost per return, claims fraud loss | Stage 4 |
| Marketplace and third-party listing integrity | Listing imagery + title and description + category taxonomy | Marketplace / seller platform | Policy-violating listings caught, notice turnaround | Stage 4 |
| Shelf, pick and pack verification | Camera stream + planogram + inventory record | Store systems / WMS | On-shelf availability, mispick rate | Stage 4 |
Catalogue enrichment is where most retailers should begin, for reasons that have nothing to do with it being the most exciting. The system of record is yours, the grounding evidence — supplier specifications, existing verified SKUs, returns reason codes — is already in the building, and the KPI moves inside a quarter. Visual search is the more visible project and the harder first one, because its quality is bounded by exactly the attribute coverage that enrichment produces: a camera entry point on a thin catalogue returns thin results, and the shopper blames the feature.
The compliance column is not a separate domain here — it runs through every row. Product identity and classification lean on GS1 standards (opens in a new tab), and the GTIN is the join key that makes attribute disagreements with suppliers and marketplaces tractable at all; GS1 Digital Link (opens in a new tab) extends the same identifier into the web resources a model can retrieve. Shopper-facing assistants and synthetic content attract transparency duties under the EU AI Act (opens in a new tab), marketplace moderation decisions attract notice-and-action and statement-of-reasons duties under the Digital Services Act (opens in a new tab), and any imagery containing a person pulls in data-protection obligations set out in the ICO's guidance on AI and data protection (opens in a new tab). None of these prohibits multimodal work. All of them assume you can explain a decision afterwards, which is a grounding property.
Why the bolt-on stage stalls: the grounding problem
Three structural patterns account for most of the plateau, none of them a modelling problem — plus the precision floors that decide which attributes may ever be automated.
The bolt-on stage stalls because the pilot was scoped to prove that the model can read an image, and reading the image was never the constraint. Modern vision-language models read product photography well. The constraint is that a retailer cannot act on a proposition it cannot check, and at stage 2 there is nothing to check against — no verified answer set, no supplier specification joined to the SKU, no record of which model version said what.
There is no verified answer set, so no threshold can be defended
Confidence scores are meaningless until they are calibrated against ground truth on your own assortment. Without that, "publish above 0.9" is a number somebody chose in a meeting, and the merchandising director who declines to sign it off is being rigorous rather than obstructive. The golden set is the cheapest artefact on this page and the one whose absence stops everything.
The pilot was scoped to a category and the business case needs the tail
Enrichment pilots are run on a category with good imagery and clean supplier data, because that is where a pilot succeeds. But the value of enrichment is concentrated precisely where the data is worst — the long tail nobody has had time to populate by hand. A pilot on the best-documented 5% of the catalogue proves the model works and proves nothing about the economics.
Content quality is discussed in taste, not in precision
Because generated copy is a matter of tone as well as fact, review meetings drift toward whether the descriptions sound right. Tone is real and it is a brand decision. It is also a different conversation from whether the fibre composition is correct, and merging the two is how factual error survives review. Separate them: tone is signed off once, per template; facts are measured per attribute, continuously.
The remedy is a per-attribute view of consequence. Not every wrong value costs the same, and treating them uniformly is what produces both over-exposure and under-automation at once. The table below is the framework we use with merchandising and compliance teams to decide, class by class, what may ever be automated and what must remain a human or supplier-supplied claim regardless of how confident a model becomes.
| Attribute class | Examples | What a wrong value costs | Floor before any auto-publish | Who signs the class off |
|---|---|---|---|---|
| Cosmetic / descriptive | Colour family, pattern, silhouette, neckline, finish | A poor facet and a slightly worse search result | High but reachable — derive from the golden set, then sample continuously | Merchandising, with sampled QA |
| Dimensional | Length, capacity, drop, weight class, pack quantity | "Not as described" returns and support contacts | Higher, and only where OCR of the supplier sheet agrees with the image | Category manager |
| Compatibility | Fits model X, socket type, thread size, cartridge series | Returns plus a support contact plus a negative review | Highest, and only with a citation to the supplier specification | Technical merchant or supplier quality |
| Composition / material | Fibre content, food ingredients, packaging material | Labelling and consumer-law exposure; sustainability claims | Never from imagery alone — document-grounded only, model may propose | Supplier data plus compliance |
| Regulated / safety | Allergens, hazard pictograms, age restriction, medical claims | Legal exposure, recall, delisting | Never auto-published; the model's only role is flagging disagreement | Compliance and legal |
| Provenance claims | Country of origin, "organic", "recycled content", certifications | Greenwashing and consumer-protection exposure | Never inferred — the claim comes from the certificate, not the pixel | Legal and sustainability |
Building the golden set — ten working days
Pick the category by consequence, not by convenience
Choose the category where missing attributes are already costing you — highest "not as described" return rate, or highest zero-result search volume. The point is not to make the model look good; it is to produce a number a category manager will act on.
Choose the twelve attributes that actually matter
Read the returns reason codes and the last quarter's search queries for that category. Twelve attributes is usually enough to cover the reasons shoppers return and fail to find, and a set small enough that human verification finishes inside a fortnight.
Verify 800–1,500 SKUs by hand against evidence, not memory
Verification means against the supplier specification, the physical sample or the packaging — not against the existing PIM value, which is frequently the thing being wrong. Sample across the tail deliberately: the well-photographed hero SKUs are not the population you need to measure.
Score per attribute and publish the confusion, not just the score
A single accuracy figure hides everything useful. Report precision and recall for each attribute, and keep the disagreements — the systematic ones (styling props read as products, lifestyle shots, multipacks) tell you what to fix in the pipeline rather than in the model.
Set the floors with merchandising and compliance in the room
Floors are a business decision informed by measurement, not a modelling output. Agree them per attribute class, write down the two or three classes that will never auto-publish, and version the document — it becomes the multimodal decision policy later.
Schedule the refresh against the assortment cycle
A golden set built in spring stops representing an autumn assortment. Tie its refresh to the buying calendar rather than to a fixed interval, and re-run it whenever the model version, the prompt or the photography style changes.
The AI system's outputs must be sufficiently accurate for the purpose, and organisations must be able to explain the reasoning behind decisions to the people affected by them.
What multimodal deployment looks like in public
Three publicly reported programmes, read against the grounding ladder. None is an Atomic Loops engagement — each links to the operator's own published material.
The most instructive public examples are not the ones with the largest models but the ones where the operator built the surrounding apparatus at the same time — the catalogue, the identity, the rights record. In each case below the differentiator is what sits behind the modality, and the lesson maps onto a specific rung of the ladder.
Three programmes read against the grounding ladder
Outcomes as reported by the operators themselves; verify figures against the linked source before reusing them, and note that dated press releases for two of these programmes no longer resolve, so the stable corporate newsroom is linked instead. Card images are generated industry scenes from our asset library, not operator photography, and imply no endorsement.
AmazonGlobal marketplace and retailer · catalogue at internet scale35
- Challenge
- Shoppers arrive with a photograph or a vague description far more often than with a product name, and a lexical index answers neither. At marketplace scale the problem compounds: much of the catalogue is seller-supplied, so imagery and attributes vary in quality across hundreds of millions of listings.
- Approach
- Amazon has publicly described two complementary multimodal entry points in its shopping app: Lens, which matches a photograph taken by the shopper against the catalogue, and Rufus, a generative shopping assistant that answers product questions grounded in catalogue data, listings and community content rather than from the model's own recall.
- Reported outcome
- Amazon reports both features as generally available to customers in its shopping app, and describes Rufus explicitly as answering from product information, listings and customer reviews — that is, as a grounded assistant rather than a free-standing chatbot.
- What it shows about the curveThe modality is the entry point; the catalogue is the product. Visual and conversational search are only as good as the attribute coverage and review corpus behind them, which is why enrichment usually has to precede the camera icon rather than follow it.
WalmartGlobal omnichannel retailer · stores plus marketplace34
- Challenge
- Search and product content across an enormous, largely supplier-fed assortment, spanning grocery, general merchandise and a third-party marketplace — categories whose attribute vocabularies have almost nothing in common with each other.
- Approach
- Walmart has publicly described building retail-specific language models trained on its own catalogue and customer data, and applying generative AI to search and shopping assistance across its properties, rather than relying solely on general-purpose models with no view of its assortment.
- Reported outcome
- Walmart publishes ongoing reporting on its AI-powered search and assistant work through its corporate technology channel; treat the specifics there as the authority rather than any secondary summary, including this one.
- What it shows about the curveThe defensible asset is the representation — catalogue, behaviour, identity — not the model. The stage-4 signature is exactly this: one representation good enough that new surfaces read from it instead of building their own.
H&M GroupGlobal fashion retailer · high assortment turnover23
- Challenge
- Producing enough on-model imagery for a fast-turning assortment across many markets and channels, where photography cost and lead time constrain how much of the catalogue can be shown well at all.
- Approach
- H&M Group publicly reported exploring AI-generated digital twins of models — likenesses created with the models' agreement, with the models retaining rights over the use of their digital twin and being remunerated for it, and the resulting imagery labelled as AI-generated where used.
- Reported outcome
- The programme was reported through H&M Group's own newsroom, which remains the authority on its current scope; the durable point is the shape of the controls rather than any volume figure.
- What it shows about the curveFor generative multimodal content, the deployment work is the governance work. Consent, rights and disclosure are what decide whether an asset can ship — capability stopped being the constraint some time ago.
Read together, the three make one argument. Amazon's assistant is described as answering from listings and reviews — grounding, stated as a product property. Walmart's investment is in a retail-specific representation rather than in a bigger general model. H&M Group's reported programme is notable less for the imagery than for the consent, rights and labelling apparatus around it, which is the part a European retailer has to build regardless of which model it picks. See About Amazon's own description of Rufus (opens in a new tab) for the grounding claim in the operator's words.