Retail & E-CommerceRegulations, Compliance & Governance
Governing AI decisions in retail and e-commerce: who may decide what, inside which bounds
Governing AI decisions in retail and e-commerce is the discipline of naming every automated decision the business takes about a customer, a seller or a colleague, declaring who — or what — is authorised to take it and inside which bounds, and keeping a record complete enough to explain, contest and reverse it.

Key takeaways
- The unit of governance is the decision, not the model. One fraud model can drive four decisions — score, hold, decline, restrict the account — with four different levels of consequence for the customer, and each needs its own authority, bound and record.
- Most retailers already run automated decisions they never chose: fraud rules that shipped configured inside the payment service provider, returns-abuse thresholds enabled by a vendor release, marketplace filters that suspend listings. Governance starts with an inventory, not a policy.
- A bound only counts when the system enforces it. An authority written in a governance document and a threshold set in a vendor console will diverge at the first release; the bound has to sit in the path the action travels, so an out-of-bounds action queues to a human instead of executing.
- Contestability is an engineering property. It needs four things that rarely exist together: a record complete enough to reconstruct the decision, a reviewer with authority to reverse, a reversal action that actually exists in the system, and a clock. A complaints inbox is none of these.
- Upheld contests are the most valuable governance dataset a retailer owns — they are labelled errors, produced free, by the people the decision landed on. Operators who watch the upheld share rather than the contest volume catch a bad bound weeks before it reaches a regulator.
Abbreviations used on this page
- ADM
- Automated decision-making (the UK GDPR / GDPR Article 22 term)
- OMS
- Order management system
- PIM
- Product information management system
- CDP
- Customer data platform
- PSP
- Payment service provider (card acquiring and fraud screening)
- CNP
- Card-not-present (the online transaction type most fraud models score)
- RMA
- Return merchandise authorisation
- SOR
- Statement of reasons (the DSA's notice for a moderation decision)
- DSA
- Digital Services Act (EU Regulation 2022/2065)
- DPIA
- Data protection impact assessment
- GMV
- Gross merchandise value
- AOV
- Average order value
Free · 8 questions · ~3 minutes
Score how your estate governs its automated decisions
Eight questions, one at a time, about three minutes. Answer them and we build your personalised decision-governance report — your stage on the ladder, your score on each of the four dimensions, and the specific gap between you and the next stage — and send it to your inbox. Your answers double as the first draft of your decision register.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised decision-governance report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the decision classes your answers suggest are least governed, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Undeclared
Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
Your next moveWrite the register: every automated decision, the system it executes in, who owns it, and who it lands on. Two weeks of walking the estate, not a governance programme.
Stage 2 · Registered
Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
Your next moveFor the three highest-consequence decisions, write the authority class and the numeric bound, then put the bound where it is enforced — in the path the action travels — not in the document.
Stage 3 · Bounded
Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
Your next moveAttach a statement of reasons and a contest route with a clock to every decision that touches a person's money, their account or their shift.
Stage 4 · Contestable
Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
Your next moveMake the evidence continuous: standing exports of decision volumes, bound-breach attempts, override rates and upheld-contest share, on the trading calendar.
Stage 5 · Self-evidencing
Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.
Your next moveTie bound review to the events that genuinely shift the distribution — market launches, vendor migrations, category expansions and peak — not only to a quarterly calendar.
0 / 24
Decision inventory
— / 6
Authority and bounds
— / 6
Oversight and contestability
— / 6
Decision evidence
— / 6
Your score maps to a rung on the decision-governance ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps you, and in retail estates it is almost always authority and bounds — the gap between what the register says and what the console does. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the decision-governance ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps you, and in retail estates it is almost always authority and bounds — the gap between what the register says and what the console does.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want your register and bounds reviewed against a live estate?
We walk your trading, engineering and legal leads through the dimension scores, open the consoles alongside the register to find where the two disagree, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan and the marked-up register either way.
How the score maps to a stage
- 0–5 — Stage 1, Undeclared. Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
- 6–11 — Stage 2, Registered. Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
- 12–16 — Stage 3, Bounded. Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
- 17–21 — Stage 4, Contestable. Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
- 22–24 — Stage 5, Self-evidencing. Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.
What governing AI decisions in retail actually means
A definition, the four authority classes, and the path a single customer-affecting decision travels from signal to reversal.
Governing AI decisions in retail and e-commerce means naming each automated decision the business takes about a person, declaring who or what may take it and within which limits, and keeping a record complete enough to explain, contest and reverse it. The unit of governance is the decision — an order held, a refund refused, a listing suppressed, a seller suspended, a shift reallocated, a price personalised — not the model that scores it and not the vendor that supplies it.
That distinction does most of the work on this page. A single fraud model in a retail estate typically drives at least four decisions of escalating consequence: it ranks a review queue, it holds an order, it declines a payment, and it restricts an account. Ranking a queue affects nobody outside the building. Restricting an account affects a person's ability to buy, may be hard to undo, and is the sort of decision that data protection regulators describe as having a legal or similarly significant effect — see the ICO's guidance on rights related to automated decision-making (opens in a new tab). Governing the model as a single object gives all four the same treatment; governing the decisions gives each the treatment its consequence deserves.
This page organises by decision. Its sibling in this cell organises by instrument — a regulatory toolkit keyed to the AI Act, the DSA, data protection and consumer law — and the two are complements: a toolkit tells you what a regime demands, a decision register tells you which of your decisions the demand lands on. Everything below assumes you would rather start from the second list, because it is the one your engineers can act on.
Decisions you could defend, as governance matures
The curve is not linear. An inventory alone changes nothing you could defend, which is why stages 1 and 2 stay flat; the inflection is at stage 3, when bounds move into the path the action travels, and it steepens at stage 4, when the record and the reversal exist. Most retail estates are on the flat part, holding a register that describes an estate the consoles no longer match.
Automated decisions you could defend on demand by stage
- Stage 1 · Undeclared — 21% of operators. Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
- Stage 2 · Registered — 38% of operators. Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
- Stage 3 · Bounded — 24% of operators. Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
- Stage 4 · Contestable — 13% of operators. Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
- Stage 5 · Self-evidencing — 4% of operators. Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the European Commission's DSA Transparency Database, where 41% of platform moderation decisions are recorded as fully automated.
One customer-affecting decision, from signal to reversal
The same checkout signals travelling three different governance paths. The stage is set by where the arrow ends: undeclared paths end at a person with no reason and no route, bounded paths end in a logged, reversible action inside a written limit, and contestable paths end with the person told, heard and — where they are right — made whole. Most estates have decisions in the top lane they could not name if asked.
- Data & feeds
- AI / model
- System-of-record action
- Where value leaks
- Human in the loop
The process, in words
- In the undeclared lane, checkout and account signals reach a vendor model that returns a score with no reason codes, a threshold inside the payment service provider declines the order, and the customer sees a generic error. Nothing is recorded that anyone inside the retailer could reconstruct, and no name sits against the threshold that caused it. This is where the money and the exposure leak simultaneously.
- In the bounded lane, the same signals belong to a registered decision with an owner and an authority class. The model returns reason codes, a bound check in the path caps what may execute automatically by order value, customer tier and market, the action is a hold written into the order management system rather than a silent refusal, and a decision record captures inputs, model and policy versions, the rule that fired and the timestamp. An agent can release the hold, and the override is logged with a reason.
- In the contestable lane, a versioned decision policy sets the bounds the path enforces. Anything out of bounds or high in consequence carries a statement of reasons to the person affected, a contest route with a published clock and a reviewer who can genuinely reverse. Upheld contests do not stop at the case file: they revise the bound, which is what closes the loop between what the estate decided and what it should have decided.
Step-by-step insights
- The signals are not the problem — the silence is
- Every lane starts from the same place: order, session, device, address and payment signals from the order management system, the customer data platform and the payment service provider. Retailers rarely have a data problem at this point; they have a disclosure problem. The undeclared lane is not worse at scoring, it is worse at admitting that a decision occurred. The single cheapest improvement available to most estates is replacing the generic checkout error with an honest, specific message and a route — which costs a sprint and immediately converts an invisible refusal into a decision someone can question.
- Reason codes are the difference between a score and a decision
- A bare score cannot be explained, bounded sensibly or contested, because nothing in it says which feature drove the outcome. Reason codes — even coarse ones such as address mismatch, velocity, device reputation or basket composition — turn a number into something a colleague can act on, a statement of reasons can quote and a reviewer can check. Ask for them contractually where the model belongs to a vendor: a payment or trust-and-safety provider that will not return reason codes is selling you a decision you cannot govern, and that limitation belongs in the register.
- Where the bound has to live
- A bound written in a governance document and a threshold set in a vendor console diverge at the first release, and nothing detects it. Put the check in the path the action travels: the service that writes the hold, the refusal or the suspension asks whether this action is inside the declared class before it executes, and routes it to a human queue when it is not. The check is usually a few dozen lines. Its value is that it makes the register true by construction, because an action that contradicts the register cannot execute.
- Hold, do not decline — the shape of a reversible action
- The difference between a hold and a decline is the difference between a decision you can undo and one you cannot. A held order preserves the basket, the price, the stock allocation and the customer's intent for as long as the review takes; a decline destroys all four and sends the customer to a competitor mid-session. Wherever the action can be reshaped into a reversible form — hold instead of decline, restrict instead of close, demote instead of delist, flag instead of refuse — the governance burden drops sharply, because reversibility is what makes automation defensible at volume.
- The record is written at decision time or not at all
- Decision records cannot be reconstructed later from application logs, because the inputs have moved on: the customer's address changed, the model was retrained, the rule was edited, the catalogue was reindexed. The record has to capture the inputs as they were, the model and policy versions in force, the rule that fired and the outcome, at the moment of the decision. Retailers that add this at stage 3 find the marginal cost trivial; retailers that defer it until an audit find that the six months under examination are exactly the six months they cannot rebuild.
- Upheld contests are the loop, not the complaint
- The final edge on this diagram — from reversal back to the policy — is the one most estates never build. Upheld contests are labelled errors from the tail of the distribution the model handles worst, produced at no cost by the people the decision landed on. Routed back, they lower a ceiling, exclude a segment or trigger a retrain. Filed as complaints, they teach the business nothing and the same bad bound keeps firing. The health metric is the upheld share by decision class, watched per market and per cohort, because a rise in one market is invisible in the aggregate.
The five stages of decision governance in detail
For each rung: what it looks like inside a trading operation, the signals a reviewer can check in an afternoon, the anti-pattern that traps estates there, and what leaving costs.
Each stage below describes an observable condition of the estate rather than an ambition. The hallmarks are things a reviewer can see, the diagnostic signals are checks you can run against your own consoles and logs this week, and the anti-pattern is the specific mistake most often made trying to leave that stage — in every case a mistake that feels like progress at the time.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Undeclared
21% of operators sit here
Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
Stage 1 is rarely a stage of inaction. Almost every retailer at this stage is already taking thousands of consequential automated decisions a day — they simply arrived with the software rather than through a decision to make them. The payment service provider ships fraud rules pre-configured. The returns platform ships an abuse score with a default action. The marketplace console ships listing filters that suppress and suspend. Nobody signed anything, and yet the estate refuses orders, refuses refunds, demotes listings and restricts accounts every hour of the trading day.
The tell is the checkout error. Ask what a customer sees when a card-not-present order is declined by the fraud model, and at stage 1 the honest answer is a generic message — 'something went wrong, please try another payment method' — chosen by whoever built the checkout, not by anyone who understood what the message was hiding. The decision was real, the consequence was real, and the person on the other end got no reason, no route and no record.
This stage is cheap to leave and expensive to sit in, and the expense is asymmetric. Nothing bad happens on most days. Then one decision class goes wrong at volume — a threshold change during peak, a vendor model update, a new market with address formats the model has never seen — and there is no register to consult, no owner to call, no bound that was breached and no log to reconstruct. The investigation starts by trying to work out which system made the decision, which is a week nobody has.
In practice
The decline nobody owned
A specialist apparel retailer's card-not-present decline rate drifted up by roughly two percentage points across a quarter. Trading blamed the new checkout, engineering blamed the traffic mix, and finance saw only a soft conversion number. The actual cause was a shared fraud model at the payment service provider that had been retuned in a routine release. Nobody inside the retailer had a name against that threshold, because as far as the business was concerned it was not a decision — it was a setting. It surfaced when a customer complaint about repeated refusals reached the data protection lead, who asked which system had decided, and got four different answers.
What it looks like
- No inventory — nobody can name the automated decisions in the estate
- Thresholds live in vendor consoles, changed by whoever holds the login
- Customers meet automated refusals as a generic error with no reason
- The first time a decision is examined is when a complaint escalates
Diagnostic signals you can check this week
- Ask for a list of the automated decisions in the estate. Time how long it takes and count how many people it takes
- Open the payment or marketplace console and look at who last changed a threshold, and when. If the field is blank or shows a vendor account, you are here
- Ask what the checkout says when the fraud model declines an order. 'Something went wrong' is the stage-1 signature
- Ask the contact centre what they do when a customer insists a refusal was wrong. If the answer is 'escalate by email', there is no route
Anti-pattern · Answering the inventory question with a list of models
Asked for the decision inventory, most teams produce a model inventory — the recommender, the fraud model, the demand forecast — because that is the list engineering already keeps. It is the wrong artefact and it hides the risk. One model routinely drives several decisions of very different consequence: the same fraud score can rank a review queue, hold an order, decline a payment and restrict an account, and only the last two will ever generate a complaint. Meanwhile the highest-consequence decisions in many estates involve no model at all — they are rules, and a rules engine is not on anybody's AI register. Inventory decisions, then attach the models to them.
What holds you here
There is no list, so there is nothing to govern — every governance conversation restarts from 'which decisions do we actually make?' and dies there.
Highest-leverage next move
Write the register: every automated decision, the system it executes in, who owns it, and who it lands on. Two weeks of walking the estate, not a governance programme.
Cost of leaving
- Effort
- 4–8 weeks
- Team
- One trading lead, one platform engineer at half time, plus legal or data protection for two workshops
- Risk
- Low — nothing changes in production. The only real risk is discovering decisions you did not know you were making
- To next stage
- 1–3 months
If this is you, the next step is
Two weeks walking the estate with your trading and engineering leads. You keep the register.
Stage 2
Registered
38% of operators sit here
Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
Stage 2 is where most retail and e-commerce estates sit, and it is a genuine achievement misread as a finished one. The register is real: someone walked the estate, found the decisions, wrote them down and put a name against each. It usually surfaces two or three decisions nobody knew existed — a dormant account-closure rule, an automatic seller warning, a price-match suppression — and that alone repays the fortnight it took.
What the register does not contain is authority. It says the returns-abuse model exists and that the returns operations manager owns it. It does not say whether that model may refuse a refund by itself, up to what order value, for which customer segments, in which markets, or what happens to the case that falls outside those limits. Authority is still whatever the vendor shipped, which means the estate's real behaviour is set by release notes rather than by anybody's decision.
The failure mode of stage 2 is slow and quiet: the document and the console diverge. Every vendor release, every threshold tweak made during a bad fraud week, every new market that inherits a default moves the live estate further from the written one. Because the register is reviewed on a governance calendar and the console changes on an engineering one, the gap is only ever discovered by an incident or an audit. Operators who have sat at stage 2 for two years typically have a register that is accurate about which decisions exist and wrong about nearly every threshold in it.
In practice
The register and the console disagreed
A fashion e-commerce operator's register recorded the returns-abuse model as 'flags high-risk returns for manual review' — which was true when it was written. A platform release eighteen months later had introduced an automatic refusal action, default-enabled below a modest order value, and an operations lead had switched it on during a bad refund quarter without anything to tell them a written bound existed. Nobody was reckless and no policy was broken, because no policy said anything about it. It came to light when a customer's complaint quoted a refusal reason the register said the system could not produce.
What it looks like
- A register exists with a named owner per decision
- Authority is inherited from vendor defaults rather than written down
- The register refreshes on a governance cadence, not on system change
- Nothing in the path blocks an action that exceeds what was intended
Diagnostic signals you can check this week
- Compare the register's last-updated date with the last release date of the platforms it describes
- Pick one decision and ask an engineer to show you where the threshold is actually set. Count the clicks from the register to the truth
- Ask whether any configuration change to a registered decision requires review, and by whom. 'Change control' that covers code but not thresholds is the common gap
- Ask which registered decisions can currently execute without any human involvement. If the answer needs research, authority is implicit
Anti-pattern · Turning the register into a risk-rating exercise
The instinctive next step after an inventory is to rate everything on it — high, medium, low; red, amber, green — because rating feels like governance and produces a slide the board understands. It changes nothing in the estate. A decision rated red behaves exactly as it did the day before, because a rating is not a constraint. The step that changes behaviour is unglamorous and specific: for each high-consequence decision, write the authority class and the numeric bound in the terms the system already uses — order value, customer tier, market, category, daily volume — and then move that bound to where it can be enforced. Rate afterwards if the board wants a colour.
What holds you here
Authority is implicit, so the live behaviour of the estate drifts away from the register with every vendor release and every threshold tweak — and nothing detects the drift.
Highest-leverage next move
For the three highest-consequence decisions, write the authority class and the numeric bound, then put the bound where it is enforced — in the path the action travels — not in the document.
Cost of leaving
- Effort
- 3–6 months
- Team
- A trading owner per decision class, one platform engineer, legal review on two classes
- Risk
- Medium — the first enforced bound will block actions that used to happen silently, and someone in operations will notice within a day
- To next stage
- 3–6 months
If this is you, the next step is
A two-week workshop per decision class, ending in bounds your engineers can implement.
Stage 3
Bounded
24% of operators sit here
Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
Stage 3 is the move from paper to path. The authority class — assist, bounded automatic, notified automatic, human-only — stops being a description and becomes a gate: the code that writes the hold, the refusal or the suspension first checks the bound, and an action outside it cannot execute. It queues. That single architectural change converts governance from an assertion into a property of the system, and it is the reason stage 3 survives staff turnover, vendor migrations and bad weeks in a way stage 2 never does.
The second thing that changes is what a threshold is. At stage 2 a threshold is configuration; at stage 3 it is a versioned artefact with an owner, a change history and a review — treated with the same seriousness as the model. This is where the operators who have imported software-engineering discipline pull away: they already know how to review a change, roll one back and tell you who approved it on which date, and they simply extend that machinery to decision bounds. Teams that keep bounds in a spreadsheet next to the console rediscover, expensively, why version control exists.
What stage 3 still does not have is the person on the other end. A bounded decision can be entirely defensible internally and still land on a customer as a silent refusal, on a seller as a listing that vanished, on a colleague as a shift that changed. The action was authorised, limited and logged — and the human affected by it received no reason and has no route back. That gap is what stage 4 exists to close, and it is also where most regulatory exposure in retail actually sits.
In practice
The bound that queued Black Friday
A grocery and general-merchandise operator declared a bound on automatic refund refusal: no automatic refusal above a set basket value, none for customers in the top loyalty tier, none in a market opened within the previous six months. The bound worked exactly as designed. What nobody had modelled was volume — peak-week returns pushed roughly four times the expected number of cases into the human queue, and the queue was staffed for a normal Tuesday. The lesson stuck: a bound is not only a limit on the machine, it is a capacity commitment for the people who catch what the machine may not do, and it has to be sized against peak rather than against the average week.
What it looks like
- Every high-consequence decision has a declared authority class
- Numeric bounds are enforced in the path the action travels
- Out-of-bounds cases queue to a named human rather than executing
- The bound is versioned, and changing it goes through review
Diagnostic signals you can check this week
- Ask an engineer to try to execute an out-of-bounds action in a test environment. If it succeeds, the bound is documentation
- Look at the change history on one bound. If there is no history, the bound is configuration, not policy
- Check how the human queue behind each bound was sized — against average volume or against last peak
- Ask what a customer, seller or colleague is told when a bounded decision goes against them. Silence here is the stage-3 signature
Anti-pattern · Setting the bound from the model's confidence rather than the action's consequence
It is tempting to bound automation by score — let the model act when it is more than ninety-something per cent confident. Confidence is a property of the model against its training distribution, and it is precisely the quantity that lies during a distribution shift: a new market, a new payment method, a new seller cohort. Bound by what the action does to the person instead. Releasing a hold is trivially reversible and can be automated widely; closing an account, refusing a refund or suspending a seller is not, and belongs behind a value cap, a segment exclusion and a human above the line — no matter how confident the model claims to be.
What holds you here
The decision is limited but still silent: the customer, seller or colleague it lands on is not told what happened, and there is no route back.
Highest-leverage next move
Attach a statement of reasons and a contest route with a clock to every decision that touches a person's money, their account or their shift.
Cost of leaving
- Effort
- 6–12 months
- Team
- Platform engineer, trading owner per class, contact-centre or trust-and-safety lead, legal review
- Risk
- Medium — enforcement creates queues, and queues need staffing decisions before peak rather than during it
- To next stage
- 6–12 months
If this is you, the next step is
We replay last peak's volumes against your declared bounds and size the human queue behind each.
Stage 4
Contestable
13% of operators sit here
Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
Stage 4 looks like a policy achievement and is almost entirely an engineering one. Contestability requires four things that rarely exist together: a record complete enough to reconstruct the decision months later, a reviewer with real authority to reverse it, a reversal action that actually exists in the system, and a clock that starts when the person asks. Retailers routinely have the first and the third and believe they have all four, because they have an inbox. An inbox is intake, not contestability.
The most underrated consequence of reaching stage 4 is the data. Upheld contests are labelled errors — cases where the automated decision was wrong, identified by the person best placed to know, at no acquisition cost to you. No offline evaluation produces a comparable signal, because upheld contests are drawn precisely from the tail the model handles worst. Operators who plumb the upheld set back into the bound and the model close the loop that every governance framework describes abstractly; operators who file them as complaints have a customer-service metric and no learning.
The discipline that fails first here is reviewer authority. It is easy to staff an appeals function and hard to give it a reverse button, because the original action was often executed by a pipeline that only runs forwards — a suspension writes a flag, a decline writes a status, and there is no supported path to unwrite either. Reviewers then 'recommend' reinstatement, an engineer runs a manual fix days later, and the published clock quietly becomes fiction. Build the reversal action at the same time as the action itself, or the appeals team is a complaints desk with a better name.
In practice
The reviewer who could not reverse
A marketplace operator published a seller-appeals process with a two-business-day target after suspending accounts on an automated risk score. The appeals team was staffed, trained and fast: most cases were assessed within a day. Reinstatement, however, averaged over a week, because the suspension had been executed by a pipeline that wrote a flag no interface could clear. Appeals could only raise an engineering ticket. The published clock measured the assessment; the seller experienced the reinstatement, and lost a week of trading. The fix was two days of engineering — a supported reverse action with its own log entry — and it should have shipped with the suspension.
What it looks like
- A statement of reasons travels with the action, not after a complaint
- Contest intake exists with a published route and a response clock
- The reviewer has authority to reverse and a reversal action that exists
- Upheld contests feed back into the bound, not just into the case file
Diagnostic signals you can check this week
- Ask a reviewer to reverse one live automated decision while you watch. Time it, and count the systems they touch
- Compare the published response clock with the elapsed time to the person actually being made whole
- Read the statement of reasons your system sends. If a colleague cannot tell from it what the person should do differently, it is a notice rather than a reason
- Ask what happened to last quarter's upheld contests. If the answer is 'they were resolved', the loop is open
Anti-pattern · Measuring contests by volume instead of by upheld share
Contest volume is read as a cost line, so the instinct is to drive it down — and the cheapest way to drive it down is to make the route harder to find. That is exactly backwards. A low contest volume with a high upheld share means the route is hidden and the decisions are wrong; a higher volume with a low upheld share means the route is visible and the bounds are about right. Watch the upheld share per decision class as the primary number, publish the route as prominently as you publish the refusal, and treat a sudden rise in upheld share as an operational alarm rather than a service-level statistic.
What holds you here
Evidence is produced on request rather than continuously, so every audit, regulator question and escalation becomes a project with a deadline someone else set.
Highest-leverage next move
Make the evidence continuous: standing exports of decision volumes, bound-breach attempts, override rates and upheld-contest share, on the trading calendar.
Cost of leaving
- Effort
- 12–18 months
- Team
- Platform team, trust-and-safety or contact-centre owner, legal, plus analytics on the upheld set
- Risk
- Higher — you are publishing reasons, and every published reason has to be true and consistent with the record
- To next stage
- 12–18 months
If this is you, the next step is
One decision class, from statement of reasons to reversal action to the bound revision it triggers.
Stage 5
Self-evidencing
4% of operators sit here
Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.
Stage 5 is narrower and less exciting than the word suggests, and it is emphatically not more autonomy. It is the opposite: an estate that continuously proves what its automation did. The decision record, the bound-breach log, the override and upheld-contest rates are produced as a matter of course, in a form a regulator, an auditor, a marketplace partner or a large customer can be handed. The characteristic experience of a stage-5 operator during an information request is a query rather than a programme.
The operating signal is the bound-review trigger. Override rate, upheld-contest share and attempted bound breaches are watched as leading indicators, keyed to the events that actually change the input distribution in retail: a market launch, a category expansion, a payment-provider migration, a peak trading week. When one of them moves, a bound review is raised automatically. This is where the discipline pays for itself, because a bound set under last year's conditions is not wrong on the day the conditions change — it is wrong quietly, for weeks, until someone complains loudly enough to be heard.
Stage 5 is also the stage most likely to regress, and the regression is always structural rather than negligent. A platform migration moves a decision out of the path where the bound was enforced. A newly acquired brand arrives with its own console and its own defaults. A vendor consolidates two decisions into one product feature. None of these looks like a governance event and all of them invalidate a bound. Maturity here is not the absence of regression; it is detecting it in days, from your own telemetry, before the people affected have to tell you.
In practice
The re-bound after a market launch
An operator running fraud declines under enforced bounds launched into a new European market with different address and postcode conventions. Within six weeks the upheld share of contested declines in that market had roughly tripled while the overall rate barely moved — the aggregate hid it, the per-market cut did not. Because the upheld share was a standing metric with a review trigger rather than a quarterly report, a bound review was raised and the market's automatic-decline ceiling was lowered while the model was retrained on local address data. The customers most affected were the newest, which is precisely the cohort a launch cannot afford to lose.
What it looks like
- Decision evidence is a standing export, not an audit deliverable
- Bound review is triggered by metric movement, not only by the calendar
- Any single decision can be reconstructed in minutes, not weeks
- A new decision class cannot go live without a register entry and a bound
Diagnostic signals you can check this week
- Ask for the reconstruction of one specific decision from six months ago and time the answer
- Check whether any bound has been revised in the last two quarters, and what triggered the revision
- Ask what happens in the register when a platform migration ships. If the answer is 'we would review it', the trigger is manual
- Look for a decision class that went live in the last year and check whether it has a register entry and a bound. New classes are where the discipline leaks
Anti-pattern · Treating the evidence pack as a report rather than a control
The pack gets built, circulated monthly and admired. Nothing in the business changes when a number in it moves, so within two quarters it is produced by a junior analyst and read by nobody. Evidence is only a control if it is wired to an action: a defined movement in upheld share or bound-breach attempts raises a bound review with a named owner and a due date, the same way a failed monitor raises an incident. If no threshold in the pack can force a change, you have documentation, and documentation regresses to decoration.
What holds you here
Sustaining it is a change-control problem: every migration, market launch, acquisition and peak silently invalidates a bound that was set under different conditions.
Highest-leverage next move
Tie bound review to the events that genuinely shift the distribution — market launches, vendor migrations, category expansions and peak — not only to a quarterly calendar.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team plus a standing decision forum with trading, legal, operations and trust and safety
- Risk
- Concentrated — infrequent, high consequence, and both regulatory and reputational when it lands
If this is you, the next step is
We stress-test the register entry, the bound, the record and the reversal against a real case.
Where retail and e-commerce operators actually sit
The distribution across the ladder, and why the registered-to-bounded step is the one most estates never take.
Most retail and e-commerce operators are at stage 2, registered but unbounded. They have done the genuinely hard organisational work of finding and listing the automated decisions in the estate, and they have stopped one step short of the change that would make the list true — moving authority out of the document and into the path the action travels.
Distribution of retail and e-commerce operators across the five stages
Stage 2 is both the mode and the plateau. The drop from registered to bounded is the largest single transition loss on the ladder, because it is the first step that changes what the estate does rather than what it says.
Share of operators
- 21% — 1 · Undeclared
- 38% — 2 · Registered (the plateau)
- 24% — 3 · Bounded
- 13% — 4 · Contestable
- 4% — 5 · Self-evidencing
That database is the closest thing retail has to a public census of automated decision-making at scale, and its shape is instructive. Two of the most-reported violation categories — unsafe, non-compliant or prohibited products, and consumer information infringements — are marketplace listing decisions, which means a large share of those automated calls are commercial decisions about someone's ability to sell. The database's own framing is worth reading in full alongside the Commission's Digital Services Act landing page (opens in a new tab), because the obligation it implements is precisely the one this page is about: telling the person what was decided and why.
The Digital Services Act (DSA) obliges providers of hosting services to inform their users of the content moderation decisions they take and explain the reasons behind those decisions in so-called statements of reasons.
The distribution above is illustrative and model-derived rather than measured — it synthesises the automation share visible in the Commission's database, the maturity signals in the ICO's guidance on AI and data protection (opens in a new tab), and the operational picture in NRF's retail research (opens in a new tab). Read it as the shape of the problem rather than as a market survey. What is not illustrative is the direction of the gap: automated decisions are being taken at a scale that only the largest platforms currently report on, by retailers whose registers, where they exist at all, describe intentions rather than thresholds.
The retail decision register: every automated decision, mapped
Eight decision classes that exist in almost every retail estate — the system each executes in, the bound that must be written, what the person is owed, and the KPI at risk.
The register is the centrepiece of this discipline, and it is a table of decisions rather than of systems or models. Each row names one automated decision, the system of record where it executes, who or what is authorised to take it today, the numeric bound that has to be written down, what the person on the other end is owed when it goes against them, and the trading KPI that moves when the bound is wrong. Eight classes cover most of what a retail and e-commerce estate actually decides.
| Decision | System of record | Bound to write | What the person is owed | KPI at risk |
|---|---|---|---|---|
| Checkout fraud hold or decline | OMS + PSP | Automatic decline only below a stated order value, outside top loyalty tiers, and not in markets opened within six months | A specific reason and an alternative route to complete the order | False-decline rate, checkout conversion |
| Returns-abuse scoring and refund refusal | OMS / RMA + CRM | No automatic refusal above a basket-value ceiling; no automatic account restriction under any score | A statement of reasons and a human review on request | Refund rate, CSAT, upheld-contest share |
| Account restriction or closure | CRM / identity platform | Human-only above a stated lifetime-value threshold; never solely automated where the effect is loss of access to essential goods | Notice, the reason, and a route to human intervention | Complaint volume, retention, regulatory exposure |
| Personalised pricing and promotion targeting | Pricing engine + CDP | Floor and ceiling per category; no personalised price without the disclosure the consumer rules require | Disclosure that the price was personalised by automated means | Gross margin, promotional leakage, trust |
| Search ranking and recommendations | Search / merchandising platform + PIM | Declared own-brand or sponsored boost; blast-radius cap on any single ranking change | The main parameters of the ranking, published | Conversion, GMV mix, category share |
| Marketplace listing removal or seller suspension | Seller console / trust-and-safety platform | Suspension only after a warning, above a stated repeat threshold, with a human above the line | A statement of reasons, internal complaint handling and an out-of-court route | Seller reinstatement time, GMV at risk |
| Colleague task allocation and scheduling | Workforce management system | No automated performance consequence; manager override always available and logged | An explanation and an override they can actually use | Attrition, task completion, on-shelf availability |
| Checkout credit or pay-later eligibility | PSP / credit partner | No solely automated refusal without a route to human review, whoever's model produced it | The reason, a human review, and the ability to make representations | Checkout conversion, complaint volume, AOV |
Three columns in that table do more work than the rest. 'System of record' decides your timeline, because a decision executing inside a vendor's platform travels through the vendor's change process rather than yours — which is why the checkout fraud decline, the decision most retailers take most often, is usually the least governed. 'Bound to write' has to be expressed in units the platform already understands, or it cannot be enforced. And 'what the person is owed' is where the regulatory frames land: the same row carries a data protection duty, a consumer-law disclosure or a platform-regulation notice depending on who the decision touches and how hard it is to undo.
One row deserves particular attention because of the money moving through it. NRF and Happy Returns put total US retail returns for 2024 at a projected $890 billion (opens in a new tab), with 93% of the retailers surveyed describing return fraud and other exploitative behaviour as a significant problem for their business. That combination — enormous legitimate volume and a real abuse problem inside it — is exactly the pressure that pushes returns decisions into automation faster than any other class, and it is why the returns row is where an unbounded automatic refusal does the most reputational damage per pound saved.
Two practical notes on making the register reconstructable. First, identify the thing the decision was about with a stable key rather than an internal code: a decision record naming a SKU that has since been reused describes a different product than the one that was suppressed, whereas GS1 identifiers for products and parties (opens in a new tab) survive catalogue churn and are the keys a marketplace partner or an auditor will recognise. Second, register per market, not per group: cross-border trading is the norm in European e-commerce — Ecommerce Europe (opens in a new tab) tracks the market's structure across member states — and a single group-level row hides the fact that the same decision runs under different thresholds and different obligations in each territory it touches.
Those duties attach per decision, not per company. Under the UK GDPR and the GDPR, a decision taken solely by automated means with a legal or similarly significant effect carries specific requirements — information, a route to human intervention, the ability to contest — set out in the ICO's guidance on automated decision-making and profiling (opens in a new tab) and interpreted across the EU by the European Data Protection Board (opens in a new tab). Under the EU AI Act (opens in a new tab), a handful of retail decisions fall into the high-risk tier — creditworthiness at checkout, and systems used for worker task allocation and monitoring — while customer-facing assistants and synthetic product imagery carry transparency duties instead. Under the Digital Services Act (opens in a new tab), a marketplace owes a statement of reasons and an internal complaint-handling route for a listing removal or seller suspension. And under EU consumer protection law (opens in a new tab), a price personalised by automated decision-making has to be disclosed as such before the customer buys.
The register is also where vendor reality gets recorded honestly. Where a decision executes inside a payment, identity or trust-and-safety platform you do not control, the row should say so, name what the contract obliges the vendor to provide — reason codes, decision logs, retention, notice of model changes — and record what it does not. Most retailers discover at least one row where the honest entry is 'we cannot currently reconstruct this decision and our contract does not require the vendor to help', and that entry is more useful than any policy statement, because it is a procurement action with a date.
Four authority classes and the four dimensions that set your rung
How to decide which decisions a machine may take alone — and why the answer comes from the consequence of the action, never from the confidence of the model.
Every automated decision in a retail estate belongs in one of four authority classes, and the class is set by two properties of the action rather than by any property of the model: how much it affects the person, and how hard it is to undo. Assist means the model advises and a person decides. Bounded automatic means the machine acts inside written numeric limits. Notified automatic means the machine acts, the person is told and a contest route opens. Human-only means no automated execution at any confidence level.
Decision inventory
Whether the estate can produce a list of the automated decisions it takes about customers, sellers and colleagues, with an owner against each. The binding test is not whether a register exists but whether it agrees with the consoles: a register that has not been reconciled against live platform configuration since the last vendor release is a historical document. This is also the dimension where a model inventory gets mistaken for a decision inventory, which understates the estate by roughly the number of rules-driven decisions in it.
Authority and bounds
Whether it is written down which decisions a machine may take alone and within what numeric limits, and whether the system enforces those limits. This is overwhelmingly the lowest-scoring dimension in retail estates, and it is the one that separates stage 2 from stage 3. The tell is simple: try to execute an out-of-bounds action in a test environment. If it succeeds, the bound is a sentence in a document rather than a property of the system.
Oversight and contestability
Whether the person a decision lands on is told, and whether a specific named human can reverse it within a stated time. The ICO's checklist for Article 22 processing (opens in a new tab) is a useful external benchmark here — it asks for a simple way for people to request reconsideration and for identified staff authorised to carry out reviews and change decisions. 'Authorised to change decisions' is the phrase most retail appeals functions cannot honestly claim, because the reversal action does not exist in the system.
Decision evidence
Whether a single decision can be reconstructed months later with its inputs, model version, policy version, rule and outcome — and whether that evidence is produced continuously or assembled on request. Reconstruction time is the most testable governance property there is, and the one every audit reduces to. The NIST AI Risk Management Framework (opens in a new tab) and its companion playbook (opens in a new tab) are the most practical published scaffolding for the measure-and-manage side of this dimension.
The four dimensions gate each other, which is why the lowest is the real rung. Excellent evidence on an unregistered decision is a log nobody will think to read. A published contest route on an unbounded decision generates contests faster than the bound generates correct outcomes. And a beautifully written authority class means nothing if the estate contains three decision classes that never made it into the register — which is the normal condition after an acquisition, a replatform or a market launch.
Which authority class does this decision belong in?
Plot the action, not the model. The vertical axis is how much the decision affects the person; the horizontal axis is how easily the action can be undone. Confidence scores do not appear on this matrix, deliberately — a model's confidence is least reliable exactly when the distribution has shifted, which is when the bound matters most.
Human-only
- Account closure, seller termination, credit refusal at checkout
- No automated execution at any confidence level
- The model ranks the queue; a named person decides and signs
Notified automatic
- Order holds, refund refusals, listing removals, shift changes
- Act automatically, but tell the person and open the contest route
- Reversal action and clock ship with the decision, not after it
Bounded automatic with a ceiling
- Price changes published to the site, markdown execution, promo suppression
- Cap the blast radius: value, category, share of catalogue, per-hour volume
- A wrong price is commercial, but it is also public and hard to unsay
Free automatic, monitored in aggregate
- Ranking, recommendations, size and fit suggestions, replenishment
- Governed by aggregate monitoring rather than per-decision review
- Publish the main ranking parameters; watch drift and category share
The bottom-left quadrant is the one retailers get wrong most often, because a price change reads as commercial and therefore low-risk. It is commercial and it is also published: a mis-bounded automatic markdown appears on the site, in the feed, in comparison engines and in customers' screenshots within minutes, and cannot be unsaid even after it is corrected. Bound it by blast radius rather than by consequence to the individual — a share-of-catalogue cap and a per-hour volume limit will save more margin than any confidence threshold.
What governed automated decisions look like in public
Two publicly reported programmes, read against the ladder. Neither is an Atomic Loops engagement — each links to the operator's own published material.
The most useful public evidence for this discipline comes from operators large enough that their automated decisions had to be described in writing. In both cases below the interesting artefact is not the model — it is what the operator published about how the decision is made, who can challenge it and what happens next, because that publication is what makes the automation defensible at the volumes involved.
Two programmes read against the decision-governance ladder
Outcomes as described in the operators' own published material; verify figures against the linked source before reusing them, and note that we have not independently audited them. The images are generated industry scenes from our library, not operator photography, and imply no endorsement.
AmazonGlobal marketplace and retailer · millions of third-party sellers34
- Challenge
- Enforcement decisions at marketplace scale — blocking suspect listings, removing counterfeit products and restricting selling accounts — cannot be taken by people at the volume required, but each one removes a seller's ability to trade and therefore has to be explicable and reversible.
- Approach
- Amazon publishes how its brand protection and enforcement machinery works through its brand services programme and its policy newsroom: automated screening applied before listings publish, brand-owner reporting routes, and defined channels through which sellers and rights owners can challenge an enforcement decision and have an account or listing reinstated.
- Reported outcome
- As reported by Amazon in its own brand protection publications, automated systems screen listings at scale before publication, and the company maintains published appeal and reinstatement routes for sellers subject to enforcement action.
- What it shows about the curveAt marketplace volume, contestability is not a nicety bolted on for regulators — it is the only mechanism that keeps large-scale automated enforcement usable, because it converts the false positives that volume guarantees into cases that can be corrected rather than into permanently lost sellers.
Amazon — Brand Services and policy newsroom (opens in a new tab)
The Home DepotHome improvement retailer · 2,000+ stores, 400k+ associates23
- Challenge
- Directing store associates to the right task at the right time — restocking, picking online orders, resolving on-shelf gaps — means an algorithm making decisions that shape a colleague's working day, which is a materially different governance problem from ranking products.
- Approach
- The Home Depot has publicly described building associate-facing tools that surface prioritised tasks on handheld devices, keeping the store leader's judgement in the loop rather than issuing instructions that cannot be varied — an assist-class decision with a human decider, in the terms of this page.
- Reported outcome
- As described in the company's own newsroom material, associate-facing task and product tools are deployed across its store estate to help colleagues find stock and prioritise work, with store teams retaining the decision on how work is sequenced.
- What it shows about the curveColleague-facing decisions belong in the assist class by default, and the override is the governance artefact: a task recommendation a manager can vary and that logs the variation stays governable, while the same recommendation wired to a performance consequence becomes a high-risk decision under EU rules on worker management.
Read together, the two cases bracket the problem. One is enforcement at a scale where automation is unavoidable and contestability is the control; the other is a workplace decision at a scale where automation is optional and the override is the control. Neither operator solved governance by writing a policy — both solved it by shipping a specific mechanism attached to a specific decision, which is the only form of this work that survives contact with a trading week. Both also chose to publish, which is worth noticing on its own: Amazon's policy newsroom (opens in a new tab) and The Home Depot's corporate newsroom are where the descriptions above come from, and a decision an operator is willing to describe in public is usually one it has already had to bound internally.