Redefining Technology

Retail & E-CommerceRegulations, Compliance & Governance

Governing AI decisions in retail and e-commerce: who may decide what, inside which bounds

Governing AI decisions in retail and e-commerce is the discipline of naming every automated decision the business takes about a customer, a seller or a colleague, declaring who — or what — is authorised to take it and inside which bounds, and keeping a record complete enough to explain, contest and reverse it.

Illustrative retail scene: a store colleague reviewing automated pricing and merchandising decisions on a tablet, with decision overlays across the shop floor
Retail & E-Commerce · Regulations, Compliance & Governance

Key takeaways

  1. The unit of governance is the decision, not the model. One fraud model can drive four decisions — score, hold, decline, restrict the account — with four different levels of consequence for the customer, and each needs its own authority, bound and record.
  2. Most retailers already run automated decisions they never chose: fraud rules that shipped configured inside the payment service provider, returns-abuse thresholds enabled by a vendor release, marketplace filters that suspend listings. Governance starts with an inventory, not a policy.
  3. A bound only counts when the system enforces it. An authority written in a governance document and a threshold set in a vendor console will diverge at the first release; the bound has to sit in the path the action travels, so an out-of-bounds action queues to a human instead of executing.
  4. Contestability is an engineering property. It needs four things that rarely exist together: a record complete enough to reconstruct the decision, a reviewer with authority to reverse, a reversal action that actually exists in the system, and a clock. A complaints inbox is none of these.
  5. Upheld contests are the most valuable governance dataset a retailer owns — they are labelled errors, produced free, by the people the decision landed on. Operators who watch the upheld share rather than the contest volume catch a bad bound weeks before it reaches a regulator.

Abbreviations used on this page

ADM
Automated decision-making (the UK GDPR / GDPR Article 22 term)
OMS
Order management system
PIM
Product information management system
CDP
Customer data platform
PSP
Payment service provider (card acquiring and fraud screening)
CNP
Card-not-present (the online transaction type most fraud models score)
RMA
Return merchandise authorisation
SOR
Statement of reasons (the DSA's notice for a moderation decision)
DSA
Digital Services Act (EU Regulation 2022/2065)
DPIA
Data protection impact assessment
GMV
Gross merchandise value
AOV
Average order value

Free · 8 questions · ~3 minutes

Score how your estate governs its automated decisions

Eight questions, one at a time, about three minutes. Answer them and we build your personalised decision-governance report — your stage on the ladder, your score on each of the four dimensions, and the specific gap between you and the next stage — and send it to your inbox. Your answers double as the first draft of your decision register.

0 of 8 answered

Question 1 of 8Decision inventory

Could you produce this week a list of every automated decision your estate takes about a customer, a seller or a colleague?

Governance has no unit of work until the decisions are named. Estates consistently contain two or three consequential decisions nobody knew were live.

How the score maps to a stage
  • 05 — Stage 1, Undeclared. Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
  • 611 — Stage 2, Registered. Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
  • 1216 — Stage 3, Bounded. Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
  • 1721 — Stage 4, Contestable. Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
  • 2224 — Stage 5, Self-evidencing. Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.

What governing AI decisions in retail actually means

A definition, the four authority classes, and the path a single customer-affecting decision travels from signal to reversal.

Governing AI decisions in retail and e-commerce means naming each automated decision the business takes about a person, declaring who or what may take it and within which limits, and keeping a record complete enough to explain, contest and reverse it. The unit of governance is the decision — an order held, a refund refused, a listing suppressed, a seller suspended, a shift reallocated, a price personalised — not the model that scores it and not the vendor that supplies it.

That distinction does most of the work on this page. A single fraud model in a retail estate typically drives at least four decisions of escalating consequence: it ranks a review queue, it holds an order, it declines a payment, and it restricts an account. Ranking a queue affects nobody outside the building. Restricting an account affects a person's ability to buy, may be hard to undo, and is the sort of decision that data protection regulators describe as having a legal or similarly significant effect — see the ICO's guidance on rights related to automated decision-making (opens in a new tab). Governing the model as a single object gives all four the same treatment; governing the decisions gives each the treatment its consequence deserves.

This page organises by decision. Its sibling in this cell organises by instrument — a regulatory toolkit keyed to the AI Act, the DSA, data protection and consumer law — and the two are complements: a toolkit tells you what a regime demands, a decision register tells you which of your decisions the demand lands on. Everything below assumes you would rather start from the second list, because it is the one your engineers can act on.

Decisions you could defend, as governance matures

The curve is not linear. An inventory alone changes nothing you could defend, which is why stages 1 and 2 stay flat; the inflection is at stage 3, when bounds move into the path the action travels, and it steepens at stage 4, when the record and the reversal exist. Most retail estates are on the flat part, holding a register that describes an estate the consoles no longer match.

Automated decisions you could defend on demand by stage

  • Stage 1 · Undeclared — 21% of operators. Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.
  • Stage 2 · Registered — 38% of operators. Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.
  • Stage 3 · Bounded — 24% of operators. Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.
  • Stage 4 · Contestable — 13% of operators. Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.
  • Stage 5 · Self-evidencing — 4% of operators. Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the European Commission's DSA Transparency Database, where 41% of platform moderation decisions are recorded as fully automated.

One customer-affecting decision, from signal to reversal

The same checkout signals travelling three different governance paths. The stage is set by where the arrow ends: undeclared paths end at a person with no reason and no route, bounded paths end in a logged, reversible action inside a written limit, and contestable paths end with the person told, heard and — where they are right — made whole. Most estates have decisions in the top lane they could not name if asked.

  • Data & feeds
  • AI / model
  • System-of-record action
  • Where value leaks
  • Human in the loop

The process, in words

  • In the undeclared lane, checkout and account signals reach a vendor model that returns a score with no reason codes, a threshold inside the payment service provider declines the order, and the customer sees a generic error. Nothing is recorded that anyone inside the retailer could reconstruct, and no name sits against the threshold that caused it. This is where the money and the exposure leak simultaneously.
  • In the bounded lane, the same signals belong to a registered decision with an owner and an authority class. The model returns reason codes, a bound check in the path caps what may execute automatically by order value, customer tier and market, the action is a hold written into the order management system rather than a silent refusal, and a decision record captures inputs, model and policy versions, the rule that fired and the timestamp. An agent can release the hold, and the override is logged with a reason.
  • In the contestable lane, a versioned decision policy sets the bounds the path enforces. Anything out of bounds or high in consequence carries a statement of reasons to the person affected, a contest route with a published clock and a reviewer who can genuinely reverse. Upheld contests do not stop at the case file: they revise the bound, which is what closes the loop between what the estate decided and what it should have decided.
Step-by-step insights
The signals are not the problem — the silence is
Every lane starts from the same place: order, session, device, address and payment signals from the order management system, the customer data platform and the payment service provider. Retailers rarely have a data problem at this point; they have a disclosure problem. The undeclared lane is not worse at scoring, it is worse at admitting that a decision occurred. The single cheapest improvement available to most estates is replacing the generic checkout error with an honest, specific message and a route — which costs a sprint and immediately converts an invisible refusal into a decision someone can question.
Reason codes are the difference between a score and a decision
A bare score cannot be explained, bounded sensibly or contested, because nothing in it says which feature drove the outcome. Reason codes — even coarse ones such as address mismatch, velocity, device reputation or basket composition — turn a number into something a colleague can act on, a statement of reasons can quote and a reviewer can check. Ask for them contractually where the model belongs to a vendor: a payment or trust-and-safety provider that will not return reason codes is selling you a decision you cannot govern, and that limitation belongs in the register.
Where the bound has to live
A bound written in a governance document and a threshold set in a vendor console diverge at the first release, and nothing detects it. Put the check in the path the action travels: the service that writes the hold, the refusal or the suspension asks whether this action is inside the declared class before it executes, and routes it to a human queue when it is not. The check is usually a few dozen lines. Its value is that it makes the register true by construction, because an action that contradicts the register cannot execute.
Hold, do not decline — the shape of a reversible action
The difference between a hold and a decline is the difference between a decision you can undo and one you cannot. A held order preserves the basket, the price, the stock allocation and the customer's intent for as long as the review takes; a decline destroys all four and sends the customer to a competitor mid-session. Wherever the action can be reshaped into a reversible form — hold instead of decline, restrict instead of close, demote instead of delist, flag instead of refuse — the governance burden drops sharply, because reversibility is what makes automation defensible at volume.
The record is written at decision time or not at all
Decision records cannot be reconstructed later from application logs, because the inputs have moved on: the customer's address changed, the model was retrained, the rule was edited, the catalogue was reindexed. The record has to capture the inputs as they were, the model and policy versions in force, the rule that fired and the outcome, at the moment of the decision. Retailers that add this at stage 3 find the marginal cost trivial; retailers that defer it until an audit find that the six months under examination are exactly the six months they cannot rebuild.
Upheld contests are the loop, not the complaint
The final edge on this diagram — from reversal back to the policy — is the one most estates never build. Upheld contests are labelled errors from the tail of the distribution the model handles worst, produced at no cost by the people the decision landed on. Routed back, they lower a ceiling, exclude a segment or trigger a retrain. Filed as complaints, they teach the business nothing and the same bad bound keeps firing. The health metric is the upheld share by decision class, watched per market and per cohort, because a rise in one market is invisible in the aggregate.

The five stages of decision governance in detail

For each rung: what it looks like inside a trading operation, the signals a reviewer can check in an afternoon, the anti-pattern that traps estates there, and what leaving costs.

Each stage below describes an observable condition of the estate rather than an ambition. The hallmarks are things a reviewer can see, the diagnostic signals are checks you can run against your own consoles and logs this week, and the anti-pattern is the specific mistake most often made trying to leave that stage — in every case a mistake that feels like progress at the time.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Undeclared

21% of operators sit here

Undeclared is the stage at which the business cannot produce a list of the automated decisions it already takes about customers, sellers and colleagues.

Stage 1 is rarely a stage of inaction. Almost every retailer at this stage is already taking thousands of consequential automated decisions a day — they simply arrived with the software rather than through a decision to make them. The payment service provider ships fraud rules pre-configured. The returns platform ships an abuse score with a default action. The marketplace console ships listing filters that suppress and suspend. Nobody signed anything, and yet the estate refuses orders, refuses refunds, demotes listings and restricts accounts every hour of the trading day.

The tell is the checkout error. Ask what a customer sees when a card-not-present order is declined by the fraud model, and at stage 1 the honest answer is a generic message — 'something went wrong, please try another payment method' — chosen by whoever built the checkout, not by anyone who understood what the message was hiding. The decision was real, the consequence was real, and the person on the other end got no reason, no route and no record.

This stage is cheap to leave and expensive to sit in, and the expense is asymmetric. Nothing bad happens on most days. Then one decision class goes wrong at volume — a threshold change during peak, a vendor model update, a new market with address formats the model has never seen — and there is no register to consult, no owner to call, no bound that was breached and no log to reconstruct. The investigation starts by trying to work out which system made the decision, which is a week nobody has.

In practice

The decline nobody owned

A specialist apparel retailer's card-not-present decline rate drifted up by roughly two percentage points across a quarter. Trading blamed the new checkout, engineering blamed the traffic mix, and finance saw only a soft conversion number. The actual cause was a shared fraud model at the payment service provider that had been retuned in a routine release. Nobody inside the retailer had a name against that threshold, because as far as the business was concerned it was not a decision — it was a setting. It surfaced when a customer complaint about repeated refusals reached the data protection lead, who asked which system had decided, and got four different answers.

What it looks like

  • No inventory — nobody can name the automated decisions in the estate
  • Thresholds live in vendor consoles, changed by whoever holds the login
  • Customers meet automated refusals as a generic error with no reason
  • The first time a decision is examined is when a complaint escalates

Diagnostic signals you can check this week

  • Ask for a list of the automated decisions in the estate. Time how long it takes and count how many people it takes
  • Open the payment or marketplace console and look at who last changed a threshold, and when. If the field is blank or shows a vendor account, you are here
  • Ask what the checkout says when the fraud model declines an order. 'Something went wrong' is the stage-1 signature
  • Ask the contact centre what they do when a customer insists a refusal was wrong. If the answer is 'escalate by email', there is no route

Anti-pattern · Answering the inventory question with a list of models

Asked for the decision inventory, most teams produce a model inventory — the recommender, the fraud model, the demand forecast — because that is the list engineering already keeps. It is the wrong artefact and it hides the risk. One model routinely drives several decisions of very different consequence: the same fraud score can rank a review queue, hold an order, decline a payment and restrict an account, and only the last two will ever generate a complaint. Meanwhile the highest-consequence decisions in many estates involve no model at all — they are rules, and a rules engine is not on anybody's AI register. Inventory decisions, then attach the models to them.

What holds you here

There is no list, so there is nothing to govern — every governance conversation restarts from 'which decisions do we actually make?' and dies there.

Highest-leverage next move

Write the register: every automated decision, the system it executes in, who owns it, and who it lands on. Two weeks of walking the estate, not a governance programme.

Cost of leaving

Effort
4–8 weeks
Team
One trading lead, one platform engineer at half time, plus legal or data protection for two workshops
Risk
Low — nothing changes in production. The only real risk is discovering decisions you did not know you were making
To next stage
1–3 months

If this is you, the next step is

Two weeks walking the estate with your trading and engineering leads. You keep the register.

Run a decision discovery

Stage 2

Registered

38% of operators sit here

Registered is the stage at which a decision inventory exists and is owned, but nothing states which decisions a model may take alone or within what limits.

Stage 2 is where most retail and e-commerce estates sit, and it is a genuine achievement misread as a finished one. The register is real: someone walked the estate, found the decisions, wrote them down and put a name against each. It usually surfaces two or three decisions nobody knew existed — a dormant account-closure rule, an automatic seller warning, a price-match suppression — and that alone repays the fortnight it took.

What the register does not contain is authority. It says the returns-abuse model exists and that the returns operations manager owns it. It does not say whether that model may refuse a refund by itself, up to what order value, for which customer segments, in which markets, or what happens to the case that falls outside those limits. Authority is still whatever the vendor shipped, which means the estate's real behaviour is set by release notes rather than by anybody's decision.

The failure mode of stage 2 is slow and quiet: the document and the console diverge. Every vendor release, every threshold tweak made during a bad fraud week, every new market that inherits a default moves the live estate further from the written one. Because the register is reviewed on a governance calendar and the console changes on an engineering one, the gap is only ever discovered by an incident or an audit. Operators who have sat at stage 2 for two years typically have a register that is accurate about which decisions exist and wrong about nearly every threshold in it.

In practice

The register and the console disagreed

A fashion e-commerce operator's register recorded the returns-abuse model as 'flags high-risk returns for manual review' — which was true when it was written. A platform release eighteen months later had introduced an automatic refusal action, default-enabled below a modest order value, and an operations lead had switched it on during a bad refund quarter without anything to tell them a written bound existed. Nobody was reckless and no policy was broken, because no policy said anything about it. It came to light when a customer's complaint quoted a refusal reason the register said the system could not produce.

What it looks like

  • A register exists with a named owner per decision
  • Authority is inherited from vendor defaults rather than written down
  • The register refreshes on a governance cadence, not on system change
  • Nothing in the path blocks an action that exceeds what was intended

Diagnostic signals you can check this week

  • Compare the register's last-updated date with the last release date of the platforms it describes
  • Pick one decision and ask an engineer to show you where the threshold is actually set. Count the clicks from the register to the truth
  • Ask whether any configuration change to a registered decision requires review, and by whom. 'Change control' that covers code but not thresholds is the common gap
  • Ask which registered decisions can currently execute without any human involvement. If the answer needs research, authority is implicit

Anti-pattern · Turning the register into a risk-rating exercise

The instinctive next step after an inventory is to rate everything on it — high, medium, low; red, amber, green — because rating feels like governance and produces a slide the board understands. It changes nothing in the estate. A decision rated red behaves exactly as it did the day before, because a rating is not a constraint. The step that changes behaviour is unglamorous and specific: for each high-consequence decision, write the authority class and the numeric bound in the terms the system already uses — order value, customer tier, market, category, daily volume — and then move that bound to where it can be enforced. Rate afterwards if the board wants a colour.

What holds you here

Authority is implicit, so the live behaviour of the estate drifts away from the register with every vendor release and every threshold tweak — and nothing detects the drift.

Highest-leverage next move

For the three highest-consequence decisions, write the authority class and the numeric bound, then put the bound where it is enforced — in the path the action travels — not in the document.

Cost of leaving

Effort
3–6 months
Team
A trading owner per decision class, one platform engineer, legal review on two classes
Risk
Medium — the first enforced bound will block actions that used to happen silently, and someone in operations will notice within a day
To next stage
3–6 months

If this is you, the next step is

A two-week workshop per decision class, ending in bounds your engineers can implement.

Write bounds for your top three decisions

Stage 3

Bounded

24% of operators sit here

Bounded is the stage at which each registered decision carries a written authority class and numeric limits, and the limits are enforced by the system rather than by the document.

Stage 3 is the move from paper to path. The authority class — assist, bounded automatic, notified automatic, human-only — stops being a description and becomes a gate: the code that writes the hold, the refusal or the suspension first checks the bound, and an action outside it cannot execute. It queues. That single architectural change converts governance from an assertion into a property of the system, and it is the reason stage 3 survives staff turnover, vendor migrations and bad weeks in a way stage 2 never does.

The second thing that changes is what a threshold is. At stage 2 a threshold is configuration; at stage 3 it is a versioned artefact with an owner, a change history and a review — treated with the same seriousness as the model. This is where the operators who have imported software-engineering discipline pull away: they already know how to review a change, roll one back and tell you who approved it on which date, and they simply extend that machinery to decision bounds. Teams that keep bounds in a spreadsheet next to the console rediscover, expensively, why version control exists.

What stage 3 still does not have is the person on the other end. A bounded decision can be entirely defensible internally and still land on a customer as a silent refusal, on a seller as a listing that vanished, on a colleague as a shift that changed. The action was authorised, limited and logged — and the human affected by it received no reason and has no route back. That gap is what stage 4 exists to close, and it is also where most regulatory exposure in retail actually sits.

In practice

The bound that queued Black Friday

A grocery and general-merchandise operator declared a bound on automatic refund refusal: no automatic refusal above a set basket value, none for customers in the top loyalty tier, none in a market opened within the previous six months. The bound worked exactly as designed. What nobody had modelled was volume — peak-week returns pushed roughly four times the expected number of cases into the human queue, and the queue was staffed for a normal Tuesday. The lesson stuck: a bound is not only a limit on the machine, it is a capacity commitment for the people who catch what the machine may not do, and it has to be sized against peak rather than against the average week.

What it looks like

  • Every high-consequence decision has a declared authority class
  • Numeric bounds are enforced in the path the action travels
  • Out-of-bounds cases queue to a named human rather than executing
  • The bound is versioned, and changing it goes through review

Diagnostic signals you can check this week

  • Ask an engineer to try to execute an out-of-bounds action in a test environment. If it succeeds, the bound is documentation
  • Look at the change history on one bound. If there is no history, the bound is configuration, not policy
  • Check how the human queue behind each bound was sized — against average volume or against last peak
  • Ask what a customer, seller or colleague is told when a bounded decision goes against them. Silence here is the stage-3 signature

Anti-pattern · Setting the bound from the model's confidence rather than the action's consequence

It is tempting to bound automation by score — let the model act when it is more than ninety-something per cent confident. Confidence is a property of the model against its training distribution, and it is precisely the quantity that lies during a distribution shift: a new market, a new payment method, a new seller cohort. Bound by what the action does to the person instead. Releasing a hold is trivially reversible and can be automated widely; closing an account, refusing a refund or suspending a seller is not, and belongs behind a value cap, a segment exclusion and a human above the line — no matter how confident the model claims to be.

What holds you here

The decision is limited but still silent: the customer, seller or colleague it lands on is not told what happened, and there is no route back.

Highest-leverage next move

Attach a statement of reasons and a contest route with a clock to every decision that touches a person's money, their account or their shift.

Cost of leaving

Effort
6–12 months
Team
Platform engineer, trading owner per class, contact-centre or trust-and-safety lead, legal review
Risk
Medium — enforcement creates queues, and queues need staffing decisions before peak rather than during it
To next stage
6–12 months

If this is you, the next step is

We replay last peak's volumes against your declared bounds and size the human queue behind each.

Stress-test your bounds against peak

Stage 4

Contestable

13% of operators sit here

Contestable is the stage at which every consequential automated decision is recorded, explained to the person it affects, and reversible on a clock.

Stage 4 looks like a policy achievement and is almost entirely an engineering one. Contestability requires four things that rarely exist together: a record complete enough to reconstruct the decision months later, a reviewer with real authority to reverse it, a reversal action that actually exists in the system, and a clock that starts when the person asks. Retailers routinely have the first and the third and believe they have all four, because they have an inbox. An inbox is intake, not contestability.

The most underrated consequence of reaching stage 4 is the data. Upheld contests are labelled errors — cases where the automated decision was wrong, identified by the person best placed to know, at no acquisition cost to you. No offline evaluation produces a comparable signal, because upheld contests are drawn precisely from the tail the model handles worst. Operators who plumb the upheld set back into the bound and the model close the loop that every governance framework describes abstractly; operators who file them as complaints have a customer-service metric and no learning.

The discipline that fails first here is reviewer authority. It is easy to staff an appeals function and hard to give it a reverse button, because the original action was often executed by a pipeline that only runs forwards — a suspension writes a flag, a decline writes a status, and there is no supported path to unwrite either. Reviewers then 'recommend' reinstatement, an engineer runs a manual fix days later, and the published clock quietly becomes fiction. Build the reversal action at the same time as the action itself, or the appeals team is a complaints desk with a better name.

In practice

The reviewer who could not reverse

A marketplace operator published a seller-appeals process with a two-business-day target after suspending accounts on an automated risk score. The appeals team was staffed, trained and fast: most cases were assessed within a day. Reinstatement, however, averaged over a week, because the suspension had been executed by a pipeline that wrote a flag no interface could clear. Appeals could only raise an engineering ticket. The published clock measured the assessment; the seller experienced the reinstatement, and lost a week of trading. The fix was two days of engineering — a supported reverse action with its own log entry — and it should have shipped with the suspension.

What it looks like

  • A statement of reasons travels with the action, not after a complaint
  • Contest intake exists with a published route and a response clock
  • The reviewer has authority to reverse and a reversal action that exists
  • Upheld contests feed back into the bound, not just into the case file

Diagnostic signals you can check this week

  • Ask a reviewer to reverse one live automated decision while you watch. Time it, and count the systems they touch
  • Compare the published response clock with the elapsed time to the person actually being made whole
  • Read the statement of reasons your system sends. If a colleague cannot tell from it what the person should do differently, it is a notice rather than a reason
  • Ask what happened to last quarter's upheld contests. If the answer is 'they were resolved', the loop is open

Anti-pattern · Measuring contests by volume instead of by upheld share

Contest volume is read as a cost line, so the instinct is to drive it down — and the cheapest way to drive it down is to make the route harder to find. That is exactly backwards. A low contest volume with a high upheld share means the route is hidden and the decisions are wrong; a higher volume with a low upheld share means the route is visible and the bounds are about right. Watch the upheld share per decision class as the primary number, publish the route as prominently as you publish the refusal, and treat a sudden rise in upheld share as an operational alarm rather than a service-level statistic.

What holds you here

Evidence is produced on request rather than continuously, so every audit, regulator question and escalation becomes a project with a deadline someone else set.

Highest-leverage next move

Make the evidence continuous: standing exports of decision volumes, bound-breach attempts, override rates and upheld-contest share, on the trading calendar.

Cost of leaving

Effort
12–18 months
Team
Platform team, trust-and-safety or contact-centre owner, legal, plus analytics on the upheld set
Risk
Higher — you are publishing reasons, and every published reason has to be true and consistent with the record
To next stage
12–18 months

If this is you, the next step is

One decision class, from statement of reasons to reversal action to the bound revision it triggers.

Design one contest route end to end

Stage 5

Self-evidencing

4% of operators sit here

Self-evidencing is the stage at which the estate produces its own decision evidence continuously and revises its own bounds when that evidence moves.

Stage 5 is narrower and less exciting than the word suggests, and it is emphatically not more autonomy. It is the opposite: an estate that continuously proves what its automation did. The decision record, the bound-breach log, the override and upheld-contest rates are produced as a matter of course, in a form a regulator, an auditor, a marketplace partner or a large customer can be handed. The characteristic experience of a stage-5 operator during an information request is a query rather than a programme.

The operating signal is the bound-review trigger. Override rate, upheld-contest share and attempted bound breaches are watched as leading indicators, keyed to the events that actually change the input distribution in retail: a market launch, a category expansion, a payment-provider migration, a peak trading week. When one of them moves, a bound review is raised automatically. This is where the discipline pays for itself, because a bound set under last year's conditions is not wrong on the day the conditions change — it is wrong quietly, for weeks, until someone complains loudly enough to be heard.

Stage 5 is also the stage most likely to regress, and the regression is always structural rather than negligent. A platform migration moves a decision out of the path where the bound was enforced. A newly acquired brand arrives with its own console and its own defaults. A vendor consolidates two decisions into one product feature. None of these looks like a governance event and all of them invalidate a bound. Maturity here is not the absence of regression; it is detecting it in days, from your own telemetry, before the people affected have to tell you.

In practice

The re-bound after a market launch

An operator running fraud declines under enforced bounds launched into a new European market with different address and postcode conventions. Within six weeks the upheld share of contested declines in that market had roughly tripled while the overall rate barely moved — the aggregate hid it, the per-market cut did not. Because the upheld share was a standing metric with a review trigger rather than a quarterly report, a bound review was raised and the market's automatic-decline ceiling was lowered while the model was retrained on local address data. The customers most affected were the newest, which is precisely the cohort a launch cannot afford to lose.

What it looks like

  • Decision evidence is a standing export, not an audit deliverable
  • Bound review is triggered by metric movement, not only by the calendar
  • Any single decision can be reconstructed in minutes, not weeks
  • A new decision class cannot go live without a register entry and a bound

Diagnostic signals you can check this week

  • Ask for the reconstruction of one specific decision from six months ago and time the answer
  • Check whether any bound has been revised in the last two quarters, and what triggered the revision
  • Ask what happens in the register when a platform migration ships. If the answer is 'we would review it', the trigger is manual
  • Look for a decision class that went live in the last year and check whether it has a register entry and a bound. New classes are where the discipline leaks

Anti-pattern · Treating the evidence pack as a report rather than a control

The pack gets built, circulated monthly and admired. Nothing in the business changes when a number in it moves, so within two quarters it is produced by a junior analyst and read by nobody. Evidence is only a control if it is wired to an action: a defined movement in upheld share or bound-breach attempts raises a bound review with a named owner and a due date, the same way a failed monitor raises an incident. If no threshold in the pack can force a change, you have documentation, and documentation regresses to decoration.

What holds you here

Sustaining it is a change-control problem: every migration, market launch, acquisition and peak silently invalidates a bound that was set under different conditions.

Highest-leverage next move

Tie bound review to the events that genuinely shift the distribution — market launches, vendor migrations, category expansions and peak — not only to a quarterly calendar.

Cost of leaving

Effort
Continuous
Team
Platform team plus a standing decision forum with trading, legal, operations and trust and safety
Risk
Concentrated — infrequent, high consequence, and both regulatory and reputational when it lands

If this is you, the next step is

We stress-test the register entry, the bound, the record and the reversal against a real case.

Audit one automated decision path end to end

Where retail and e-commerce operators actually sit

The distribution across the ladder, and why the registered-to-bounded step is the one most estates never take.

Most retail and e-commerce operators are at stage 2, registered but unbounded. They have done the genuinely hard organisational work of finding and listing the automated decisions in the estate, and they have stopped one step short of the change that would make the list true — moving authority out of the document and into the path the action travels.

Distribution of retail and e-commerce operators across the five stages

Stage 2 is both the mode and the plateau. The drop from registered to bounded is the largest single transition loss on the ladder, because it is the first step that changes what the estate does rather than what it says.

Share of operators

  • 21% — 1 · Undeclared
  • 38% — 2 · Registered (the plateau)
  • 24% — 3 · Bounded
  • 13% — 4 · Contestable
  • 4% — 5 · Self-evidencing

Source: Illustrative distribution, synthesised from European Commission DSA transparency reporting, ICO automated decision-making guidance and NRF retail research

That database is the closest thing retail has to a public census of automated decision-making at scale, and its shape is instructive. Two of the most-reported violation categories — unsafe, non-compliant or prohibited products, and consumer information infringements — are marketplace listing decisions, which means a large share of those automated calls are commercial decisions about someone's ability to sell. The database's own framing is worth reading in full alongside the Commission's Digital Services Act landing page (opens in a new tab), because the obligation it implements is precisely the one this page is about: telling the person what was decided and why.

The Digital Services Act (DSA) obliges providers of hosting services to inform their users of the content moderation decisions they take and explain the reasons behind those decisions in so-called statements of reasons.

The distribution above is illustrative and model-derived rather than measured — it synthesises the automation share visible in the Commission's database, the maturity signals in the ICO's guidance on AI and data protection (opens in a new tab), and the operational picture in NRF's retail research (opens in a new tab). Read it as the shape of the problem rather than as a market survey. What is not illustrative is the direction of the gap: automated decisions are being taken at a scale that only the largest platforms currently report on, by retailers whose registers, where they exist at all, describe intentions rather than thresholds.

The retail decision register: every automated decision, mapped

Eight decision classes that exist in almost every retail estate — the system each executes in, the bound that must be written, what the person is owed, and the KPI at risk.

The register is the centrepiece of this discipline, and it is a table of decisions rather than of systems or models. Each row names one automated decision, the system of record where it executes, who or what is authorised to take it today, the numeric bound that has to be written down, what the person on the other end is owed when it goes against them, and the trading KPI that moves when the bound is wrong. Eight classes cover most of what a retail and e-commerce estate actually decides.

DecisionSystem of recordBound to writeWhat the person is owedKPI at risk
Checkout fraud hold or declineOMS + PSPAutomatic decline only below a stated order value, outside top loyalty tiers, and not in markets opened within six monthsA specific reason and an alternative route to complete the orderFalse-decline rate, checkout conversion
Returns-abuse scoring and refund refusalOMS / RMA + CRMNo automatic refusal above a basket-value ceiling; no automatic account restriction under any scoreA statement of reasons and a human review on requestRefund rate, CSAT, upheld-contest share
Account restriction or closureCRM / identity platformHuman-only above a stated lifetime-value threshold; never solely automated where the effect is loss of access to essential goodsNotice, the reason, and a route to human interventionComplaint volume, retention, regulatory exposure
Personalised pricing and promotion targetingPricing engine + CDPFloor and ceiling per category; no personalised price without the disclosure the consumer rules requireDisclosure that the price was personalised by automated meansGross margin, promotional leakage, trust
Search ranking and recommendationsSearch / merchandising platform + PIMDeclared own-brand or sponsored boost; blast-radius cap on any single ranking changeThe main parameters of the ranking, publishedConversion, GMV mix, category share
Marketplace listing removal or seller suspensionSeller console / trust-and-safety platformSuspension only after a warning, above a stated repeat threshold, with a human above the lineA statement of reasons, internal complaint handling and an out-of-court routeSeller reinstatement time, GMV at risk
Colleague task allocation and schedulingWorkforce management systemNo automated performance consequence; manager override always available and loggedAn explanation and an override they can actually useAttrition, task completion, on-shelf availability
Checkout credit or pay-later eligibilityPSP / credit partnerNo solely automated refusal without a route to human review, whoever's model produced itThe reason, a human review, and the ability to make representationsCheckout conversion, complaint volume, AOV
The retail and e-commerce decision register. 'Bound to write' is expressed in the units the platform already uses — order value, customer tier, market, category, daily volume — because a bound in any other unit cannot be enforced in the path.

Three columns in that table do more work than the rest. 'System of record' decides your timeline, because a decision executing inside a vendor's platform travels through the vendor's change process rather than yours — which is why the checkout fraud decline, the decision most retailers take most often, is usually the least governed. 'Bound to write' has to be expressed in units the platform already understands, or it cannot be enforced. And 'what the person is owed' is where the regulatory frames land: the same row carries a data protection duty, a consumer-law disclosure or a platform-regulation notice depending on who the decision touches and how hard it is to undo.

One row deserves particular attention because of the money moving through it. NRF and Happy Returns put total US retail returns for 2024 at a projected $890 billion (opens in a new tab), with 93% of the retailers surveyed describing return fraud and other exploitative behaviour as a significant problem for their business. That combination — enormous legitimate volume and a real abuse problem inside it — is exactly the pressure that pushes returns decisions into automation faster than any other class, and it is why the returns row is where an unbounded automatic refusal does the most reputational damage per pound saved.

Two practical notes on making the register reconstructable. First, identify the thing the decision was about with a stable key rather than an internal code: a decision record naming a SKU that has since been reused describes a different product than the one that was suppressed, whereas GS1 identifiers for products and parties (opens in a new tab) survive catalogue churn and are the keys a marketplace partner or an auditor will recognise. Second, register per market, not per group: cross-border trading is the norm in European e-commerce — Ecommerce Europe (opens in a new tab) tracks the market's structure across member states — and a single group-level row hides the fact that the same decision runs under different thresholds and different obligations in each territory it touches.

Those duties attach per decision, not per company. Under the UK GDPR and the GDPR, a decision taken solely by automated means with a legal or similarly significant effect carries specific requirements — information, a route to human intervention, the ability to contest — set out in the ICO's guidance on automated decision-making and profiling (opens in a new tab) and interpreted across the EU by the European Data Protection Board (opens in a new tab). Under the EU AI Act (opens in a new tab), a handful of retail decisions fall into the high-risk tier — creditworthiness at checkout, and systems used for worker task allocation and monitoring — while customer-facing assistants and synthetic product imagery carry transparency duties instead. Under the Digital Services Act (opens in a new tab), a marketplace owes a statement of reasons and an internal complaint-handling route for a listing removal or seller suspension. And under EU consumer protection law (opens in a new tab), a price personalised by automated decision-making has to be disclosed as such before the customer buys.

The register is also where vendor reality gets recorded honestly. Where a decision executes inside a payment, identity or trust-and-safety platform you do not control, the row should say so, name what the contract obliges the vendor to provide — reason codes, decision logs, retention, notice of model changes — and record what it does not. Most retailers discover at least one row where the honest entry is 'we cannot currently reconstruct this decision and our contract does not require the vendor to help', and that entry is more useful than any policy statement, because it is a procurement action with a date.

Four authority classes and the four dimensions that set your rung

How to decide which decisions a machine may take alone — and why the answer comes from the consequence of the action, never from the confidence of the model.

Every automated decision in a retail estate belongs in one of four authority classes, and the class is set by two properties of the action rather than by any property of the model: how much it affects the person, and how hard it is to undo. Assist means the model advises and a person decides. Bounded automatic means the machine acts inside written numeric limits. Notified automatic means the machine acts, the person is told and a contest route opens. Human-only means no automated execution at any confidence level.

  • Decision inventory

    Whether the estate can produce a list of the automated decisions it takes about customers, sellers and colleagues, with an owner against each. The binding test is not whether a register exists but whether it agrees with the consoles: a register that has not been reconciled against live platform configuration since the last vendor release is a historical document. This is also the dimension where a model inventory gets mistaken for a decision inventory, which understates the estate by roughly the number of rules-driven decisions in it.

  • Authority and bounds

    Whether it is written down which decisions a machine may take alone and within what numeric limits, and whether the system enforces those limits. This is overwhelmingly the lowest-scoring dimension in retail estates, and it is the one that separates stage 2 from stage 3. The tell is simple: try to execute an out-of-bounds action in a test environment. If it succeeds, the bound is a sentence in a document rather than a property of the system.

  • Oversight and contestability

    Whether the person a decision lands on is told, and whether a specific named human can reverse it within a stated time. The ICO's checklist for Article 22 processing (opens in a new tab) is a useful external benchmark here — it asks for a simple way for people to request reconsideration and for identified staff authorised to carry out reviews and change decisions. 'Authorised to change decisions' is the phrase most retail appeals functions cannot honestly claim, because the reversal action does not exist in the system.

  • Decision evidence

    Whether a single decision can be reconstructed months later with its inputs, model version, policy version, rule and outcome — and whether that evidence is produced continuously or assembled on request. Reconstruction time is the most testable governance property there is, and the one every audit reduces to. The NIST AI Risk Management Framework (opens in a new tab) and its companion playbook (opens in a new tab) are the most practical published scaffolding for the measure-and-manage side of this dimension.

The four dimensions gate each other, which is why the lowest is the real rung. Excellent evidence on an unregistered decision is a log nobody will think to read. A published contest route on an unbounded decision generates contests faster than the bound generates correct outcomes. And a beautifully written authority class means nothing if the estate contains three decision classes that never made it into the register — which is the normal condition after an acquisition, a replatform or a market launch.

Which authority class does this decision belong in?

Plot the action, not the model. The vertical axis is how much the decision affects the person; the horizontal axis is how easily the action can be undone. Confidence scores do not appear on this matrix, deliberately — a model's confidence is least reliable exactly when the distribution has shifted, which is when the bound matters most.

Human-only

  • Account closure, seller termination, credit refusal at checkout
  • No automated execution at any confidence level
  • The model ranks the queue; a named person decides and signs

Notified automatic

  • Order holds, refund refusals, listing removals, shift changes
  • Act automatically, but tell the person and open the contest route
  • Reversal action and clock ship with the decision, not after it

Bounded automatic with a ceiling

  • Price changes published to the site, markdown execution, promo suppression
  • Cap the blast radius: value, category, share of catalogue, per-hour volume
  • A wrong price is commercial, but it is also public and hard to unsay

Free automatic, monitored in aggregate

  • Ranking, recommendations, size and fit suggestions, replenishment
  • Governed by aggregate monitoring rather than per-decision review
  • Publish the main ranking parameters; watch drift and category share
Effect on the person — top: Legal or similarly significant — money, access, livelihood, bottom: Commercial only — ranking, recommendations, stock allocation
Reversibility of the action — left: Hard to undo — account closed, seller suspended, price already shown, right: Trivially reversible — hold released, listing restored, refund reissued

The bottom-left quadrant is the one retailers get wrong most often, because a price change reads as commercial and therefore low-risk. It is commercial and it is also published: a mis-bounded automatic markdown appears on the site, in the feed, in comparison engines and in customers' screenshots within minutes, and cannot be unsaid even after it is corrected. Bound it by blast radius rather than by consequence to the individual — a share-of-catalogue cap and a per-hour volume limit will save more margin than any confidence threshold.

What governed automated decisions look like in public

Two publicly reported programmes, read against the ladder. Neither is an Atomic Loops engagement — each links to the operator's own published material.

The most useful public evidence for this discipline comes from operators large enough that their automated decisions had to be described in writing. In both cases below the interesting artefact is not the model — it is what the operator published about how the decision is made, who can challenge it and what happens next, because that publication is what makes the automation defensible at the volumes involved.

Two programmes read against the decision-governance ladder

Outcomes as described in the operators' own published material; verify figures against the linked source before reusing them, and note that we have not independently audited them. The images are generated industry scenes from our library, not operator photography, and imply no endorsement.

Illustrative fulfilment-centre scene with automated goods movement and colleagues reviewing exceptions on tablets — a generated industry image, not Amazon photographyAmazonGlobal marketplace and retailer · millions of third-party sellers34
Challenge
Enforcement decisions at marketplace scale — blocking suspect listings, removing counterfeit products and restricting selling accounts — cannot be taken by people at the volume required, but each one removes a seller's ability to trade and therefore has to be explicable and reversible.
Approach
Amazon publishes how its brand protection and enforcement machinery works through its brand services programme and its policy newsroom: automated screening applied before listings publish, brand-owner reporting routes, and defined channels through which sellers and rights owners can challenge an enforcement decision and have an account or listing reinstated.
Reported outcome
As reported by Amazon in its own brand protection publications, automated systems screen listings at scale before publication, and the company maintains published appeal and reinstatement routes for sellers subject to enforcement action.
What it shows about the curveAt marketplace volume, contestability is not a nicety bolted on for regulators — it is the only mechanism that keeps large-scale automated enforcement usable, because it converts the false positives that volume guarantees into cases that can be corrected rather than into permanently lost sellers.

Amazon — Brand Services and policy newsroom (opens in a new tab)

Illustrative warehouse and store back-of-house scene with colleagues reviewing task allocation on handheld devices — a generated industry image, not Home Depot photographyThe Home DepotHome improvement retailer · 2,000+ stores, 400k+ associates23
Challenge
Directing store associates to the right task at the right time — restocking, picking online orders, resolving on-shelf gaps — means an algorithm making decisions that shape a colleague's working day, which is a materially different governance problem from ranking products.
Approach
The Home Depot has publicly described building associate-facing tools that surface prioritised tasks on handheld devices, keeping the store leader's judgement in the loop rather than issuing instructions that cannot be varied — an assist-class decision with a human decider, in the terms of this page.
Reported outcome
As described in the company's own newsroom material, associate-facing task and product tools are deployed across its store estate to help colleagues find stock and prioritise work, with store teams retaining the decision on how work is sequenced.
What it shows about the curveColleague-facing decisions belong in the assist class by default, and the override is the governance artefact: a task recommendation a manager can vary and that logs the variation stays governable, while the same recommendation wired to a performance consequence becomes a high-risk decision under EU rules on worker management.

The Home Depot — corporate newsroom (opens in a new tab)

Read together, the two cases bracket the problem. One is enforcement at a scale where automation is unavoidable and contestability is the control; the other is a workplace decision at a scale where automation is optional and the override is the control. Neither operator solved governance by writing a policy — both solved it by shipping a specific mechanism attached to a specific decision, which is the only form of this work that survives contact with a trading week. Both also chose to publish, which is worth noticing on its own: Amazon's policy newsroom (opens in a new tab) and The Home Depot's corporate newsroom are where the descriptions above come from, and a decision an operator is willing to describe in public is usually one it has already had to bound internally.

The decision record, layer by layer

What has to exist for a decision to be explainable, contestable and reversible — and which layer you can defer.

A defensible automated decision needs six layers, and the order in which you build them decides whether the estate compounds or stalls. The architecture below is deliberately vendor-neutral: every layer is defined by what it must guarantee rather than by which product supplies it, because in most retail estates several of these layers already exist in the order management, payments or trust-and-safety platforms and the work is connecting them rather than buying them.

Layers required by stage

Each layer is annotated with the stage that first requires it. An estate trying to reach stage 4 without the record layer is running an appeals process over decisions it cannot reconstruct — which is worse than having no appeals process, because it makes a promise it cannot keep.

  1. Decision inventory

    Stage 2+

    • Register entryOne row per decision, not per model
    • Named ownerA trading owner, not only an engineering one
    • System of recordWhere the action actually executes
  2. Authority and bounds

    Stage 3+

    • Authority classAssist, bounded, notified or human-only
    • Numeric boundValue, tier, market, category, volume
    • Bound check in the pathOut-of-bounds actions queue, not execute
  3. Action and notice

    Stage 3+

    • Action in the system of recordA hold or flag, reversible by design
    • Statement of reasonsSpecific enough to act on, not a template
    • Disclosure surfaceWhere personalisation or automation is declared
  4. Record

    Stage 3+

    • Input snapshotThe values as they were at decision time
    • Version stampsModel, policy and rule versions in force
    • Reason codes and retention clockWhy it fired, and how long it is kept
  5. Contest and reversal

    Stage 4+

    • Intake with a clockPublished route, started when they ask
    • Reviewer authorityCan change the decision, not recommend it
    • Reversal actionA supported operation with its own log entry
  6. Evidence and review

    Stage 5+

    • Standing exportVolumes, breaches, overrides, upheld share
    • Review triggerMetric movement raises a bound review
    • Change gateNo new decision class without a register entry

Pipeline described

  1. Decision inventory (stage 2+) — Register entry: One row per decision, not per model; Named owner: A trading owner, not only an engineering one; System of record: Where the action actually executes
  2. Authority and bounds (stage 3+) — Authority class: Assist, bounded, notified or human-only; Numeric bound: Value, tier, market, category, volume; Bound check in the path: Out-of-bounds actions queue, not execute
  3. Action and notice (stage 3+) — Action in the system of record: A hold or flag, reversible by design; Statement of reasons: Specific enough to act on, not a template; Disclosure surface: Where personalisation or automation is declared
  4. Record (stage 3+) — Input snapshot: The values as they were at decision time; Version stamps: Model, policy and rule versions in force; Reason codes and retention clock: Why it fired, and how long it is kept
  5. Contest and reversal (stage 4+) — Intake with a clock: Published route, started when they ask; Reviewer authority: Can change the decision, not recommend it; Reversal action: A supported operation with its own log entry
  6. Evidence and review (stage 5+) — Standing export: Volumes, breaches, overrides, upheld share; Review trigger: Metric movement raises a bound review; Change gate: No new decision class without a register entry
Step-by-step insights
Decision inventory — one row per decision, always
The temptation is to collapse rows: 'fraud decisioning' as a single entry covering scoring, holding, declining and restricting. Collapsing hides exactly the variation that matters, because those four actions differ by an order of magnitude in consequence and belong in different authority classes. Split until each row has a single action, a single system of record and a single owner. Retail estates typically end up with between fifteen and forty rows, which is a manageable artefact and a genuinely surprising number to most executives who expected five.
Authority and bounds — expressed in the platform's own units
A bound stated as 'low-risk cases only' cannot be enforced, because no field in the order management system holds 'low risk'. Write it in units the platform already carries: order value in the trading currency, loyalty tier, market code, category node, decisions per hour. This constraint is a feature — it forces the conversation between trading, legal and engineering to happen in numbers, and it means the check that enforces the bound is a comparison against fields that already exist rather than a new data model.
Action and notice — reshape the action before you write the notice
Before drafting a statement of reasons, ask whether the action can be made reversible: hold rather than decline, restrict rather than close, demote rather than delist, flag rather than refuse. A reversible action lowers the notice burden, shortens the contest clock and reduces the cost of being wrong, all at once. Where the action genuinely cannot be reversed — a personalised price already shown, a marketplace listing removed during a promotional window — that irreversibility is the argument for a lower automatic ceiling, not for a better-worded apology.
Record — written at decision time, or not at all
The record captures the inputs as they were, the model, policy and rule versions in force, the reason codes and the outcome, at the instant the decision executes. It cannot be reconstructed afterwards from application logs because everything moves: addresses change, models retrain, catalogues reindex, rules get edited. Set the retention clock deliberately against both the contest window and the data minimisation duty — long enough that a customer who complains three months later can be answered, and no longer than you can justify, which is a conversation to have with legal once and then encode.
Contest and reversal — the reverse action is the whole layer
Intake and reviewer authority are organisational; the reverse action is engineering, and it is the piece that is almost always missing. If suspending a seller writes a flag that no interface can clear, then the appeals team's authority is theoretical and the published clock measures the wrong thing. Build the reversal at the same time as the action, give it its own log entry naming the reviewer and the reason, and rehearse it once a quarter on a live case — the same way a rollback is rehearsed, and for the same reason.
Evidence and review — wire the numbers to an action or they decay
A standing pack of decision volumes, bound-breach attempts, override rates and upheld-contest share is only a control if a defined movement forces something to happen. Set thresholds per decision class, and have a breach raise a bound review with a named owner and a due date, exactly as a failed monitor raises an incident. Tie the review calendar to the events that shift retail distributions — peak, market launch, category expansion, payment or platform migration — rather than to the financial quarter, which correlates with nothing the models care about.

The layer most often skipped is the record, and skipping it quietly caps the estate at stage 3 forever. It is invisible while things go well, costs a sprint to add and cannot be retrofitted for the period you actually need it for. If you build one thing from this section before the next peak, build the record — everything above it becomes possible afterwards, and nothing above it is possible without it.

A 90-day plan: bringing checkout fraud declines under a decision record

The registered-to-bounded transition made concrete on the single decision most retailers automate most often — and govern least. Contains no model development.

Moving one rung takes about 90 days when it is scoped to a single decision class, and several years when it is scoped to the estate. To make that concrete, the plan below runs the transition on the automated checkout fraud decline: the decision a retailer takes tens of thousands of times a week, that executes inside a payment service provider, that no one internally usually owns, and whose errors cost conversion and goodwill simultaneously. The model is the vendor's and stays untouched — the quarter is spent on ownership, bounds, records and a route back.

Registered to bounded on checkout declines, in one quarter

One market, one payment flow, one owner. If a phase needs more than its window, narrow the scope — one market rather than three, one card scheme rather than all — rather than extending the plan.

  1. Days 1–15

    Name the decision and baseline it

    Pull 90 days of decline data from the payment service provider and the order management system for one market. Split declines into issuer declines, vendor-model declines and your own rules. Compute a false-decline proxy: declined orders where the same customer succeeded on a retry, on another card, or through the contact centre. Name two owners — a fraud or payments owner for the threshold and a trading owner for the conversion consequence — and write the register row.

    A register row, two named owners, a decline baseline by reason

  2. Days 16–45

    Write the authority class and put the bound in the path

    Declare the class: automatic decline permitted below a stated order value, outside the top loyalty tiers, and not in markets opened within six months; everything else becomes a hold routed to the review queue. Implement the check where the decline is written, not in a document. Size the queue against last peak, not last Tuesday. Replace the generic checkout error with an honest, specific message and a route — the cheapest conversion work in the plan.

    An enforced bound, a sized queue, an honest checkout message

  3. Days 46–70

    Write the record and open the contest route

    Every automatic decline and hold writes a decision record: input snapshot, vendor model version, your policy version, reason codes, outcome, timestamp. Give contact-centre agents a supported one-action release with a mandatory reason code, and publish a route for customers who believe a refusal was wrong, with a stated response time. Log every release and every contest outcome against the decision record it belongs to.

    Reconstructable decisions, a live reversal action, a published clock

  4. Days 71–90

    Attribute, then re-bound

    Hold a comparable customer segment on the previous policy so the difference is attributable rather than asserted. Report recovered orders and recovered revenue, false-decline proxy, agent release rate, contest volume and upheld share — not model accuracy. Then revise the bound using the upheld contests, and put the next review on the trading calendar against peak and the next market launch.

    An attributable revenue delta and a first bound revision

The order matters

  1. Ownership before bounds

    A bound with no owner is a setting. Name the fraud owner and the trading owner in week one, because the argument you are really resolving — how much fraud loss the business will accept to recover declined good orders — is a commercial decision that nobody in engineering can make and nobody in fraud should make alone.

  2. Bounds before records

    It is tempting to start with logging because it is the least contentious work. Start with the bound instead: it is the only step that changes what the estate does, and the volume of records worth keeping is much clearer once the automatic scope has been narrowed to the cases you actually chose to automate.

  3. Records before routes

    Do not publish a contest route until a decision can be reconstructed. A published route over unreconstructable decisions produces cases your reviewers cannot answer, which converts a governance improvement into a customer-service failure and a written record of your own inability to explain the decision.

  4. One market before the estate

    Run the whole loop in one market with one payment flow before extending. The second market costs a fraction of the first if the bound check and the record are shared, and roughly the same as the first if each market implements its own — which is the difference between a repeatable capability and a project you will run four times.

The commercial case for this quarter does not rest on compliance. Baymard's checkout research puts the recoverable opportunity from better checkout design across the US and EU at around $260 billion, from an average 35.26% potential conversion uplift on large sites (opens in a new tab), and a badly bounded fraud decline is a checkout failure with none of the usability upside — the customer wanted to buy, had a valid card and was refused without a reason. Every recovered good order in the day-71 report is revenue that was already inside the funnel, which is why this plan tends to survive a budget conversation that a governance programme would not.

Instrumenting decision governance: formula, source, cadence

The metrics that prove decisions are governed rather than described — where each comes from, how often to read it, and the rung at which it first means something.

A governance claim you cannot instrument is an opinion with a template. Every metric below reduces to counts and timestamps that the order management system, the payment platform, the seller console or your own decision log already records — the work is joining them, not creating them. The table is the build sheet: formula, source, cadence, and the rung at which the number first measures something real.

MetricFormula / readSourceCadenceHonest from
Decision coverageRegistered decisions ÷ automated decisions found in a console walkRegister + platform configurationQuarterly and after every releaseStage 2
Register driftRegister rows whose live threshold differs from the recorded oneRegister vs console exportMonthlyStage 2
Bound-breach attemptsActions blocked by the bound check ÷ actions attemptedDecision service logWeeklyStage 3
Automatic share by classActions executed without a human ÷ actions in that decision classDecision logWeeklyStage 3
Override rateColleague reversals ÷ automated actions surfaced to colleaguesOMS / seller console override logWeeklyStage 3
False-decline proxyDeclined orders later completed by the same customer ÷ declinesPSP + OMS order historyWeeklyStage 3
Time to noticeAction timestamp → statement of reasons sentNotification servicePer decisionStage 4
Contest upheld shareContests decided in the person's favour ÷ contests closedContest case systemWeekly, by market and cohortStage 4
Time to make wholeContest raised → the person's position actually restoredCase system + system of recordPer caseStage 4
Reconstruction timeElapsed time to return inputs, versions and rule for one past decisionDecision record storeQuarterly drillStage 4
Bound ageDays since the bound was last reviewed against live evidencePolicy version historyMonthlyStage 5
Instrumentation build sheet for retail decision governance. 'Honest from' is the rung at which the metric stops being a proxy and starts being a measurement.

Two of those metrics carry more diagnostic weight than the rest. Register drift is the fastest way to tell a live governance system from a documentary one — if it is above zero and nobody knew, the register is describing an estate that no longer exists. And the contest upheld share, cut by market and by customer cohort rather than reported in aggregate, is the earliest available warning that a bound has stopped fitting the world: it moves weeks before complaint volume does, and months before anything reaches a regulator or a marketplace partner.

Bounded-decision readiness checklist

Run this against one decision class — the checkout decline is the usual candidate. If you cannot tick all seven, that class is not yet bounded, however good the model behind it is. Tick as you go; this list works with no JavaScript.

0 of 7 ticked

Nothing ticked — and that is the honest starting point

Most estates start here on their highest-volume decision, because that decision arrived configured inside a vendor platform and was never chosen. Do not start with tooling. Start with the register row and the two owners: an afternoon of work that converts a setting into a decision somebody is accountable for, which is the precondition for everything else on this list.

Failure modes that send decision governance backwards

Governance is not monotonic. Four regressions account for almost all of it, and none of them looks like a governance event on the day it happens.

Decision governance regresses quietly, because every mechanism that regresses keeps producing output. The bound still exists in the document, the appeals inbox still receives mail, the register still lists the decision — and the estate has moved out from under all three. Four patterns account for almost every regression we see in retail.

Likelihood: highImpact: high

A platform migration moves the decision out from under the bound

The bound check lives in the service that wrote the action. Replatform the checkout, migrate to a new payment provider or consolidate two trust-and-safety tools, and the action gets written somewhere else — with the vendor's defaults, not yours. Nothing alerts, because from the register's point of view the decision is unchanged.

PreventionAdd a migration gate: no decision-writing service goes live without its bound check and its record, verified by an out-of-bounds test in staging.

Likelihood: highImpact: medium

An acquisition or new market arrives with its own consoles

A newly acquired brand or a market launched on a local platform brings decisions nobody registered, thresholds nobody wrote and a checkout message nobody read. It is the single most reliable way for an estate at stage 4 to be operating at stage 1 in one of its territories without anybody noticing for a year.

PreventionMake register coverage a day-one integration deliverable, ahead of catalogue and payments work, and re-baseline the upheld share per market.

Likelihood: mediumImpact: medium

The bound is widened during a bad week and never narrowed

Fraud spikes, the queue backs up, someone widens the automatic ceiling to clear it — a correct operational call. The ceiling then stays wide, because nothing schedules the narrowing and the queue is comfortable. Two quarters later the automatic scope bears no relationship to the class that was agreed.

PreventionEvery emergency bound change carries an expiry date and reverts automatically unless it is re-approved through the normal review.

Likelihood: mediumImpact: high

The appeals function loses its reverse action

A refactor removes the interface that let a reviewer clear a flag, or a new decision class ships without one. Reviewers keep assessing and keep 'recommending' — and the published clock silently starts measuring the assessment rather than the outcome, while the person waits days to be made whole.

PreventionRehearse one live reversal per decision class each quarter and measure time-to-make-whole, not time-to-decide.

What connects all four is that none of them is negligent and none is visible from a governance document. They are ordinary trading and engineering events — a migration, an acquisition, a bad fraud week, a refactor — whose governance consequence is invisible unless something in the estate is watching for it. That is the whole argument for stage 5: not more automation, but continuous evidence, so that the day a bound stops meaning what it meant is a day you find out from your own telemetry rather than from a customer, a seller, a marketplace partner or a regulator.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Automated decision
A decision taken by a system about a customer, a seller or a colleague without a person deciding the individual case. The unit of governance on this page: an order held, a refund refused, a listing removed, a shift reallocated, a price personalised.
Decision register
The inventory of automated decisions in an estate — one row per decision rather than per model — recording the system of record, the owner, the authority class, the bound, what the affected person is owed and the trading KPI at risk.
Authority class
Which of four levels of autonomy a decision is permitted: assist (the model advises, a person decides), bounded automatic (acts inside written limits), notified automatic (acts, tells the person, opens a contest route) or human-only.
Bound
The numeric limit on what a decision may do automatically, expressed in fields the platform already holds — order value, loyalty tier, market, category, actions per hour — so it can be enforced in the path rather than asserted in a document.
Bound check
The test executed in the service that writes the action, which routes an out-of-bounds case to a human queue instead of executing it. Its presence is what distinguishes a governed estate from a documented one.
Decision record
The artefact written at decision time capturing the inputs as they were, the model and policy versions in force, the rule that fired, the reason codes and the outcome. It cannot be reconstructed after the fact, which is why it is written or lost.
Statement of reasons
The notice explaining what was decided and why, sent to the person affected. Under the Digital Services Act, hosting providers must issue one for content moderation decisions and submit it to the Commission's transparency database.
Contestability
The property of a decision being challengeable in practice: a record complete enough to reconstruct it, a reviewer with authority to reverse, a reversal action that exists in the system, and a clock that starts when the person asks.
Upheld-contest share
Contests decided in the affected person's favour divided by contests closed, read per decision class, market and cohort. The earliest reliable signal that a bound has stopped fitting the world, and a free source of labelled errors.
False-decline proxy
Declined orders later completed by the same customer — on a retry, another card or through the contact centre — divided by total declines. The practical stand-in for a true false-positive rate, which no retailer can observe directly.
Register drift
The count of register rows whose live platform threshold differs from the recorded one. Above zero and unnoticed, it means the register describes an estate that no longer exists — the characteristic stage-2 failure.
Time to make whole
Elapsed time from a contest being raised to the person's position actually being restored — the order released, the listing reinstated, the refund paid. Distinct from, and usually much longer than, the time taken to decide the contest.

Frequently asked questions

The questions retail and e-commerce teams ask most often when they start governing decisions rather than models.

What counts as an automated decision in a retail estate?

Any decision the system takes about a person without someone deciding the individual case: holding or declining an order, refusing a refund, restricting an account, removing a listing, suspending a seller, personalising a price, allocating a colleague's tasks. Note that a rules engine with no model in it produces automated decisions too — which is why a model inventory understates the estate, often badly. Inventory by action and consequence, then attach whichever model or rule drives each one.

Why govern decisions rather than models?

Because consequence attaches to actions, not to scores. One fraud model routinely drives four decisions of escalating severity — ranking a queue, holding an order, declining a payment, restricting an account — and governing the model as a single object gives all four the same treatment. Governing decisions lets you automate freely where the action is reversible and commercially bounded, while holding account closure or credit refusal behind a person. It is also the only framing your engineers can implement, because a bound has to attach to an action.

How do we decide which decisions a model may take alone?

Plot the action on two axes: how much it affects the person, and how easily it can be undone. Reversible, commercially bounded actions such as ranking or releasing a hold can run automatically with aggregate monitoring. Significant but reversible actions — order holds, refund refusals, listing removals — can run automatically provided the person is told and a contest route exists. Significant and hard to undo — account closure, seller termination, credit refusal — should stay human-decided regardless of model confidence.

Where should the bound actually live?

In the service that writes the action, not in a policy document or a vendor console. The code that creates the hold, the refusal or the suspension checks whether this action falls inside the declared authority class, and routes it to a human queue when it does not. This is usually a small piece of engineering, and it is the change that makes the register true by construction: an action contradicting the register cannot execute, so document and estate cannot silently diverge.

What does GDPR Article 22 mean for e-commerce decisions?

It sets additional rules for decisions taken solely by automated means that have a legal or similarly significant effect on someone. The ICO's guidance requires a valid basis — contract necessity, explicit consent or authorisation in law — plus information about the processing, a simple way to request human intervention or challenge the decision, and regular checks that the system works as intended. In retail this bites hardest on credit and pay-later refusals, account restrictions and any refund refusal that materially affects the customer.

What does the Digital Services Act require of an online marketplace?

Marketplaces are online platforms, so moderation decisions — removing a listing, restricting a seller's visibility, suspending an account — require a statement of reasons to the affected user and submission of that statement to the Commission's transparency database. Platforms must also run an internal complaint-handling system and point users toward out-of-court dispute settlement. Practically, this means the decision record and the contest route are not optional design choices; they are the artefacts the obligation is made of.

Which retail AI decisions are high-risk under the EU AI Act?

Two families matter most for retailers. Systems evaluating creditworthiness — which covers pay-later eligibility presented at your checkout even where a partner's model produces the answer — and systems used for recruitment, task allocation, monitoring or evaluation of workers, which covers algorithmic scheduling and task direction for store colleagues. Customer-facing assistants and synthetic product imagery fall under transparency duties rather than the high-risk tier. Check the Commission's regulatory-framework page for the current application dates before planning against them.

Do we have to tell customers when a price is personalised?

In the EU, yes. Consumer protection rules require traders to inform consumers when the price presented has been personalised on the basis of automated decision-making, before the purchase. Practically this means the disclosure belongs on the product and checkout surfaces rather than buried in a privacy notice, and that the pricing engine has to record which prices were personalised and on what basis — which is the same decision record this page argues for everywhere else, applied to a commercial rather than a rights-affecting decision.

How long does it take to move from registered to bounded?

About 90 days when scoped to a single decision class with two named owners and one market, as in the plan above. The work is ownership, bounds, records and a reversal path rather than model development, because the model is usually the vendor's and stays untouched. Scoping the transition to the whole estate rather than one decision is what turns 90 days into two years, largely because every additional platform adds its own change-approval path.

How do we govern a decision that executes inside a vendor's platform?

Record the truth in the register first: name what the contract obliges the vendor to provide — reason codes, decision logs, retention periods, notice of model or threshold changes — and what it does not. Then close the gap contractually at the next renewal, and in the meantime move what you can into your own path: hold rather than let the vendor decline outright, log what you receive, and put your own bound check in front of the vendor's action wherever the integration allows it.

What should we report to the board about automated decisions?

Five numbers per decision class, quarterly: automatic share, bound-breach attempts, override rate, contest volume and upheld-contest share, with the last two cut by market and customer cohort. Add register drift for the estate as a whole. Boards respond to the upheld share because it is unambiguous — it counts decisions your own reviewers judged wrong — and to register drift because it answers the only governance question a board really has, which is whether the estate still behaves the way the paperwork says.

Does an ISO or NIST framework replace this work?

No, but both help you structure it. The NIST AI Risk Management Framework and its playbook give a well-tested vocabulary for governing, mapping, measuring and managing AI risk, and ISO/IEC 42001 provides an auditable management-system shell. Neither tells you which of your decisions may be automated, at what order value, in which markets — those are commercial judgements specific to your estate. Use the frameworks for the scaffolding and the register for the decisions; auditors will ask for both.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for retail, e-commerce, logistics and manufacturing operators — pricing, ranking, fraud screening, returns and demand systems running against live trading data, wired into the OMS, PIM and payment estate rather than delivered as dashboards. Governance work here means the decision path, the bound and the record, not a policy document.

  • · Production decision systems across checkout, returns, pricing and marketplace operations
  • · Decision registers and authority bounds built with trading, legal and engineering in the same room
  • · Integration-first delivery: OMS write-back, decision logging, tested reversal paths
  • · 18 cited sources on this page

Sources

  1. European CommissionRegulatory framework for artificial intelligence (EU AI Act) (opens in a new tab)
  2. European CommissionDigital Services Act package (opens in a new tab)
  3. European CommissionDSA Transparency Database — statements of reasons (opens in a new tab)
  4. Information Commissioner's OfficeRights related to automated decision making including profiling (opens in a new tab)
  5. Information Commissioner's OfficeGuidance on AI and data protection (opens in a new tab)
  6. EDPBEuropean Data Protection Board (opens in a new tab)
  7. NISTAI Risk Management Framework (opens in a new tab)
  8. NISTAI RMF Playbook (opens in a new tab)
  9. National Retail Federation and Happy Returns2024 Consumer Returns in the Retail Industry (opens in a new tab)
  10. National Retail FederationResearch and insights (opens in a new tab)
  11. Baymard InstituteCart abandonment rate statistics and checkout research (opens in a new tab)
  12. GS1GS1 standards for product and party identification (opens in a new tab)
  13. European CommissionConsumer protection law (opens in a new tab)
  14. Ecommerce EuropeEuropean e-commerce policy and market reporting (opens in a new tab)
  15. ISOArtificial intelligence standards (ISO/IEC 42001) (opens in a new tab)
  16. AmazonBrand Services — brand protection and enforcement (opens in a new tab)
  17. AmazonPolicy news and views (opens in a new tab)
  18. The Home DepotCorporate newsroom (opens in a new tab)

Find out which of your decisions you could not currently defend

We run the assessment with your trading, engineering and legal leads, open your consoles alongside your register to find where the two disagree, and leave you with a costed 90-day plan for the weakest dimension. You keep the plan and the marked-up register whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.