ReturnWarden

Why return reason codes lie, and what to track instead

Ask a returns team what their top return reason is and the answer is almost always some version of "didn't fit" or "not as described." Ask them whether those codes are true and you get a long pause. Return reason codes are self-reported by customers who have every incentive to pick the most friction-free option, and the dropdown you designed to generate clean analytics mostly generates clean-looking fiction.

This matters because brands build real decisions on this data: which products to fix, which sizes to adjust, where the quality problems are. When the underlying codes are unreliable, the decisions inherit the unreliability. A product team that "fixes" sizing based on reason codes is often fixing a problem that does not exist, while the real problem, wardrobing in occasion categories, bracketing in footwear, hides behind the same bland codes.

The codes fail in predictable ways. Customers pick the first reasonable option to get through the flow faster, which inflates whatever sits at the top of your dropdown. They pick the reason least likely to trigger questions, so "didn't fit" absorbs everything from "I wore it to a wedding" to "I found it cheaper elsewhere." And they pick strategically: experienced abusers learn which reasons route to instant approval and which trigger review, and they select accordingly. Your reason-code analytics are, in part, a record of what your abusers have figured out about your process.

The fix is not better dropdowns. It is treating the reason code as one weak signal among several strong ones, and building your analytics on the strong ones. Warehouse condition at intake is stronger: tags intact or removed, wear marks, odor, resalable as-is. Timing is stronger: returns filed the Monday after every holiday, the same category returned every quarter. Account history is strongest of all: the customer's lifetime return rate, the concentration of their returns in high-abuse categories, the trend over time. None of these come from the customer. All of them are harder to fake.

Cross the reason code against these signals and the lies become visible. A "didn't fit" return on an account with a 60 percent lifetime return rate in occasion dresses, filed the Monday after a wedding-heavy weekend, with warehouse notes saying tags removed cleanly, is not a sizing problem. It is wardrobing wearing a costume. Your reason-code dashboard will file it under fit. Your cross-referenced data will file it under fraud. One of those is useful.

Redesign the flow with this in mind. Keep the reason dropdown short, because long lists get gamed more, not less. Put the most honest options where they are easy to find, and stop treating "other" as a failure: a free-text "other" that a human occasionally reads is more informative than a precise-sounding code everyone lies on. Most importantly, stop routing decisions on the code alone. The reason code can decide the exchange offer or the refund method. It should never decide whether the return gets reviewed.

For product teams, replace reason-code analytics with return-rate-by-SKU analytics. A SKU with a 40 percent return rate has a problem regardless of what the codes say, and a SKU with a 5 percent rate does not, even if its codes look alarming. Then layer in the cross-referenced signals to separate product problems from customer problems: if one SKU's returns concentrate in five accounts, that is not a product issue. If they spread evenly across hundreds of accounts with consistent condition notes, it is.

Reason codes will always be part of the flow, because customers expect to say why they are returning something. Let them. Just stop believing them. The data you generate yourself, from the warehouse, the clock, and the account history, is the data worth building on.

Get a free returns fraud scan