Key takeaways
- RTO risk scoring reads the signals already inside an order, pincode, phone, value, address and payment, and predicts return likelihood before you spend a rupee on shipping.
- No single signal decides anything. A high-return pincode is a nudge, not a verdict, until it stacks with a weak address and a first-time COD buyer.
- The strongest signal in India is a phone with a repeat-RTO history across stores, because behaviour predicts behaviour better than geography does.
- A score is only useful if it maps to an action: allow, WhatsApp-confirm, OTP-gate, force prepaid, or hold. The bands are the point, not the number.
- Start with transparent rules, watch whether the scores match real returns, then add ML for the ambiguous middle. Scoring badly is worse than not scoring.
Here is the uncomfortable truth about COD in India: you can usually tell which orders will come back before you ship a single one. Not perfectly, but well enough to change what you do. The information is right there in the order, the pincode, the phone number, the address, the cart value, the fact that it is COD. RTO risk scoring is just the discipline of reading those signals and acting on them instead of shipping everything blind and hoping.
I want to be precise about what scoring is and is not, because a lot of brands either overtrust it or dismiss it. A risk score is not a lie detector and it is not a reason to reject customers. It is a router. Its entire job is to send each order down the right path so your good buyers sail through and your risky ones get a little friction. Let me show you the signals, how they combine, and what to do with the answer.
The signals that actually predict RTO
Start with the raw inputs. Each one carries a bit of predictive weight on its own, and a lot more when they stack. These are the ones that matter in the Indian context, roughly in order of how much they move the needle.
| Signal | What it tells you | Weight |
|---|---|---|
| Repeat-RTO phone (cross-store) | This buyer has bounced orders before, often across brands | Very high |
| Pincode RTO history | Some pincodes return far above average, structurally | High |
| Address completeness | Short or junk addresses fail delivery constantly | High |
| COD flag | COD orders RTO several times more than prepaid | High |
| Phone validity | Fake or unreachable numbers cannot be worked on NDR | Medium-high |
| Order value | Very high-value COD carries more refusal risk | Medium |
| Order velocity / fraud pattern | Many orders, one address, odd timing signals abuse | Medium |
| New vs returning buyer | A proven buyer on your store is far safer | Medium |
The one people underrate is the repeat-RTO phone. Geography is a blunt instrument, an entire pincode is not risky, some houses in it are. Behaviour is sharp. A phone number that has refused deliveries before, especially across multiple stores in a shared network, is the single cleanest predictor you have. That is why serial returner detection and COD fraud detection sit at the top of any serious model. Pincode history comes next, and it is real, but handle it with more care. Some pincodes in tier-3 India run structurally higher RTO because of access, local address conventions, and patchy courier reach, not because the buyers there are dishonest. You will find the patterns in high-RTO pincodes in India. Use it as a weight in the score, never as a blanket ban, or you will punish honest buyers for the accident of their postcode.
How the signals become a score
There are two layers, and you want both. The first is plain rules. The second is machine learning. Do not skip straight to ML, because rules give you something ML cannot: you can explain exactly why an order was flagged, which matters when you are deciding whether to refuse a customer.
Layer one: transparent rules
Rules are simple, fast, and auditable. You assign points and add them up. Something like:
- Repeat-RTO phone on record: +40
- Address missing house or area: +25
- Pincode in your worst-return decile: +20
- COD selected: +15
- Order value above a threshold you set for your catalogue: +10
- New buyer, never ordered before: +10
- Phone fails a basic validity check: +15
Add the points, and the total lands the order in a band. The beauty of rules is that when an order scores 65, you can say precisely which signals got it there. That transparency is what lets you act confidently. It also lets you tune, because you can see which rule is over-firing and dial it back.
Layer two: machine learning for the middle
Rules are great at the extremes. A repeat-RTO phone on a junk address is obviously high risk, and a returning prepaid buyer is obviously low. The problem is the murky middle, the first-time COD buyer with a decent address in an average pincode. Is that a 12 percent RTO order or a 22 percent one? Rules shrug. A model trained on your own delivered-versus-returned history can read the subtle combinations a human never would. How that model is built and trained is covered in RTO prediction with machine learning.
What to do per risk band
This is the part that matters, and the part most guides skip. A score with no action attached is a vanity metric. The whole reason you scored the order is to treat it differently. Here is a sane banding and the action for each. Tune the thresholds to your own catalogue and margins.
| Band | Score range | Action | Why |
|---|---|---|---|
| Green | 0-25 | Allow, ship as-is | Proven or clean order, do not add friction |
| Yellow | 26-50 | WhatsApp COD confirmation before dispatch | Cheap filter for impulse and accidental orders |
| Orange | 51-70 | OTP-gate COD, or nudge hard to prepaid | Real risk, verify intent or shift the payment |
| Red | 71-100 | Force prepaid or partial-COD, else hold | Too risky to ship COD without skin in the game |
Walk through the logic. Green orders get nothing added, because friction on a good buyer is just lost conversion. Yellow gets a WhatsApp "reply YES to confirm", the cheapest filter there is, detailed in the WhatsApp COD confirmation sequence. Orange orders you either verify with OTP or push toward prepaid with a real incentive. Red orders you do not ship on plain COD at all, you convert them to prepaid or partial-COD so the buyer has something at stake, and if they will not, you hold or drop. That move from orange and red toward prepaid is where scoring pays for itself, because a converted order leaves the RTO pool entirely, you get the cash up front instead of waiting fifteen days on the courier, and the return risk collapses. The mechanics of the shift are in COD-to-prepaid conversion and the offers that actually work are in prepaid incentives that work.
The mistakes that make scoring backfire
First mistake: treating a score as a rejection. It is not. Most of your flagged orders are still worth shipping, just with a confirmation or a prepaid nudge attached. If your scoring is quietly cancelling orders, you have built a revenue leak and called it risk management. Second mistake: over-weighting pincode. Punishing an entire postcode drops good buyers along with bad ones. A postcode does not refuse a delivery, a specific buyer does. Weight the pincode, do not ban on it, and let behaviour signals like the repeat-RTO phone always outrank raw geography.
Third mistake, and the one that quietly rots a good model: scoring and never checking. A score that is never validated against real returns drifts. Festive-season buyers behave nothing like your January buyers, new pincodes open up as couriers expand, and a fraud pattern you tuned for six months ago mutates. You have to feed actual delivered-versus-RTO outcomes back into the score every few weeks, which is the whole idea behind the RTO feedback loop. Score, ship, observe, correct, repeat. A model you set and forget is worse than honest rules you keep an eye on.
Where scoring fits in the bigger picture
Risk scoring is one step in a longer chain, not a standalone fix. It does not repair a broken address field, it just reads it, so address quality and phone verification have to come first or the score is reading garbage in and confidently reporting garbage out. And scoring only creates value if the bands are wired to real actions downstream, the confirmations, the prepaid nudges, the courier choices in courier allocation. The full sequence is in the 9-step RTO reduction playbook, where scoring is step three, sitting exactly between fixing your inputs and acting on the output. Get that placement right and the two ₹1,299 orders I opened with stop being a coin toss: the clean one ships green and untouched, the risky one gets a WhatsApp confirm or a prepaid nudge before it ever leaves your warehouse. You did not refuse anyone. You just stopped shipping blind.
Score every order before it ships
Kwikfy scores RTO risk on every order using pincode, phone reputation, address and payment signals, then routes each one to allow, verify or prepaid automatically.
Start Free →