Route Optimization Is a Data Problem Before It's an Algorithm Problem
Teams shop for routing algorithms while their address data, service times, and vehicle constraints are fiction. Route optimization starts with data plumbing.
A distribution operator we worked with had already bought a route optimization engine before we met them. Good one, too. Six months in, drivers were ignoring its routes, dispatchers were re-sequencing stops by hand every morning, and management had concluded that optimization 'doesn't work here.' The algorithm was fine. Its inputs were fiction. Depot-to-stop distances came from addresses that geocoded to the wrong block, service times were a flat 10 minutes whether the stop was a kiosk or a hypermarket with a 40-minute receiving queue, and half the fleet's stated capacities didn't match the trucks. Route optimization software development fails this way constantly, and the pattern is worth understanding before you spend anything on algorithms.
Garbage Coordinates, Garbage Routes
The solver optimizes the world you describe to it, not the world your drivers actually visit. Start with addresses. In GCC and South Asian markets especially, free-text addresses geocode badly: on that operator's data, we measured about 18 percent of delivery points landing more than 200 meters from the true location, which in a dense souq district can mean the wrong side of a one-way system and a 15-minute detour the plan knows nothing about. The fix wasn't clever. Every completed delivery already produced a GPS fix from the driver's phone, so we started correcting stored coordinates from actual delivery events, flagging points where the geocode and the delivery history disagreed. Within eight weeks the location data was self-healing. That one loop did more for route quality than any solver parameter we ever touched.
Service Time Is Where Plans Quietly Die
Drive time between stops is the part software estimates well. Time at the stop is the part everyone makes up. A flat per-stop estimate is off by a factor of four across a mixed customer base, and the errors compound: by stop nine, the plan says 11:40 and the truck arrives at 13:05, misses a receiving window, and the day unravels. The fix, again, is measurement over assumption. Timestamps drivers were already generating, arrival, proof-of-delivery, departure, gave us per-customer service time distributions in a few weeks. Some stops averaged six minutes. Some averaged 47. Feeding real distributions into the solver, plus honest time windows collected from customers rather than guessed, is the difference between a plan drivers follow and a plan they abandon by mid-morning, at which point you're paying for optimization nobody uses.
Constraints Live in Dispatchers' Heads
Every veteran dispatcher holds a constraint model no one ever wrote down. Truck 7's tail lift is broken so it can't take cage deliveries. That client only accepts before 9. Don't send new drivers to the port pickup. When the software doesn't know these things, its routes are technically optimal and practically wrong, and every manual correction teaches the operation to distrust the system a little more. On our project we spent two full days interviewing dispatchers and drivers before writing anything, and turned about 60 unwritten rules into structured constraint data with an owner and a review cadence. Unglamorous work. It's also the reason adoption held: the day the plans stopped needing manual fixes was the day dispatchers stopped overriding them.
Drivers Close the Loop, or Nobody Does
One more thing the algorithm-first crowd misses: your feedback loop is a person in a truck, and the product has to respect that. If reporting a wrong location or an unrealistic time window takes more than two taps, drivers won't do it, and your data stops improving the day the consultants leave. We put a 'this stop is wrong' button directly in the driver app: wrong pin, gate moved, customer closed, receiving takes longer than planned. Each report flowed into a review queue that corrected the master data within a day or two, and drivers could see their reports getting fixed, which mattered more than any incentive scheme we considered. Within three months, drivers had corrected more location and constraint data than the initial cleanup project had. The operations lesson generalizes: the people executing the plan are the only ones positioned to tell you where the plan's model of the world is wrong. Build them the shortest possible path to say so.
The Order of Operations
- Verify delivery coordinates against actual GPS delivery events, and build the correction loop before anything else
- Replace flat service-time estimates with measured per-customer distributions from driver timestamps
- Write down the constraints living in dispatchers' heads, and give the list an owner
- Reconcile vehicle capacities and equipment status against reality, weekly
- Only then evaluate solvers, and judge them on plan adherence, not on simulated distance savings
What Good Looks Like
With the data layer fixed, that operator's outcome was unremarkable in the best way: plan adherence rose past 85 percent from roughly half, fuel and overtime dropped by double digits, and, telling for us, the eventual routing engine mattered less than expected. A decent open-source solver on honest data beat the premium engine on bad data by a wide margin, and it wasn't close. If you're budgeting a routing project, put the first 60 to 70 percent of effort into the data pipeline and the feedback loops, and treat the algorithm as the last mile. Ask any prospective vendor how they'll validate your coordinates and measure your service times before the solver ever runs. The good ones have a rehearsed answer, because they've been burned too. It's the part vendors love to demo, and the part that's least likely to be your actual problem.
Have a project that needs this kind of thinking applied to it?