Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty
| dc.contributor.author | Hou, Shixuan | |
| dc.contributor.author | Hu, Chengming | |
| dc.contributor.author | Chowdary, Annabathuni Sandeep | |
| dc.contributor.author | Ghaddar, Bissan | |
| dc.contributor.author | Naoum-Sawaya, Joe | |
| dc.contributor.author | Gao, Jie | |
| dc.contributor.ror | https://ror.org/02jjdwm75 | |
| dc.date.accessioned | 2026-09-02T09:36:40Z | |
| dc.date.issued | 2026-12 | |
| dc.description.abstract | As last-mile delivery demand continues to grow and customer preferences become increasingly diverse, traditional single-mode logistics systems are inadequate in addressing the complexity and variability of modern urban delivery environments. In this paper, we propose an innovative hybrid last-mile delivery system that leverages mobile depots as intermediate hubs, integrating crowd-shipping, and customer self-pickup, to enhance cost-effectiveness. However multiple stakeholders in the system have conflicting interests, ignoring their behavioral uncertainties can impair operational efficiency and increase delivery costs. For example, customers and crowd-shippers may compete for the limited storage capacity of mobile depots, and crowd-shippers may be unwilling to deliver orders that customers decline to pick up. Specifically, the study captures crowd-shippers’ order-acceptance behavior and customers’ self-pickup choice through two binomial logit models, and embeds these behavioral components into a stochastic mixed-integer programming model. This model jointly optimizes mobile depots allocation and order dispatch decisions, aiming to minimize the expected total operational cost. Considering computational complexity of this model, we propose PRiME-DQN, a reinforcement learning framework that incorporates behavior-prioritized resource filtering, action biasing, and model-efficient temperature tuning to accelerate convergence and improve solution quality. A case study on Amazon’s last-mile delivery in Seattle shows that the proposed hybrid model achieves up to 49.5% cost reduction in low-shipper settings and 11.3% in higher-shipper settings compared to a crowd-shipping system, while also consistently surpassing the self-pickup baseline by 3–6%. Extensive scaling experiments show that PRiME-DQN produces high-quality feasible solutions efficiently. For the largest instances, it achieves delivery costs 1.97% lower than Gurobi’s best incumbent solution obtained within the 3,600-second time limit, while reducing the solution time to approximately 1400 seconds. | |
| dc.description.peerreviewed | Yes | |
| dc.description.status | Published | |
| dc.format | application/pdf | |
| dc.identifier.citation | Hou, S., Hu, C., Chowdary, A. S., Ghaddar, B., Naoum-Sawaya, J., & Gao, J. (2026). Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty. Transportation Research Part E: Logistics and Transportation Review, 216, https://doi.org/10.1016/j.tre.2026.105170 | |
| dc.identifier.doi | https://doi.org/10.1016/j.tre.2026.105170 | |
| dc.identifier.issn | 1878-5794 | |
| dc.identifier.officialurl | https://www.sciencedirect.com/science/article/pii/S1366554526005089 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14417/4471 | |
| dc.journal.title | Transportation Research Part E: Logistics and Transportation Review | |
| dc.language.iso | eng | |
| dc.publisher | Elsevier | |
| dc.relation.department | Accounting & Management Control | |
| dc.relation.entity | IE University | |
| dc.relation.school | IE School of Science & Technology | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | |
| dc.rights.accessRights | info:eu-repo/semantics/openAccess | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject.keywords | Transportation and logistics | |
| dc.subject.keywords | Mixed integer programming | |
| dc.subject.keywords | Behavior uncertainty | |
| dc.subject.keywords | Reinforcement learning | |
| dc.subject.ods | ODS 9 - Industria, innovación e infraestructura | |
| dc.subject.unesco | 33 Ciencias Tecnológicas | |
| dc.title | Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty | |
| dc.type | info:eu-repo/semantics/article | |
| dc.version.type | info:eu-repo/semantics/publishedVersion | |
| dc.volume.number | 216 | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 3e8d108e-2dfb-4db4-bc22-f229f807562f | |
| relation.isAuthorOfPublication | 9454bcb1-3635-4138-a1a0-e399b46d1d90 | |
| relation.isAuthorOfPublication.latestForDiscovery | 3e8d108e-2dfb-4db4-bc22-f229f807562f |
Bloque original
1 - 1 de 1
Cargando...
- Nombre:
- 1-s2.0-S1366554526005089-main.pdf
- Tamaño:
- 6.66 MB
- Formato:
- Adobe Portable Document Format
Bloque de licencias
1 - 1 de 1
Cargando...
- Nombre:
- license.txt
- Tamaño:
- 2.89 KB
- Formato:
- Item-specific license agreed to upon submission
- Descripción:
