Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty

dc.contributor.authorHou, Shixuan
dc.contributor.authorHu, Chengming
dc.contributor.authorChowdary, Annabathuni Sandeep
dc.contributor.authorGhaddar, Bissan
dc.contributor.authorNaoum-Sawaya, Joe
dc.contributor.authorGao, Jie
dc.contributor.rorhttps://ror.org/02jjdwm75
dc.date.accessioned2026-09-02T09:36:40Z
dc.date.issued2026-12
dc.description.abstractAs last-mile delivery demand continues to grow and customer preferences become increasingly diverse, traditional single-mode logistics systems are inadequate in addressing the complexity and variability of modern urban delivery environments. In this paper, we propose an innovative hybrid last-mile delivery system that leverages mobile depots as intermediate hubs, integrating crowd-shipping, and customer self-pickup, to enhance cost-effectiveness. However multiple stakeholders in the system have conflicting interests, ignoring their behavioral uncertainties can impair operational efficiency and increase delivery costs. For example, customers and crowd-shippers may compete for the limited storage capacity of mobile depots, and crowd-shippers may be unwilling to deliver orders that customers decline to pick up. Specifically, the study captures crowd-shippers’ order-acceptance behavior and customers’ self-pickup choice through two binomial logit models, and embeds these behavioral components into a stochastic mixed-integer programming model. This model jointly optimizes mobile depots allocation and order dispatch decisions, aiming to minimize the expected total operational cost. Considering computational complexity of this model, we propose PRiME-DQN, a reinforcement learning framework that incorporates behavior-prioritized resource filtering, action biasing, and model-efficient temperature tuning to accelerate convergence and improve solution quality. A case study on Amazon’s last-mile delivery in Seattle shows that the proposed hybrid model achieves up to 49.5% cost reduction in low-shipper settings and 11.3% in higher-shipper settings compared to a crowd-shipping system, while also consistently surpassing the self-pickup baseline by 3–6%. Extensive scaling experiments show that PRiME-DQN produces high-quality feasible solutions efficiently. For the largest instances, it achieves delivery costs 1.97% lower than Gurobi’s best incumbent solution obtained within the 3,600-second time limit, while reducing the solution time to approximately 1400 seconds.
dc.description.peerreviewedYes
dc.description.statusPublished
dc.formatapplication/pdf
dc.identifier.citationHou, S., Hu, C., Chowdary, A. S., Ghaddar, B., Naoum-Sawaya, J., & Gao, J. (2026). Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty. Transportation Research Part E: Logistics and Transportation Review, 216, https://doi.org/10.1016/j.tre.2026.105170
dc.identifier.doihttps://doi.org/10.1016/j.tre.2026.105170
dc.identifier.issn1878-5794
dc.identifier.officialurlhttps://www.sciencedirect.com/science/article/pii/S1366554526005089
dc.identifier.urihttps://hdl.handle.net/20.500.14417/4471
dc.journal.titleTransportation Research Part E: Logistics and Transportation Review
dc.language.isoeng
dc.publisherElsevier
dc.relation.departmentAccounting & Management Control
dc.relation.entityIE University
dc.relation.schoolIE School of Science & Technology
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 International
dc.rights.accessRightsinfo:eu-repo/semantics/openAccess
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subject.keywordsTransportation and logistics
dc.subject.keywordsMixed integer programming
dc.subject.keywordsBehavior uncertainty
dc.subject.keywordsReinforcement learning
dc.subject.odsODS 9 - Industria, innovación e infraestructura
dc.subject.unesco33 Ciencias Tecnológicas
dc.titleDeep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty
dc.typeinfo:eu-repo/semantics/article
dc.version.typeinfo:eu-repo/semantics/publishedVersion
dc.volume.number216
dspace.entity.typePublication
relation.isAuthorOfPublication3e8d108e-2dfb-4db4-bc22-f229f807562f
relation.isAuthorOfPublication9454bcb1-3635-4138-a1a0-e399b46d1d90
relation.isAuthorOfPublication.latestForDiscovery3e8d108e-2dfb-4db4-bc22-f229f807562f

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
1-s2.0-S1366554526005089-main.pdf
Tamaño:
6.66 MB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
2.89 KB
Formato:
Item-specific license agreed to upon submission
Descripción: