Opponent aware reinforcement learning

dc.contributor.authorGallego, Víctor
dc.contributor.authorNaveiro, Roi
dc.contributor.authorRíos Insua, David
dc.contributor.authorGómez-Ullate, David
dc.contributor.funderNational Science Foundation (NSF)
dc.contributor.funderStatistical and Applied Mathematical Sciences Institute (SAMSI)
dc.contributor.funderSevero Ochoa Excellence Programme
dc.contributor.funderMinisterio de Ciencia, Innovación y Universidades
dc.contributor.funderEuropean Union
dc.contributor.funderHorizon 2020
dc.contributor.funderBBVA Foundation (FBBVA)
dc.contributor.rorhttps://ror.org/02jjdwm75
dc.date.accessioned2026-10-01T13:40:13Z
dc.date.issued2027-01-01
dc.description.abstractIn certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit. We introduce Threatened Markov Decision Processes (TMDPs) as a framework to support an agent against potential opponents in an RL context as well as schemes resulting in novel learning approaches to deal with TMDPs. After introducing our framework and deriving theoretical results, empirical evidence is given via extensive experiments, showing the importance for an RL agent of acknowledging adversarial awareness.
dc.description.peerreviewedYes
dc.description.sponsorshipThis manuscript was initially prepared while the authors were visiting SAMSI within the Games and Decisions in Risk and Reliability program, under National Science Foundation (NSF) Grant DMS-1638521 (Statistical and Applied Mathematical Sciences Institute). The authors also acknowledge support from the Severo Ochoa Excellence Programme CEX-2023-001347-S, Ministry of Science programs PID2021-124662OB-I00 and PID2025-172412OB-C21, the European Union’s Horizon 2020 Research and Innovation Programme under Grant Agreement No. 101021797 (STARLIGHT), the FBBVA AMALFI project, and the SEDIA Excellent AI project EMOROBCARE. VG acknowledges support from grant FPU16-05034 and PTQ2021-011758, RN acknowledges support from grant FPU15-03636. RN acknowledges the 2026 Leonardo Grant for Scientific Research and Cultural Creation from the BBVA Foundation. The BBVA Foundation accepts no responsibility for the opinions, comments, and content included in the project and/or the results derived from it, which are the sole and absolute responsibility of their authors. DRI is grateful to the AXA-ICMAT Chair in Adversarial Risk Analysis.
dc.description.statusPublished
dc.formatapplication/pdf
dc.identifier.citationGallego, V., Naveiro, R., Ríos Insua, D., & Gómez-Ullate, D. (2027). Opponent aware reinforcement learning. European Journal of Operational Research, 336(1), 296–309. https://doi.org/10.1016/j.ejor.2026.08.031
dc.identifier.doihttps://doi.org/10.1016/j.ejor.2026.08.031
dc.identifier.issn1872-6860
dc.identifier.officialurlhttps://www.sciencedirect.com/science/article/pii/S0377221726007150?getft_integrator=scopus&pes=vor&utm_source=scopus
dc.identifier.urihttps://hdl.handle.net/20.500.14417/4552
dc.issue.number1
dc.journal.titleEuropean Journal of Operational Research
dc.language.isoeng
dc.page.final309
dc.page.initial296
dc.page.total14
dc.publisherElsevier
dc.relation.departmentApplied Mathematics
dc.relation.entityIE University
dc.relation.projectidDMS-1638521
dc.relation.projectidCEX-2023-001347-S
dc.relation.projectidPID2021-124662OB-I00
dc.relation.projectidPID2025-172412OB-C21
dc.relation.projectid101021797
dc.relation.projectidFPU16-05034
dc.relation.projectidPTQ2021-011758
dc.relation.projectidFPU15-03636
dc.relation.schoolIE School of Science & Technology
dc.rightsAttribution 4.0 International
dc.rights.accessRightsinfo:eu-repo/semantics/openAccess
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subject.keywordsArtificial intelligence
dc.subject.keywordsMulti-agent reinforcement learning
dc.subject.keywordsGame theory
dc.subject.keywordsLevel-k thinking
dc.subject.odsODS 8 - Trabajo decente y crecimiento económico
dc.subject.unesco53 Ciencias Económicas::5308 Economía general
dc.titleOpponent aware reinforcement learning
dc.typeinfo:eu-repo/semantics/article
dc.version.typeinfo:eu-repo/semantics/publishedVersion
dc.volume.number336
dspace.entity.typePublication
relation.isAuthorOfPublicationd0525f43-b84b-4613-9984-4324ddf81556
relation.isAuthorOfPublication.latestForDiscoveryd0525f43-b84b-4613-9984-4324ddf81556

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
1-s2.0-S0377221726007150-main.pdf
Tamaño:
2.87 MB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
2.89 KB
Formato:
Item-specific license agreed to upon submission
Descripción: