Opponent aware reinforcement learning
| dc.contributor.author | Gallego, Víctor | |
| dc.contributor.author | Naveiro, Roi | |
| dc.contributor.author | Ríos Insua, David | |
| dc.contributor.author | Gómez-Ullate, David | |
| dc.contributor.funder | National Science Foundation (NSF) | |
| dc.contributor.funder | Statistical and Applied Mathematical Sciences Institute (SAMSI) | |
| dc.contributor.funder | Severo Ochoa Excellence Programme | |
| dc.contributor.funder | Ministerio de Ciencia, Innovación y Universidades | |
| dc.contributor.funder | European Union | |
| dc.contributor.funder | Horizon 2020 | |
| dc.contributor.funder | BBVA Foundation (FBBVA) | |
| dc.contributor.ror | https://ror.org/02jjdwm75 | |
| dc.date.accessioned | 2026-10-01T13:40:13Z | |
| dc.date.issued | 2027-01-01 | |
| dc.description.abstract | In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit. We introduce Threatened Markov Decision Processes (TMDPs) as a framework to support an agent against potential opponents in an RL context as well as schemes resulting in novel learning approaches to deal with TMDPs. After introducing our framework and deriving theoretical results, empirical evidence is given via extensive experiments, showing the importance for an RL agent of acknowledging adversarial awareness. | |
| dc.description.peerreviewed | Yes | |
| dc.description.sponsorship | This manuscript was initially prepared while the authors were visiting SAMSI within the Games and Decisions in Risk and Reliability program, under National Science Foundation (NSF) Grant DMS-1638521 (Statistical and Applied Mathematical Sciences Institute). The authors also acknowledge support from the Severo Ochoa Excellence Programme CEX-2023-001347-S, Ministry of Science programs PID2021-124662OB-I00 and PID2025-172412OB-C21, the European Union’s Horizon 2020 Research and Innovation Programme under Grant Agreement No. 101021797 (STARLIGHT), the FBBVA AMALFI project, and the SEDIA Excellent AI project EMOROBCARE. VG acknowledges support from grant FPU16-05034 and PTQ2021-011758, RN acknowledges support from grant FPU15-03636. RN acknowledges the 2026 Leonardo Grant for Scientific Research and Cultural Creation from the BBVA Foundation. The BBVA Foundation accepts no responsibility for the opinions, comments, and content included in the project and/or the results derived from it, which are the sole and absolute responsibility of their authors. DRI is grateful to the AXA-ICMAT Chair in Adversarial Risk Analysis. | |
| dc.description.status | Published | |
| dc.format | application/pdf | |
| dc.identifier.citation | Gallego, V., Naveiro, R., Ríos Insua, D., & Gómez-Ullate, D. (2027). Opponent aware reinforcement learning. European Journal of Operational Research, 336(1), 296–309. https://doi.org/10.1016/j.ejor.2026.08.031 | |
| dc.identifier.doi | https://doi.org/10.1016/j.ejor.2026.08.031 | |
| dc.identifier.issn | 1872-6860 | |
| dc.identifier.officialurl | https://www.sciencedirect.com/science/article/pii/S0377221726007150?getft_integrator=scopus&pes=vor&utm_source=scopus | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14417/4552 | |
| dc.issue.number | 1 | |
| dc.journal.title | European Journal of Operational Research | |
| dc.language.iso | eng | |
| dc.page.final | 309 | |
| dc.page.initial | 296 | |
| dc.page.total | 14 | |
| dc.publisher | Elsevier | |
| dc.relation.department | Applied Mathematics | |
| dc.relation.entity | IE University | |
| dc.relation.projectid | DMS-1638521 | |
| dc.relation.projectid | CEX-2023-001347-S | |
| dc.relation.projectid | PID2021-124662OB-I00 | |
| dc.relation.projectid | PID2025-172412OB-C21 | |
| dc.relation.projectid | 101021797 | |
| dc.relation.projectid | FPU16-05034 | |
| dc.relation.projectid | PTQ2021-011758 | |
| dc.relation.projectid | FPU15-03636 | |
| dc.relation.school | IE School of Science & Technology | |
| dc.rights | Attribution 4.0 International | |
| dc.rights.accessRights | info:eu-repo/semantics/openAccess | |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject.keywords | Artificial intelligence | |
| dc.subject.keywords | Multi-agent reinforcement learning | |
| dc.subject.keywords | Game theory | |
| dc.subject.keywords | Level-k thinking | |
| dc.subject.ods | ODS 8 - Trabajo decente y crecimiento económico | |
| dc.subject.unesco | 53 Ciencias Económicas::5308 Economía general | |
| dc.title | Opponent aware reinforcement learning | |
| dc.type | info:eu-repo/semantics/article | |
| dc.version.type | info:eu-repo/semantics/publishedVersion | |
| dc.volume.number | 336 | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | d0525f43-b84b-4613-9984-4324ddf81556 | |
| relation.isAuthorOfPublication.latestForDiscovery | d0525f43-b84b-4613-9984-4324ddf81556 |
Bloque original
1 - 1 de 1
Cargando...
- Nombre:
- 1-s2.0-S0377221726007150-main.pdf
- Tamaño:
- 2.87 MB
- Formato:
- Adobe Portable Document Format
Bloque de licencias
1 - 1 de 1
Cargando...
- Nombre:
- license.txt
- Tamaño:
- 2.89 KB
- Formato:
- Item-specific license agreed to upon submission
- Descripción:
