MDPRP: A Q-learning approach for the joint control of beaconing rate and transmission power in VANETs

Vehicular ad-hoc communications rely on periodic broadcast beacons as the basis for most of their safety applications, allowing vehicles to be aware of their surroundings. However, an excessive beaconing load might compromise the proper operation of these crucial applications, especially regarding t...

Descripción completa

Detalles Bibliográficos
Autores: Aznar Poveda, Juan, García Sánchez, Antonio Javier, Egea López, Esteban, García Haro, Juan
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2021
País:España
Institución:Universidad Politécnica de Cartagena(UPCT)
Repositorio:Repositorio Digital UPCT
OAI Identifier:oai:repositorio.upct.es:10317/11432
Acceso en línea:http://hdl.handle.net/10317/11432
https://ieeexplore.ieee.org/document/9319141
Access Level:acceso abierto
Palabra clave:Vehicular ad-hoc networks
Connected vehicles
Vehicle-to-vehicle (V2V) communications
Congestion control
Power control
Rate control
Reinforcement learning
IEEE 802.11p,
SAE J2945/1
Ingeniería Telemática
3325 Tecnología de las Telecomunicaciones
Descripción
Sumario:Vehicular ad-hoc communications rely on periodic broadcast beacons as the basis for most of their safety applications, allowing vehicles to be aware of their surroundings. However, an excessive beaconing load might compromise the proper operation of these crucial applications, especially regarding the exchange of emergency messages. Therefore, congestion control can play an important role. In this article, we propose joint beaconing rate and transmission power control based on policy evaluation. To this end, a Markov Decision Process (MDP) is modeled by making a set of reasonable simplifying assumptions which are resolved using Q-learning techniques. This MDP characterization, denoted as MDPRP (indicating Rate and Power), leverages the trade-off between beaconing rate and transmission power allocation. Moreover, MDPRP operates in a non-cooperative and distributed fashion, without requiring additional information from neighbors, which makes it suitable for use in infrastructureless (ad-hoc) networks. The results obtained reveal that MDPRP not only balances the channel load successfully but also provides positive outcomes in terms of packet delivery ratio. Finally, the robustness of the solution is shown since the algorithm works well even in those cases where none of the assumptions made to derive the MDP model apply.