Decentralized Federated Policy Gradient with Byzantine Fault-Tolerance and Provably Fast Convergence

Philip Jordan, Florian Grötschla, Flint Xiaofeng Fan, Roger Wattenhofer

AAMAS 2024 · Oral

Agents collaborate on reinforcement learning while acting in local environments and without exchanging raw trajectories. This work combines robust aggregation with Byzantine-resilient agreement to remove dependence on a trusted central aggregator. The policy-gradient methods provide convergence guarantees and are evaluated under honest participation and several Byzantine attacks.

Decentralized federated reinforcement learning with potentially faulty or adversarial participants.
Decentralized federated reinforcement learning with potentially faulty or adversarial participants.

← All publications

➤