Inverse Reinforcement Learning for Mean-field Games with Average Reward Criterion
Description
We study the inverse reinforcement learning (IRL) problem for discrete-time, infinite-horizon mean-field games (MFGs) with an average-reward criterion. Unlike the forward setting, where the reward function is known, IRL assumes access only to expert demonstrations that are optimal under some unknown reward function. The objective is to recover a policy that explains the expert behavior under some unknown reward, selected via the maximum causal entropy principle. Our approach is based on the maximum causal entropy principle, which selects the least biased policy among those consistent with the observed demonstrations. We show that the resulting non-convex formulation is equivalently reformulated as a convex optimization problem over occupation measures. Furthermore, we establish that the dual objective is smooth and strongly convex over compact sets and derive a variational representation using a log-partition formulation. Finally, we propose a first-order algorithm for solving the dual problem and recovering an entropy-maximizing equilibrium policy.
Files
bib-14f7b655-2669-4887-b02b-c63865d98096.txt
Files
(168 Bytes)
| Name | Size | Download all |
|---|---|---|
|
md5:0da6c53f9d5a0a3f8d4d7341d0e1ce4f
|
168 Bytes | Preview Download |