Hi, I noticed in TorchRL, there's normalize_advantage_exclude_dims which normalize advantage without mixing agent dimention; however, BenchMARL doesn't seem to have adopted it. I also notice an imbalance loss_entropy and loss_objective when training with MAPPO and IPPO.
Is this implementation detail deliberate?
Hi, I noticed in TorchRL, there's
normalize_advantage_exclude_dimswhich normalize advantage without mixing agent dimention; however, BenchMARL doesn't seem to have adopted it. I also notice an imbalanceloss_entropyandloss_objectivewhen training with MAPPO and IPPO.Is this implementation detail deliberate?