2
Multimodal Reward Hacking in Reinforcement Learning
This paper investigates the phenomenon of reward hacking in multimodal large language models aligned via reinforcement learning across various tasks and model scales.
Impact
48/100
Current rank score
2.28
Source tier
Tier 1
Category
Research
Firefly links to the original publisher. The summary above is AI-generated for orientation and may differ from the source. The “current rank score” decays over time so newer significant stories surface first.