llama.cpp released version b11430, which includes updates to matmul and flash-atten scalability for hexagon. The changes aim to improve performance in multi-core and multi-device scenarios.
What happened
llama.cpp has released version b11430, which includes updates to matmul and flash-atten scalability for hexagon. The updates aim to improve performance in multi-core and multi-device scenarios. The release includes head-parallel flash_attn partitioning for row-split multicore, better work splitting in multi-device scenarios, and improvements to matmul solver and gating based on model sweeps. These changes are aimed at enhancing the efficiency and scalability of the hexagon implementation.
What to weigh
- The updates are aimed at improving performance in multi-core and multi-device scenarios.
- The release includes head-parallel flash_attn partitioning for row-split multicore.
Source: llama.cpp
