b8839
ggml: backend-agnostic tensor parallelism (experimental) (#19378)
CUDA: manage NCCL communicators in context (#21891)
ggml-backend-meta: add multi-segment read support in get_tensor (#22063)
vulkan : cmake integration (#8119)
[SYCL] Fix Q8_0 reorder: garbage on 2nd prompt + crash on full VRAM (#21638)