master
scripts : make the shell scripts cross-platform (#14341)
ggml-cuda: Add generic NVFP4 MMQ kernel (#21074)
ci : switch from pyright to ty (#20826)
hexagon: improved Op queuing, buffer and cache management (#21705)
benches : update models + numbers (#19359)
llama : reorganize source code + improve CMake (#8006)
scripts: add sqlite3 check for compare-commits.sh (#15633)
llama-bench: add `-fitc` and `-fitt` to arguments (#21304)
scripts: update corpus of compare-logprobs (#19326)
Docs: add instructions for adding backends (#14889)
refactor : remove libcurl, use OpenSSL when available (#18828)
server: Add cached_tokens info to oaicompat responses (#19361)
ci : bump ty to 0.0.26 (#21156)
build : pass all warning flags to nvcc via -Xcompiler (#5570)
scripts : update get-hellaswag.sh and get-winogrande.sh (#20542)
scripts : improve get-wikitext-2.sh (#19952)
scripts: corrected encoding when getting chat template (#11866) (#11907)
llama: end-to-end tests (#19802)
support SYCL backend windows build (#5208)
chore : correct typos [no ci] (#20041)
scripts: add function call test script (#21234)
Autoparser - complete refactoring of parser architecture (#18675)
scripts : update sync scripts
sync : ggml
vendor : update cpp-httplib to 0.42.0 (#21781)
convert.py : add python logging instead of print() (#6511)
llama : move end-user examples to tools directory (#13249)