v0.3.4
[Infer] Serving example w/ ray-serve (multiple GPU case) (#4841)
[Inference]ADD Bench Chatglm2 script (#4963)
[Kernels]Updated Triton kernels into 2.1.0 and adding flash-decoding for llama token attention (#4965)
[inference] Add smmoothquant for llama (#4904)