v0.0.2-alpha12
Add safety server to inference (#2449)
Improvements to the HF worker container (#2339)
changed basic hf server to support quantization and streaming (#2293)
Introduce model configs to abstract pairings of models and hardware (#2194)
Fixed worker requirements wrt transformers (#2621)
Added docker image for standalone worker (#2300)
Safety level control (#2516)