v0.0.3-alpha17
Add safety server to inference (#2449)
Emit connection retry message only when actually retrying on error (#2804)
changed basic hf server to support quantization and streaming (#2293)
Introduce model configs to abstract pairings of models and hardware (#2194)
Fixed worker requirements wrt transformers (#2621)
Improvements to the HF worker container (#2339)
Added docker image for standalone worker (#2300)
Safety level control (#2516)