v0.0.3-alpha15
Add safety server to inference (#2449)
Added max messages and max message length settings for inference (#2774)
Introduce model configs to abstract pairings of models and hardware (#2194)
Emit connection retry message only when actually retrying on error (#2804)
adjusted token buffer to pop EOS
fixed inference deploy
inference: allow user change chat title (#2496)