v0.0.3-alpha28
Inference worker inform backend on safety intervention (#2505)
Add allow_data_use to ChatRead model (#3107)
Introduce model configs to abstract pairings of models and hardware (#2194)
Added max messages and max message length settings for inference (#2774)
Use the tool with the highest similarity instead of the first one over 75% (#3100)
adjusted token buffer to pop EOS
Add inference api docs (#3059)
revert postgres port in `full-dev-setup` (#3022)