main
Add llama training support (#2055)
[Fix] Consume much more gpt memory running eval_rm (#3614)
sft8 training preparation (#2988)
Improve scores of small 1.4B reward model.. (#2329)
Orca chat Dataloader (#3583)
Add megacode3 dataset (#3656)
fixed "TypeError: 'NoneType' object is not iterable" for reward model… (#3587)
Add chunking of pretrain text modeling datasets (#3586)
Changes for orcacode experiment (#3612)
Feature/remove reward instructor (#2289)
Add HFSummaryPairs class & fix AnthropicRLHF parsing (#2362)
Resolve duplication issue with filtered prosocial (#3202)
Choice to add global system-prefix to the assistant during changes (#2053)