main
fix style
fix readme
Add GRPO and Support RLVR for PPO (#6186)
fix ci; remove test cases that failed on 3080 (those with tps), can pass locally
[ColossalChat] Update RLHF V2 (#5286)
[application] Update README (#6196)
[feat[ Support one-behind to reduce bubble time. Add profiling code (#6353)
tested after rebasing, fix importance sampling bug
Add new implementations of RL algorithms (#6383)
all tests passed
fix code evaluation