/
githubmirror
/
ColossalAI
Обзор
Документация
Войти
/
githubmirror
/
ColossalAI
Код
Запросы
0
Пакеты
0
Релизы
0
Аналитика
Безопасность
v0.3.7
ColossalAI
/
...
/
ColossalChat
/
examples
/
..
community
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
data_preparation_scripts
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
inference
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
ray
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
training_scripts
[devops] remove post commit ci (#5566)
2 года назад
README.md
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
requirements.txt
[ColossalChat] Update RLHF V2 (#5286)
2 года назад
Examples
Table of Contents
Install requirements
Get Start with ColossalRun
Training Configuration
RLHF Training Stage1 - Supervised Instructs Tuning
Step 1: Data Collection
Step 2: Preprocessing
Step 3: Training
RLHF Training Stage2 - Training Reward Model
Step 1: Data Collection
Step 2: Preprocessing
Step 3: Training
Features and Tricks in RM Training
Note on Reward Model Training
RLHF Training Stage3 - Proximal Policy Optimization
Step 1: Data Collection
Step 2: Preprocessing
Step 3: Training
Sample Training Results Using Default Script
Reward
Note on PPO Training
Q1: My reward is negative
Q2: My actor loss is negative
Q3: My reward doesn't go up (decreases)
Q4: Generation is garbage
Alternative Option For RLHF: Direct Preference Optimization
DPO Training Stage1 - Supervised Instructs Tuning
DPO Training Stage2 - DPO Training
Step 1: Data Collection & Preparation
Step 2: Training
DPO Result
Hardware Requirements
Inference example
Attention
README.md