ColossalAI

Author	SHA1	Message	Date
BlueRum	7548ca5a54	[chatgpt]Reward Model Training Process update (#3133 ) * add normalize function to value_head in bloom rm * add normalization to value_function in gpt_rm * add normalization to value_head of opt_rm * add Anthropic/hh-rlhf dataset * Update __init__.py * Add LogExpLoss in RM training * Update __init__.py * update rm trainer to use acc as target * update example/train_rm * Update train_rm.sh * code style * Update README.md * Update README.md * add rm test to ci * fix tokenier * fix typo * change batchsize to avoid oom in ci * Update test_ci.sh	2023-03-20 09:59:06 +08:00
BlueRum	0672b5afac	[chatgpt] fix lora support for gpt (#3113 ) * fix gpt-actor * fix gpt-critic * fix opt-critic	2023-03-13 10:37:41 +08:00
hiko2MSP	191daf7411	[chatgpt] type miss of kwargs (#3107 )	2023-03-13 00:00:02 +08:00
Fazzie-Maqianli	02ae80bf9c	[chatgpt]add flag of action mask in critic(#3086 )	2023-03-10 14:40:14 +08:00
Fazzie-Maqianli	c21b11edce	change nn to models (#3032 )	2023-03-07 16:34:22 +08:00