Commit Graph

3877 Commits (feature/zerobubble)
 

Author SHA1 Message Date
YeAnbang 8a3ff4f315 fix style
4 months ago
zhurunhua ad35a987d3
[Feature] Add a switch to control whether the model checkpoint needs to be saved after each epoch ends (#5941)
4 months ago
Edenzzzz 2069472e96
[Hotfix] Fix ZeRO typo #5936
4 months ago
Hongxin Liu 5fd0592767
[fp8] support all-gather flat tensor (#5932)
4 months ago
Gao, Ruiyuan 5fb958cc83
[FIX BUG] convert env param to int in (#5934)
4 months ago
Insu Jang a521ffc9f8
Add n_fused as an input from native_module (#5894)
4 months ago
YeAnbang 9688e19b32 remove real data path
4 months ago
YeAnbang b0e15d563e remove real data path
4 months ago
YeAnbang 12fe8b5858 refactor evaluation
4 months ago
YeAnbang c5f582f666 fix test data
4 months ago
zhurunhua 4ec17a7cdf
[FIX BUG] UnboundLocalError: cannot access local variable 'default_conversation' where it is not associated with a value (#5931)
4 months ago
YeAnbang 150505cbb8 Merge branch 'kto' of https://github.com/hpcaitech/ColossalAI into kto
4 months ago
YeAnbang d49550fb49 refactor tokenization
4 months ago
Tong Li d08c99be0d
Merge branch 'main' into kto
4 months ago
Tong Li f585d4e38e
[ColossalChat] Hotfix for ColossalChat (#5910)
4 months ago
Edenzzzz 8cc8f645cd
[Examples] Add lazy init to OPT and GPT examples (#5924)
4 months ago
YeAnbang 544b7a38a1 fix style, add kto data sample
4 months ago
Guangyao Zhang 62661cde22
Merge pull request #5921 from BurkeHulk/fp8_fix
4 months ago
YeAnbang 845ea7214e Merge branch 'main' of https://github.com/hpcaitech/ColossalAI into kto
4 months ago
YeAnbang 09d5ffca1a add kto
4 months ago
Hongxin Liu e86127925a
[plugin] support all-gather overlap for hybrid parallel (#5919)
4 months ago
GuangyaoZhang 5b969fd831 fix shardformer fp8 communication training degradation
4 months ago
Guangyao Zhang d0bdb51f48
Merge pull request #5899 from BurkeHulk/SP_fp8
4 months ago
Hongxin Liu 73494de577
[release] update version (#5912)
4 months ago
GuangyaoZhang 6a20f07b80 remove all to all
4 months ago
GuangyaoZhang 5a310b9ee1 fix rebase
4 months ago
GuangyaoZhang 457a0de79f shardformer fp8
4 months ago
Hongxin Liu 27a72f0de1 [misc] support torch2.3 (#5893)
4 months ago
アマデウス 530283dba0 fix object_to_tensor usage when torch>=2.3.0 (#5820)
4 months ago
Guangyao Zhang 2e28c793ce [compatibility] support torch 2.2 (#5875)
4 months ago
Hanks 9470701110
Merge pull request #5885 from BurkeHulk/feature/fp8_comm
4 months ago
YeAnbang d8bf7e09a2
Merge pull request #5901 from hpcaitech/colossalchat
4 months ago
Guangyao Zhang 1c961b20f3
[ShardFormer] fix qwen2 sp (#5903)
4 months ago
Stephan Kö 45c49dde96
[Auto Parallel]: Speed up intra-op plan generation by 44% (#5446)
4 months ago
YeAnbang b3594d4d68 fix orpo cross entropy loss
4 months ago
pre-commit-ci[bot] 51f916b11d [pre-commit.ci] auto fixes from pre-commit.com hooks
5 months ago
BurkeHulk 1f1b856354 Merge remote-tracking branch 'origin/feature/fp8_comm' into feature/fp8_comm
5 months ago
BurkeHulk 66018749f3 add fp8_communication flag in the script
5 months ago
BurkeHulk e88190184a support fp8 communication in pipeline parallelism
5 months ago
BurkeHulk 1e1959467e fix scaling algorithm in FP8 casting
5 months ago
Hongxin Liu c068ef0fa0
[zero] support all-gather overlap (#5898)
5 months ago
YeAnbang 115c4cc5a4 hotfix citation
5 months ago
YeAnbang e7a8634636 fix eval
5 months ago
YeAnbang dd9e1cdafe
Merge pull request #5850 from hpcaitech/rlhf_SimPO
5 months ago
pre-commit-ci[bot] 8a9721bafe [pre-commit.ci] auto fixes from pre-commit.com hooks
5 months ago
YeAnbang 33f15203d3 Merge branch 'main' of https://github.com/hpcaitech/ColossalAI into rlhf_SimPO
5 months ago
YeAnbang f6ef5c3609 fix style
5 months ago
YeAnbang d888c3787c add benchmark for sft, dpo, simpo, orpo. Add benchmarking result. Support lora with gradient checkpoint
5 months ago
GuangyaoZhang dbfa7d39fc fix typo
5 months ago
Guangyao Zhang 669849d74b
[ShardFormer] Add Ulysses Sequence Parallelism support for Command-R, Qwen2 and ChatGLM (#5897)
5 months ago