Commit Graph

3681 Commits (f20b066c591fc69bac41757de7d3cfa1e081f688)
 

Author SHA1 Message Date
Guangyao Zhang f20b066c59
[fp8] Disable all_gather intranode. Disable Redundant all_gather fp8 (#6059)
2 months ago
botbw 696fced0d7
[fp8] fix missing fp8_comm flag in mixtral (#6057)
2 months ago
flybird11111 a35a078f08
[doc] update sp doc (#6055)
2 months ago
Hongxin Liu 13946c4448
[fp8] hotfix backward hook (#6053)
2 months ago
botbw c54c4fcd15
[hotfix] moe hybrid parallelism benchmark & follow-up fix (#6048)
3 months ago
Wenxuan Tan 8fd25d6e09
[Feature] Split cross-entropy computation in SP (#5959)
3 months ago
Hongxin Liu b3db1058ec
[release] update version (#6041)
3 months ago
Hanks 5ce6dd75bf
[fp8] disable all_to_all_fp8 in intranode (#6045)
3 months ago
Hongxin Liu 26e553937b
[fp8] fix linear hook (#6046)
3 months ago
Hongxin Liu c3b5caff0e
[fp8] optimize all-gather (#6043)
3 months ago
Tong Li c650a906db
[Hotfix] Remove deprecated install (#6042)
3 months ago
Gao, Ruiyuan e9032fb0b2
[colossalai/checkpoint_io/...] fix bug in load_state_dict_into_model; format error msg (#6020)
3 months ago
Guangyao Zhang e96a0761ea
[FP8] unsqueeze scale to make it compatible with torch.compile (#6040)
3 months ago
Tong Li 0d3a85d04f
add fused norm (#6038)
3 months ago
Tong Li 4a68efb7da
[Colossal-LLaMA] Refactor latest APIs (#6030)
3 months ago
Hongxin Liu cc1b0efc17
[plugin] hotfix zero plugin (#6036)
3 months ago
Wenxuan Tan d383449fc4
[CI] Remove triton version for compatibility bug; update req torch >=2.2 (#6018)
3 months ago
Hongxin Liu 17904cb5bf
Merge pull request #6012 from hpcaitech/feature/fp8_comm
3 months ago
Wang Binluo 4a6f31eb0c
Merge pull request #6033 from wangbluo/fix
3 months ago
pre-commit-ci[bot] 80d24ae519 [pre-commit.ci] auto fixes from pre-commit.com hooks
3 months ago
wangbluo dae39999d7 fix
3 months ago
Wenxuan Tan 7cf9df07bc
[Hotfix] Fix llama fwd replacement bug (#6031)
3 months ago
Wang Binluo 0bf46c54af
Merge pull request #6029 from hpcaitech/flybird11111-patch-1
3 months ago
flybird11111 9e767643dd
Update low_level_zero_plugin.py
3 months ago
pre-commit-ci[bot] 3b0df30362 [pre-commit.ci] auto fixes from pre-commit.com hooks
3 months ago
flybird11111 0bc9a870c0
Update train_dpo.py
3 months ago
Hongxin Liu caab4a307f
Merge branch 'main' into feature/fp8_comm
3 months ago
Wang Binluo afe845ff15
Merge pull request #6024 from wangbluo/fix_merge
3 months ago
pre-commit-ci[bot] a292554179 [pre-commit.ci] auto fixes from pre-commit.com hooks
3 months ago
wangbluo 971b16a74f fix
3 months ago
Wang Binluo d77e66a577
Merge pull request #6023 from wangbluo/fp8_merge
3 months ago
Wang Binluo eea37da6fa
[fp8] Merge feature/fp8_comm to main branch of Colossalai (#6016)
3 months ago
wangbluo 8b8e282441 fix
3 months ago
wangbluo 698c8b9804 fix
3 months ago
wangbluo 6aface9316 fix
3 months ago
wangbluo 193030f696 fix
3 months ago
wangbluo eb5ba40def fix the merge
3 months ago
Tong Li 39e2597426
[ColossalChat] Add PP support (#6001)
3 months ago
Hongxin Liu 0d3b0bd864
[plugin] add cast inputs option for zero (#6003) (#6022)
3 months ago
wangbluo 2d362ac090 fix merge
3 months ago
wangbluo 2e4cbe3a2d fix
3 months ago
wangbluo 2ee6235cfa fix
3 months ago
wangbluo f7acfa1bd5 fix
3 months ago
wangbluo 53823118f2 fix
3 months ago
Edenzzzz dcc44aab8d
[misc] Use dist logger in plugins (#6011)
3 months ago
wangbluo 1f703e0ef4 fix
3 months ago
wangbluo 88b3f0698c fix the merge
3 months ago
wangbluo 2eb36839c6 fix
3 months ago
wangbluo 12b44012d9 fix
3 months ago
wangbluo 0d8e82a024 Merge branch 'fp8_merge' of https://github.com/wangbluo/ColossalAI into fp8_merge
3 months ago