ColossalAI

Commit Graph

Author	SHA1	Message	Date
genghaozhe	1ec92d29af	remove perf log, unrelated file and so on	2024-05-20 05:23:26 +00:00
genghaozhe	5c6c5d6be3	remove comments	2024-05-20 05:23:12 +00:00
genghaozhe	7416e4943b	fix conflicts to beautify the code	2024-05-20 04:09:51 +00:00
botbw	f5a5287f87	Merge pull request #5731 from botbw/prefetch [gemini] prefetch for auto policy	2024-05-20 12:04:33 +08:00
genghaozhe	d22bf30ca6	implement auto policy prefetch and modify a little origin code.	2024-05-20 04:01:53 +00:00
pre-commit-ci[bot]	f1918e18a5	[pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci	2024-05-20 03:00:07 +00:00
hxwang	a55a9e298b	[gemini] init auto policy prefetch	2024-05-20 02:21:17 +00:00
Haze188	c5ddf17c76	Merge branch 'hpcaitech:feature/prefetch' into feature/prefetch	2024-05-17 18:58:53 +08:00
genghaozhe	06a3a100b3	remove unrelated code	2024-05-17 10:57:49 +00:00
genghaozhe	3d625ca836	add some todo Message	2024-05-17 10:55:28 +00:00
botbw	9690981601	Merge pull request #5722 from botbw/prefetch [gemini] prefetch chunks	2024-05-17 13:46:18 +08:00
botbw	e57812c672	[chore] Update placement_policy.py	2024-05-17 13:42:18 +08:00
genghaozhe	013690a86b	remove set(all_chunks)	2024-05-16 09:57:51 +00:00
hxwang	6efbadba25	[chore] remove debugging info	2024-05-16 16:46:39 +08:00
hxwang	20701d4533	[chore] remove print	2024-05-16 16:45:50 +08:00
hxwang	f45f8a2aa7	[gemini] maxprefetch means maximum work to keep	2024-05-16 16:12:53 +08:00
genghaozhe	fc2248cf99	Merge branch 'prefetch' of github.com:botbw/ColossalAI into feature/prefetch	2024-05-16 08:05:32 +00:00
genghaozhe	5470e5f94e	a commit for fake push test	2024-05-16 08:03:40 +00:00
pre-commit-ci[bot]	6bbe956316	[pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci	2024-05-16 07:26:20 +00:00
hxwang	82b25524ff	Merge branch 'prefetch' of github.com:botbw/ColossalAI into prefetch	2024-05-16 07:25:22 +00:00
genghaozhe	1f6b57099c	Merge branch 'prefetch' of github.com:botbw/ColossalAI into botbw-prefetch	2024-05-16 07:23:40 +00:00
hxwang	2e68eebdfe	[chore] refactor & sync	2024-05-16 07:22:10 +00:00
binmakeswell	2011b1356a	[misc] Update PyTorch version in docs (#5724 ) * [misc] Update PyTorch version in docs * [misc] Update PyTorch version in docs	2024-05-16 13:54:32 +08:00
pre-commit-ci[bot]	5bedea6e10	[pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci	2024-05-16 05:20:01 +00:00
hxwang	4148ceed9f	[gemini] use compute_chunk to find next chunk	2024-05-16 13:17:26 +08:00
hxwang	b2e9745888	[chore] sync	2024-05-16 04:45:06 +00:00
hxwang	6e38eafebe	[gemini] prefetch chunks	2024-05-15 16:51:44 +08:00
Tong Li	913c920ecc	[Colossal-LLaMA] Fix sft issue for llama2 (#5719 ) * fix minor issue * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-05-15 10:52:11 +08:00
Edenzzzz	43995ee436	[Feature] Distributed optimizers: Lamb, Galore, CAME and Adafactor (#5694 ) * [feat] Add distributed lamb; minor fixes in DeviceMesh (#5476) * init: add dist lamb; add debiasing for lamb * dist lamb tester mostly done * all tests passed * add comments * all tests passed. Removed debugging statements * moved setup_distributed inside plugin. Added dist layout caching * organize better --------- Co-authored-by: Edenzzzz <wtan45@wisc.edu> * [hotfix] Improve tester precision by removing ZeRO on vanilla lamb (#5576) Co-authored-by: Edenzzzz <wtan45@wisc.edu> * [optim] add distributed came (#5526) * test CAME under LowLevelZeroOptimizer wrapper * test CAME TP row and col pass * test CAME zero pass * came zero add master and worker param id convert * came zero test pass * came zero test pass * test distributed came passed * reform code, Modify some expressions and add comments * minor fix of test came * minor fix of dist_came and test * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * minor fix of dist_came and test * rebase dist-optim * rebase dist-optim * fix remaining comments * add test dist came using booster api --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * [optim] Distributed Adafactor (#5484) * [feature] solve conflict; update optimizer readme; * [feature] update optimize readme; * [fix] fix testcase; * [feature] Add transformer-bert to testcase;solve a bug related to indivisible shape (induction in use_zero and tp is row parallel); * [feature] Add transformers_bert model zoo in testcase; * [feature] add user documentation to docs/source/feature. * [feature] add API Reference & Sample to optimizer Readme; add state check for bert exam; * [feature] modify user documentation; * [fix] fix readme format issue; * [fix] add zero=0 in testcase; cached augment in dict; * [fix] fix percision issue; * [feature] add distributed rms; * [feature] remove useless comment in testcase; * [fix] Remove useless test; open zero test; remove fp16 test in bert exam; * [feature] Extract distributed rms function; * [feature] add booster + lowlevelzeroPlugin in test; * [feature] add Start_with_booster_API case in md; add Supporting Information in md; * [fix] Also remove state movement in base adafactor; * [feature] extract factor function; * [feature] add LowLevelZeroPlugin test; * [fix] add tp=False and zero=True in logic; * [fix] fix use zero logic; * [feature] add row residue logic in column parallel factor; * [feature] add check optim state func; * [feature] Remove duplicate logic; * [feature] update optim state check func and percision test bug; * [fix] update/fix optim state; Still exist percision issue; * [fix] Add use_zero check in _rms; Add plugin support info in Readme; Add Dist Adafactor init Info; * [feature] removed print & comments in utils; * [feature] uodate Readme; * [feature] add LowLevelZeroPlugin test with Bert model zoo; * [fix] fix logic in _rms; * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [fix] remove comments in testcase; * [feature] add zh-Han Readme; --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * [Feature] refractor dist came; fix percision error; add low level zero test with bert model zoo; (#5676) * [feature] daily update; * [fix] fix dist came; * [feature] refractor dist came; fix percision error; add low level zero test with bert model zoo; * [fix] open rms; fix low level zero test; fix dist came test function name; * [fix] remove redundant test; * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * [Feature] Add Galore (Adam, Adafactor) and distributed GaloreAdamW8bit (#5570) * init: add dist lamb; add debiasing for lamb * dist lamb tester mostly done * all tests passed * add comments * all tests passed. Removed debugging statements * moved setup_distributed inside plugin. Added dist layout caching * organize better * update comments * add initial distributed galore * add initial distributed galore * add galore set param utils; change setup_distributed interface * projected grad precision passed * basic precision tests passed * tests passed; located svd precision issue in fwd-bwd; banned these tests * Plugin DP + TP tests passed * move get_shard_dim to d_tensor * add comments * remove useless files * remove useless files * fix zero typo * improve interface * remove moe changes * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix import * fix deepcopy * update came & adafactor to main * fix param map * fix typo --------- Co-authored-by: Edenzzzz <wtan45@wisc.edu> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> * [Hotfix] Remove one buggy test case from dist_adafactor for now (#5692) Co-authored-by: Edenzzzz <wtan45@wisc.edu> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> --------- Co-authored-by: Edenzzzz <wtan45@wisc.edu> Co-authored-by: chongqichuizi875 <107315010+chongqichuizi875@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: duanjunwen <54985467+duanjunwen@users.noreply.github.com> Co-authored-by: Hongxin Liu <lhx0217@gmail.com>	2024-05-14 13:52:45 +08:00
hugo-syn	393c8f5b7f	[hotfix] fix inference typo (#5438 )	2024-05-13 21:06:44 +08:00
Edenzzzz	785cd9a9c9	[misc] Update PyTorch version in docs (#5711 ) Co-authored-by: Edenzzzz <wtan45@wisc.edu>	2024-05-13 12:02:52 +08:00
Wang Binluo	537f6a3855	[Shardformer]fix the num_heads assert for llama model and qwen model (#5704 ) * fix the num_heads assert * fix the transformers import * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix the import --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-05-10 15:33:39 +08:00
Wang Binluo	a3cc68ca93	[Shardformer] Support the Qwen2 model (#5699 ) * feat: support qwen2 model * fix: modify model config and add Qwen2RMSNorm * fix qwen2 model conflicts * test: add qwen2 shard test * to: add qwen2 auto policy * support qwen model * fix the conflicts * add try catch * add transformers version for qwen2 * add the ColoAttention for the qwen2 model * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * add the unit test version check * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix the test input bug * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fix the version check * fix the version check --------- Co-authored-by: Wenhao Chen <cwher@outlook.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-05-09 20:04:25 +08:00
flybird11111	d4c5ef441e	[gemini]remove registered gradients hooks (#5696 ) * fix gemini fix gemini * fix fix	2024-05-09 10:29:49 +08:00
Wang Binluo	22297789ab	Merge pull request #5684 from wangbluo/parallel_output [Shardformer] Add Parallel output for shardformer models	2024-05-07 22:59:42 -05:00
wangbluo	4e50cce26b	fix the mistral model	2024-05-07 09:17:56 +00:00
wangbluo	a8408b4d31	remove comment code	2024-05-07 07:08:56 +00:00
pre-commit-ci[bot]	ca56b93d83	[pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci	2024-05-07 07:07:09 +00:00
wangbluo	108ddfb795	add parallel_output for the opt model	2024-05-07 07:05:53 +00:00
pre-commit-ci[bot]	88f057ce7c	[pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci	2024-05-07 07:03:47 +00:00
Edenzzzz	58954b2986	[misc] Add an existing issue checkbox in bug report (#5691 ) Co-authored-by: Wenxuan(Eden) Tan <wtan45@wisc.edu>	2024-05-07 12:18:50 +08:00
flybird11111	77ec773388	[zero]remove registered gradients hooks (#5687 ) * remove registered hooks fix fix fix zero fix fix fix fix fix zero fix zero fix fix fix * fix fix fix	2024-05-07 12:01:38 +08:00
Edenzzzz	c25f83c85f	fix missing pad token (#5690 ) Co-authored-by: Edenzzzz <wtan45@wisc.edu>	2024-05-06 18:17:26 +08:00
wangbluo	2632916329	remove useless code	2024-05-01 09:23:43 +00:00
wangbluo	9efc79ef24	add parallel output for mistral model	2024-04-30 08:10:20 +00:00
Wang Binluo	d3f34ee8cc	[Shardformer] add assert for num of attention heads divisible by tp_size (#5670 ) * add assert for num of attention heads divisible by tp_size * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-04-29 18:47:47 +08:00
flybird11111	6af6d6fc9f	[shardformer] support bias_gelu_jit_fused for models (#5647 ) * support gelu_bias_fused for gpt2 * support gelu_bias_fused for gpt2 fix fix fix * fix fix * fix	2024-04-29 15:33:51 +08:00
Hongxin Liu	7f8b16635b	[misc] refactor launch API and tensor constructor (#5666 ) * [misc] remove config arg from initialize * [misc] remove old tensor contrusctor * [plugin] add npu support for ddp * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * [devops] fix doc test ci * [test] fix test launch * [doc] update launch doc --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-04-29 10:40:11 +08:00
linsj20	91fa553775	[Feature] qlora support (#5586 ) * [feature] qlora support * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * qlora follow commit * migrate qutization folder to colossalai/ * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * minor fixes --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2024-04-28 10:51:27 +08:00
flybird11111	8954a0c2e2	[LowLevelZero] low level zero support lora (#5153 ) * low level zero support lora low level zero support lora * add checkpoint test * add checkpoint test * fix * fix * fix * fix fix fix fix * fix * fix fix fix fix fix fix fix * fix * fix fix fix fix fix fix fix * fix * test ci * git # This is a combination of 3 commits. Update low_level_zero_plugin.py Update low_level_zero_plugin.py fix fix fix * fix naming fix naming fix naming fix	2024-04-28 10:51:27 +08:00

1 2 3 4 5 ...

3126 Commits (1ec92d29af16fcfc1b641e61eded877c5680cc47) All Branches Search

3126 Commits (1ec92d29af16fcfc1b641e61eded877c5680cc47)

All Branches