ColossalAI

Commit Graph

Author	SHA1	Message	Date
Frank Lee	86ac782d7c	[test] added timm models to test model zoo (#3129 ) * [test] added timm models to test model zoo * polish code * polish code * polish code * polish code * polish code	2023-03-14 14:29:18 +08:00
Xuanlei Zhao	30dd13c450	[autochunk] support complete benchmark (#3121 ) * refact memory code * dont log free var memory * add memory align * update chunk target * update setting for new memory * finish test * update tracer * update typo * update test * add unet test * add bench * update bench * update bench * init * support vit * move to cpu * add cpu benchmark	2023-03-13 17:42:37 +08:00
Super Daniel	fff98f06ed	[analyzer] a minimal implementation of static graph analyzer (#2852 ) * [hotfix] meta tensor default device. * [siu] add experimental submodules to main branch. * [siu] * [siu] * [analyzer] init. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [test] add test. * Update symbolic_trace.py * mark skip tests. * try except. * try except. * try except. * s * init * init * fix * skip * skip --------- Co-authored-by: Daniel Shao <superdainiu@MININT-PVARVID.fareast.corp.microsoft.com> Co-authored-by: Daniel Shao <superdainiu@Daniels-Mac.local>	2023-03-10 13:21:05 +08:00
Xuanlei Zhao	10c61de2f7	[autochunk] support vit (#3084 ) support vit for autochunk * support some new ops for vit * fix some bugs * add test for vit	2023-03-10 10:23:26 +08:00
YuliangLiu0306	8e4e8601b7	[DTensor] implement layout converter (#3055 ) * [DTensor] refactor LayoutConverter for DTensor * polish code * polish docstring	2023-03-10 09:53:52 +08:00
Xuanlei Zhao	2ca9728cbb	[autochunk] refactor chunk memory estimation (#2762 ) * refact memory code * dont log free var memory * add memory align * update chunk target * update setting for new memory * finish test * update tracer * update typo * update test	2023-03-08 16:22:30 +08:00
YuliangLiu0306	29386a54e6	[DTensor] refactor CommSpec (#3034 )	2023-03-08 10:45:31 +08:00
YuliangLiu0306	4269196c79	[hotfix] skip auto checkpointing tests (#3029 ) * [hotfix] skip auto checkpointing tests * fix test name issue	2023-03-07 15:50:00 +08:00
YuliangLiu0306	cd2b0eaa8d	[DTensor] refactor sharding spec (#2987 ) * [autoparallel] refactor sharding spec * rename function name	2023-03-07 11:08:11 +08:00
YuliangLiu0306	e414e4092b	[DTensor] implementation of dtensor (#2946 ) * [DTensor] implementation of dtensor * test layout convert * polish	2023-03-01 16:34:58 +08:00
YuliangLiu0306	197d0bf4ed	[autoparallel] apply repeat block to reduce solving time (#2912 )	2023-02-28 11:03:30 +08:00
YuliangLiu0306	819e25d8b1	[hotfix] fix autoparallel compatibility test issues (#2754 )	2023-02-23 17:28:36 +08:00
YuliangLiu0306	0f392d7403	[autoparallel] find repeat blocks (#2854 ) * [autoparallel] find repeat blocks * polish * polish * polish	2023-02-23 17:28:19 +08:00
Boyuan Yao	c7764d3f22	[autoparallel] Patch meta information of `torch.where` (#2822 ) * [autoparallel] patch meta information of torch.where * [autoparallel] pre-commit modified	2023-02-22 10:28:21 +08:00
Boyuan Yao	fcc4097efa	[autoparallel] Patch meta information of `torch.tanh()` and `torch.nn.Dropout` (#2773 ) * [autoparallel] tanh meta information * [autoparallel] remove redundant code * [autoparallel] patch meta information of torch.nn.Dropout	2023-02-22 10:27:59 +08:00
Boyuan Yao	7ea6bc7f69	[autoparallel] Patch tensor related operations meta information (#2789 ) * [autoparallel] tensor related meta information prototype * [autoparallel] tensor related meta information * [autoparallel] tensor related meta information * [autoparallel] tensor related meta information * [autoparallel] tensor related meta information	2023-02-20 17:38:55 +08:00
HELSON	56ddc9ca7a	[hotfix] add correct device for fake_param (#2796 )	2023-02-17 15:29:07 +08:00
Boyuan Yao	a2b43e393d	[autoparallel] Patch meta information of `torch.nn.Embedding` (#2760 ) * [autoparallel] embedding metainfo * [autoparallel] fix function name in test_activation_metainfo * [autoparallel] undo changes in activation metainfo and related tests	2023-02-17 10:39:48 +08:00
YuliangLiu0306	1dc003c169	[autoparallel] distinguish different parallel strategies (#2699 )	2023-02-15 22:28:28 +08:00
YuliangLiu0306	21d6a48f4d	[autoparallel] add shard option (#2696 ) * [autoparallel] add shard option * polish	2023-02-15 13:48:28 +08:00
YuliangLiu0306	cb2c6a2415	[autoparallel] refactor runtime pass (#2644 ) * [autoparallel] refactor runtime pass * add unit test * polish	2023-02-15 10:36:19 +08:00
YuliangLiu0306	0b2a738393	[autoparallel] remove deprecated codes (#2664 )	2023-02-15 09:54:32 +08:00
YuliangLiu0306	7fa6be49d2	[autoparallel] test compatibility for gemini and auto parallel (#2700 )	2023-02-15 09:43:29 +08:00
Boyuan Yao	40c916b192	[autoparallel] Patch meta information of `torch.nn.functional.softmax` and `torch.nn.Softmax` (#2674 ) * [autoparallel] softmax metainfo * [autoparallel] softmax metainfo	2023-02-13 16:09:22 +08:00
HELSON	8213f89fd2	[gemini] add fake_release_chunk for keep-gathered chunk in the inference mode (#2671 )	2023-02-13 14:35:32 +08:00
Boyuan Yao	0385b26ebf	[autoparallel] Patch meta information of `torch.nn.LayerNorm` (#2647 ) * [autoparallel] layernorm metainfo patch * [autoparallel] polish test	2023-02-10 14:29:24 +08:00
YuliangLiu0306	37df666f38	[autoparallel] refactor handlers which reshape input tensors (#2615 ) * [autoparallel] refactor handlers which reshape input tensors * polish	2023-02-08 15:02:49 +08:00
YuliangLiu0306	cb3d1bef62	[autoparallel] adapt autoparallel tests with latest api (#2626 )	2023-02-08 15:02:12 +08:00
Boyuan Yao	90a9fdd91d	[autoparallel] Patch meta information of `torch.matmul` (#2584 ) * [autoparallel] matmul metainfo * [auto_parallel] remove unused print * [tests] skip test_matmul_handler when torch version is lower than 1.12.0	2023-02-08 11:05:31 +08:00
oahzxl	6ba8364881	[autochunk] support diffusion for autochunk (#2621 ) * add alphafold benchmark * renae alphafold test * rename tests * rename diffuser * renme * rename * update transformer * update benchmark * update benchmark * update bench memory * update transformer benchmark * rename * support diffuser * support unet metainfo prop * fix bug and simplify code * update linear and support some op * optimize max region search, support conv * update unet test * support some op * support groupnorm and interpolate * update flow search * add fix dim in node flow * fix utils * rename * support diffusion * update diffuser * update chunk search * optimize imports * import * finish autochunk	2023-02-07 16:32:45 +08:00
oahzxl	c4b15661d7	[autochunk] add benchmark for transformer and alphafold (#2543 )	2023-02-02 15:06:43 +08:00
oahzxl	05671fcb42	[autochunk] support multi outputs chunk search (#2538 ) Support multi outputs chunk search. Previously we only support single output chunk search. It is more flexible and improve performance by a large margin. For transformer, we reduce memory by 40% than previous search strategy. 1. rewrite search strategy to support multi outputs chunk search 2. fix many, many bugs 3. update tests	2023-02-01 13:18:51 +08:00
oahzxl	63199c6687	[autochunk] support transformer (#2526 )	2023-01-31 16:00:06 +08:00
Frank Lee	b55deb0662	[workflow] only report coverage for changed files (#2524 ) * [workflow] only report coverage for changed files * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file * polish file	2023-01-30 21:28:27 +08:00
HELSON	b528eea0f0	[zero] add zero wrappers (#2523 ) * [zero] add zero wrappers * change names * add wrapper functions to init	2023-01-29 17:52:58 +08:00
HELSON	077a5cdde4	[zero] fix gradient clipping in hybrid parallelism (#2521 ) * [zero] fix gradient clipping in hybrid parallelism * [testing] change model name to avoid pytest warning * [hotfix] fix unit testing	2023-01-29 15:09:57 +08:00
HELSON	707b11d4a0	[gemini] update ddp strict mode (#2518 ) * [zero] add strict ddp mode for chunk init * [gemini] update gpt example	2023-01-28 14:35:25 +08:00
HELSON	2d1a7dfe5f	[zero] add strict ddp mode (#2508 ) * [zero] add strict ddp mode * [polish] add comments for strict ddp mode * [zero] fix test error	2023-01-20 14:04:38 +08:00
oahzxl	c04f183237	[autochunk] support parsing blocks (#2506 )	2023-01-20 11:18:17 +08:00
oahzxl	72341e65f4	[auto-chunk] support extramsa (#3 ) (#2504 )	2023-01-20 10:13:03 +08:00
oahzxl	ecccc91f21	[autochunk] support autochunk on evoformer (#2497 )	2023-01-19 11:41:00 +08:00
HELSON	d565a24849	[zero] add unit testings for hybrid parallelism (#2486 )	2023-01-18 10:36:10 +08:00
oahzxl	4953b4ace1	[autochunk] support evoformer tracer (#2485 ) support full evoformer tracer, which is a main module of alphafold. previously we just support a simplifed version of it. 1. support some evoformer's op in fx 2. support evoformer test 3. add repos for test code	2023-01-16 19:25:05 +08:00
YuliangLiu0306	67e1912b59	[autoparallel] support origin activation ckpt on autoprallel system (#2468 )	2023-01-16 16:25:13 +08:00
HELSON	21c88220ce	[zero] add unit test for low-level zero init (#2474 )	2023-01-15 10:42:01 +08:00
HELSON	a5dc4253c6	[zero] polish low level optimizer (#2473 )	2023-01-13 14:56:17 +08:00
Jiarui Fang	867c8c2d3a	[zero] low level optim supports ProcessGroup (#2464 )	2023-01-13 10:05:58 +08:00
YuliangLiu0306	8221fd7485	[autoparallel] update binary elementwise handler (#2451 ) * [autoparallel] update binary elementwise handler * polish	2023-01-12 09:35:10 +08:00
HELSON	5521af7877	[zero] fix state_dict and load_state_dict for ddp ignored parameters (#2443 ) * [ddp] add is_ddp_ignored [ddp] rename to is_ddp_ignored * [zero] fix state_dict and load_state_dict * fix bugs * [zero] update unit test for ZeroDDP	2023-01-11 14:55:41 +08:00
YuliangLiu0306	41429b9b28	[autoparallel] add shard option (#2423 )	2023-01-11 13:40:33 +08:00
HELSON	bb4e9a311a	[zero] add inference mode and its unit test (#2418 )	2023-01-11 10:07:37 +08:00
oahzxl	61fdd3464a	update doc	2023-01-10 12:29:09 +08:00
oahzxl	36ab2cb783	change import	2023-01-10 12:20:40 +08:00
oahzxl	7ab2db206f	adapt new fx	2023-01-10 11:56:00 +08:00
oahzxl	e532679c95	Merge branch 'main' of https://github.com/oahzxl/ColossalAI into chunk	2023-01-10 11:29:01 +08:00
oahzxl	c1492e5013	add test in import	2023-01-10 11:20:28 +08:00
HELSON	ea13a201bb	[polish] polish code for get_static_torch_model (#2405 ) * [gemini] polish code * [testing] remove code * [gemini] make more robust	2023-01-09 17:41:38 +08:00
oahzxl	212b5b1b5f	add comments	2023-01-09 16:29:33 +08:00
oahzxl	aafc3516a5	add available	2023-01-09 15:32:19 +08:00
oahzxl	d5c4f0bf95	code style	2023-01-09 15:22:09 +08:00
oahzxl	d106b271f8	add chunk search test	2023-01-09 15:19:08 +08:00
oahzxl	a005965d2d	update codegen test	2023-01-09 14:57:47 +08:00
oahzxl	3abbaf8bc6	update codegen test	2023-01-09 14:53:04 +08:00
oahzxl	74b81395a2	update codegen test	2023-01-09 14:26:22 +08:00
oahzxl	18a51c87fe	rename test	2023-01-09 14:20:54 +08:00
oahzxl	cb68ee864a	set benchmark	2023-01-09 14:20:41 +08:00
Jiarui Fang	4e96039649	[device] find best logical mesh	2023-01-07 14:04:30 +08:00
Frank Lee	40d376c566	[setup] support pre-build and jit-build of cuda kernels (#2374 ) * [setup] support pre-build and jit-build of cuda kernels * polish code * polish code * polish code * polish code * polish code * polish code	2023-01-06 20:50:26 +08:00
oahzxl	a6cdbf9161	seperate trace flow	2023-01-06 17:24:23 +08:00
oahzxl	da4076846d	rename	2023-01-06 17:09:37 +08:00
oahzxl	fd87d78a28	rename ambiguous variable	2023-01-06 14:28:04 +08:00
oahzxl	8a634af2f5	close mem and code print	2023-01-06 14:19:45 +08:00
oahzxl	1a6d2a740b	take apart chunk code gen	2023-01-06 14:14:45 +08:00
HELSON	48d33b1b17	[gemini] add get static torch model (#2356 )	2023-01-06 13:41:19 +08:00
oahzxl	d1f0773182	rename	2023-01-06 11:48:33 +08:00
oahzxl	06a5355d98	update test	2023-01-06 11:44:01 +08:00
oahzxl	efb1c64c30	restruct dir	2023-01-06 11:39:26 +08:00
YuliangLiu0306	b5a3a4a65f	[device] find best logical mesh	2023-01-05 17:21:29 +08:00
YuliangLiu0306	9c9246c0d9	[device] alpha beta profiler (#2311 ) * [device] alpha beta profiler * add usage * fix variable name	2023-01-05 16:39:55 +08:00
Jiarui Fang	db6eea3583	[builder] reconfig op_builder for pypi install (#2314 )	2023-01-04 16:32:32 +08:00
HELSON	5d3a2be3af	[amp] add gradient clipping for unit tests (#2283 ) * [amp] add gradient clipping in unit tests * fix bugs	2023-01-04 11:59:56 +08:00
zbian	e94c79f15b	improved allgather & reducescatter for 3d	2023-01-03 17:46:08 +08:00
YuliangLiu0306	fb87322773	[autoparallel] fix spelling error (#2270 )	2023-01-03 16:13:00 +08:00
YuliangLiu0306	8897b8f753	[autoparallel] autoparallel initialize (#2238 )	2022-12-31 01:02:14 +08:00
YuliangLiu0306	3b1b91eaf4	[autoparallel] record parameter attribute in colotracer (#2217 ) * [autoparallel] record parameter attribute in collotracer * [autoparallel] fix construct_meta_info bug	2022-12-28 19:29:08 +08:00
Boyuan Yao	24246f7aa5	[autoparallel] Attach input, buffer and output tensor to MetaInfo class (#2162 ) * [fx] metainfo class for auto parallel * [fx] add unit test for linear metainfo * [fx] fix bwd param for linear * [fx] modify unit test * [fx] modify unit test * [fx] modify import * [fx] modify import * [fx] modify import * [fx] move meta profiler to auto parallel * [fx] add conv metainfo class * [fx] restore profiler * [fx] restore meta profiler * [autoparallel] modify unit test * [fx] modify unit test * [autoparallel] add batchnorm metainfo class * [autoparallel] fix batchnorm unit test function declaration * [fx] restore profiler * [fx] add relu metainfo class * [fx] restore profiler * [autoparallel] modify metainfo input * [autoparallel] add pooling metainfo * [autoparallel] add F.linear metainfo generator * [autoparallel] add binary elementwise metainfo * [fx] recover profiler * [autoparallel] fix forward memory calculation * [autoparallel] modify constants.py * [autoparallel] remove redundant print * [autoparallel] add F.conv metainfo * [autoparallel] linear fix * [autoparallel] memory estimation for communication actions * [autoparallel] fix docstring * [autoparallel] fix variables name * [autoparallel] attach tensor to metainfo class * [autoparallel] fix dangerous try except * [autoparallel] attach memory cost to shape consistency node * [autoparallel] attach shape consistency node's metainfo to the node * [autoparallel] remove todo in shape consistency memory estimation * [autoparallel] fix the annotation	2022-12-28 13:37:40 +08:00
YuliangLiu0306	78509124d3	[autoparallel] update getitem handler (#2207 )	2022-12-27 19:58:32 +08:00
YuliangLiu0306	4851f2d607	[autoparallel] update_getattr_handler (#2193 )	2022-12-26 21:57:39 +08:00
YuliangLiu0306	f10ce01e31	[autoparallel] add gpt2 performance test code (#2194 )	2022-12-26 21:56:58 +08:00
HELSON	a3100bd50d	[testing] add beit model for unit testings (#2196 ) * [testing] add beit model * [beit] fix bugs * [beit] fix bugs * [testing] fix bugs	2022-12-26 17:35:36 +08:00
HELSON	2458659919	[zero] fix error for BEiT models (#2169 ) * [zero] fix error for BEiT models * [ColoParameter] add unpack operation for tuple arguments * fix bugs * fix chunkv2 unit testing * add assertion for gradient state	2022-12-26 15:03:54 +08:00
Jiarui Fang	355ffb386e	[builder] unified cpu_optim fused_optim inferface (#2190 )	2022-12-23 20:57:41 +08:00
Jiarui Fang	9587b080ba	[builder] use runtime builder for fused_optim (#2189 )	2022-12-23 17:07:03 +08:00
Jiarui Fang	bc0e271e71	[buider] use builder() for cpu adam and fused optim in setup.py (#2187 )	2022-12-23 16:05:13 +08:00
Jiarui Fang	d42afd30f8	[builder] runtime adam and fused_optim builder (#2184 )	2022-12-23 14:14:21 +08:00
YuliangLiu0306	550f8f8905	[autoparallel] integrate_gpt_related_tests (#2134 ) * [autoparallel] integrate_gpt_related_tests * polish code * polish code * add GPT2Model into runtime test	2022-12-23 12:36:59 +08:00
Jiarui Fang	27327a4c90	[example] add palm pytorch version (#2172 )	2022-12-22 10:15:34 +08:00
Jiarui Fang	b87496a66b	[hotfix] fix auto policy of test_sharded_optim_v2 (#2157 )	2022-12-20 23:03:18 +08:00
YuliangLiu0306	16335cb537	[hotfix] fix aten default bug (#2158 )	2022-12-20 22:40:46 +08:00
Jiarui Fang	2827f41898	[Gemini] GeminiDPP convert to PyTorch Module. (#2151 )	2022-12-20 10:19:36 +08:00

1 2 3 4 5 ...

770 Commits (b29e1f07224298aea35aab7ee83284beac28e0d8)