ColossalAI

Commit Graph

Author	SHA1	Message	Date
Jiarui Fang	a0e5971692	[Tensor] test model check results for a simple net (#887 )	2022-04-27 12:00:18 +08:00
Jiarui Fang	72cdc06875	[Tensor] make ColoTensor more robust for getattr (#886 ) * [Tensor] make ColoTensor more robust for getattr * polish * polish	2022-04-27 10:57:49 +08:00
Ziyue Jiang	9bc5a77c31	[tensor] wrap function in the torch_tensor to ColoTensor (#881 )	2022-04-26 20:13:56 +08:00
ver217	4df6471f5d	fix import error (#880 )	2022-04-26 19:28:40 +08:00
Jiarui Fang	7f76517a85	[Tensor] make a simple net works with 1D row TP (#879 )	2022-04-26 18:11:47 +08:00
ver217	c4d903e64a	[gemini] accelerate adjust_layout() (#878 ) * add lru cache * polish code * update unit test * fix sharded optim	2022-04-26 18:08:31 +08:00
Jiarui Fang	909211453b	[Tensor] Add some attributes to ColoTensor (#877 ) * [Tensor] add some function to ColoTensor * torch.allclose * rm torch.add	2022-04-26 15:10:47 +08:00
HELSON	425b4a96b8	[gemini] polish stateful_tensor_mgr (#876 )	2022-04-26 15:05:03 +08:00
Jiarui Fang	e43f83aa5c	[Tensor] get named parameters for model using ColoTensors (#874 )	2022-04-26 14:08:01 +08:00
LuGY	2883040286	[example] change qkv processing (#870 )	2022-04-26 13:33:27 +08:00
Jiarui Fang	96211c2cc8	[tensor] customized op returns ColoTensor (#875 ) * [tensor] customized op returns ColoTensor * polish * polish code	2022-04-26 13:23:59 +08:00
Ziyue Jiang	26d4ab8b03	[Tensor] Add function to spec and update linear 1Drow and unit tests (#869 )	2022-04-26 10:15:26 +08:00
Frank Lee	11f54c7b6b	[doc] improved docstring and assertion messages for the engine module (#871 )	2022-04-26 10:00:18 +08:00
Frank Lee	1c34382678	[doc] improved assertion messages in trainer (#873 )	2022-04-26 10:00:12 +08:00
Frank Lee	7a64fae33a	[doc] improved error messages in initialize (#872 )	2022-04-26 10:00:03 +08:00
Jiarui Fang	1190b2c4a4	[tensor] add cross_entrophy_loss (#868 )	2022-04-25 16:01:52 +08:00
HELSON	3107817172	[gemini] add stateful tensor container (#867 )	2022-04-25 14:58:16 +08:00
Jiarui Fang	d01d3b8cb0	colo init context add device attr. (#866 )	2022-04-25 14:24:26 +08:00
Frank Lee	2238758c2e	[usability] improved error messages in the context module (#856 )	2022-04-25 13:42:31 +08:00
Frank Lee	9fdebadd69	[doc] improved docstring in the amp module (#857 )	2022-04-25 13:42:17 +08:00
Frank Lee	b862d89d00	[doc] improved docstring in the logging module (#861 )	2022-04-25 13:42:00 +08:00
Frank Lee	8004c8e938	[doc] improved docstring in the communication module (#863 )	2022-04-25 13:41:43 +08:00
Jiarui Fang	8af5f7423d	[tensor] an initial dea of tensor spec (#865 ) * a initial dea of tensor spec * polish * polish	2022-04-25 13:33:52 +08:00
Jiarui Fang	126ba573a8	[Tensor] add layer norm Op (#852 )	2022-04-25 11:49:20 +08:00
Frank Lee	a82da26f7e	[cli] refactored micro-benchmarking cli and added more metrics (#858 )	2022-04-25 11:48:07 +08:00
Frank Lee	ee222dfbf3	[usability] added assertion message in registry (#864 )	2022-04-25 11:45:15 +08:00
HELSON	f0e654558f	[gemini] polish code (#855 )	2022-04-25 10:40:14 +08:00
Jiarui Fang	29159d9b5b	hotfix tensor unittest bugs (#862 )	2022-04-25 10:06:53 +08:00
Frank Lee	1258af71cc	[ci] cache cuda extension (#860 )	2022-04-25 10:03:47 +08:00
YuliangLiu0306	c6930d8ddf	[pipelinable]use ColoTensor to replace dummy tensor. (#853 )	2022-04-24 18:31:22 +08:00
Ziyue Jiang	bcc8655021	[Tensor ] Add 1Drow weight reshard by spec (#854 )	2022-04-24 18:30:20 +08:00
ver217	d7e0303d1e	[zero] use GeminiMemoryManager when sampling model data (#850 )	2022-04-24 17:17:22 +08:00
ver217	232142f402	[utils] refactor profiler (#837 ) * add model data profiler * add a subclass of torch.profiler.profile * refactor folder structure * remove redundant codes * polish code * use GeminiMemoryManager * fix import path * fix stm profiler ext * polish comments * remove useless file	2022-04-24 17:03:59 +08:00
Jiarui Fang	62f059251b	[Tensor] init a tp network training unittest (#849 )	2022-04-24 16:43:44 +08:00
ver217	0dea140760	[hotfix] add deconstructor for stateful tensor (#848 ) * add deconstructor for stateful tensor * fix colo init context	2022-04-24 15:03:04 +08:00
ver217	0f7ed8c192	fix _post_init_method of zero init ctx (#847 )	2022-04-24 14:16:50 +08:00
Ziyue Jiang	2a0a427e04	[tensor]add assert for colo_tensor 1Drow (#846 )	2022-04-24 14:12:45 +08:00
Ziyue Jiang	05023ecfee	[Tensor] TP Linear 1D row (#843 )	2022-04-24 13:43:12 +08:00
Frank Lee	cf6d1c9284	[CLI] refactored the launch CLI and fixed bugs in multi-node launching (#844 ) * [cli] fixed multi-node job launching * [cli] fixed a bug in version comparison * [cli] support launching with env var * [cli] fixed multi-node job launching * [cli] fixed a bug in version comparison * [cli] support launching with env var * added docstring * [cli] added extra launch arguments * [cli] added default launch rdzv args * [cli] fixed version comparison * [cli] added docstring examples and requierment * polish docstring * polish code * polish code	2022-04-24 13:26:26 +08:00
HELSON	e5ea3fdeef	[gemini] add GeminiMemoryManger (#832 ) * refactor StatefulTensor, tensor utilities * add unitest for GeminiMemoryManager	2022-04-24 13:08:48 +08:00
YuliangLiu0306	35ea6e1023	[pipelinable]use pipelinable context to initialize non-pipeline model (#816 ) * [CLI] add CLI launcher * Revert "[CLI] add CLI launcher" This reverts commit `df7e6506d4`. * [pipeline]add module lazy init feature to support large model initization. * [pipeline]add to_layer_list and partition method to support arbitrary non-pp model * refactor the module structure * polish * [pipelinable]add unit test for pipelinable * polish * polish * Fix CodeFactor issues.	2022-04-24 13:03:12 +08:00
Jiarui Fang	ea0a2ed25f	[hotfix] the bug of numel() in ColoTensor (#845 )	2022-04-24 12:32:10 +08:00
LuGY	c1e8d2001e	modefied the pp build for ckpt adaptation (#803 )	2022-04-24 12:23:16 +08:00
Jiarui Fang	8789850eea	Init Conext supports lazy allocate model memory (#842 )	2022-04-22 18:03:35 +08:00
Jiarui Fang	4575a3298b	[hotfix] ColoTensor pin_memory (#840 )	2022-04-22 17:07:46 +08:00
Frank Lee	9f6f656952	[setup] use env var instead of option for cuda ext (#839 )	2022-04-22 15:44:56 +08:00
Frank Lee	943982d29a	[unittest] refactored unit tests for change in dependency (#838 )	2022-04-22 15:39:07 +08:00
github-actions[bot]	f271f34716	Automated submodule synchronization (#827 ) Co-authored-by: github-actions <github-actions@github.com>	2022-04-22 15:24:58 +08:00
Frank Lee	01e9f834f5	[dependency] removed torchvision (#833 ) * [dependency] removed torchvision * fixed transforms	2022-04-22 15:24:35 +08:00
Jiarui Fang	cb5a4778e1	Revert "[WIP] Applying ColoTensor on TP-1D-row Linear. (#831 )" (#835 ) This reverts commit `ac88de6dfc`.	2022-04-22 14:45:57 +08:00

1 2 3 4 5 ...

545 Commits (a0e59716920b4225fe0c226a9618c46f8ff25f0d) All Branches Search

545 Commits (a0e59716920b4225fe0c226a9618c46f8ff25f0d)

All Branches