ColossalAI

Commit Graph

Author	SHA1	Message	Date
Jiarui Fang	126ba573a8	[Tensor] add layer norm Op (#852 )	2022-04-25 11:49:20 +08:00
Frank Lee	a82da26f7e	[cli] refactored micro-benchmarking cli and added more metrics (#858 )	2022-04-25 11:48:07 +08:00
Frank Lee	ee222dfbf3	[usability] added assertion message in registry (#864 )	2022-04-25 11:45:15 +08:00
HELSON	f0e654558f	[gemini] polish code (#855 )	2022-04-25 10:40:14 +08:00
Jiarui Fang	29159d9b5b	hotfix tensor unittest bugs (#862 )	2022-04-25 10:06:53 +08:00
Frank Lee	1258af71cc	[ci] cache cuda extension (#860 )	2022-04-25 10:03:47 +08:00
YuliangLiu0306	c6930d8ddf	[pipelinable]use ColoTensor to replace dummy tensor. (#853 )	2022-04-24 18:31:22 +08:00
Ziyue Jiang	bcc8655021	[Tensor ] Add 1Drow weight reshard by spec (#854 )	2022-04-24 18:30:20 +08:00
ver217	d7e0303d1e	[zero] use GeminiMemoryManager when sampling model data (#850 )	2022-04-24 17:17:22 +08:00
ver217	232142f402	[utils] refactor profiler (#837 ) * add model data profiler * add a subclass of torch.profiler.profile * refactor folder structure * remove redundant codes * polish code * use GeminiMemoryManager * fix import path * fix stm profiler ext * polish comments * remove useless file	2022-04-24 17:03:59 +08:00
Jiarui Fang	62f059251b	[Tensor] init a tp network training unittest (#849 )	2022-04-24 16:43:44 +08:00
ver217	0dea140760	[hotfix] add deconstructor for stateful tensor (#848 ) * add deconstructor for stateful tensor * fix colo init context	2022-04-24 15:03:04 +08:00
ver217	0f7ed8c192	fix _post_init_method of zero init ctx (#847 )	2022-04-24 14:16:50 +08:00
Ziyue Jiang	2a0a427e04	[tensor]add assert for colo_tensor 1Drow (#846 )	2022-04-24 14:12:45 +08:00
Ziyue Jiang	05023ecfee	[Tensor] TP Linear 1D row (#843 )	2022-04-24 13:43:12 +08:00
Frank Lee	cf6d1c9284	[CLI] refactored the launch CLI and fixed bugs in multi-node launching (#844 ) * [cli] fixed multi-node job launching * [cli] fixed a bug in version comparison * [cli] support launching with env var * [cli] fixed multi-node job launching * [cli] fixed a bug in version comparison * [cli] support launching with env var * added docstring * [cli] added extra launch arguments * [cli] added default launch rdzv args * [cli] fixed version comparison * [cli] added docstring examples and requierment * polish docstring * polish code * polish code	2022-04-24 13:26:26 +08:00
HELSON	e5ea3fdeef	[gemini] add GeminiMemoryManger (#832 ) * refactor StatefulTensor, tensor utilities * add unitest for GeminiMemoryManager	2022-04-24 13:08:48 +08:00
YuliangLiu0306	35ea6e1023	[pipelinable]use pipelinable context to initialize non-pipeline model (#816 ) * [CLI] add CLI launcher * Revert "[CLI] add CLI launcher" This reverts commit `df7e6506d4`. * [pipeline]add module lazy init feature to support large model initization. * [pipeline]add to_layer_list and partition method to support arbitrary non-pp model * refactor the module structure * polish * [pipelinable]add unit test for pipelinable * polish * polish * Fix CodeFactor issues.	2022-04-24 13:03:12 +08:00
Jiarui Fang	ea0a2ed25f	[hotfix] the bug of numel() in ColoTensor (#845 )	2022-04-24 12:32:10 +08:00
LuGY	c1e8d2001e	modefied the pp build for ckpt adaptation (#803 )	2022-04-24 12:23:16 +08:00
Jiarui Fang	8789850eea	Init Conext supports lazy allocate model memory (#842 )	2022-04-22 18:03:35 +08:00
Jiarui Fang	4575a3298b	[hotfix] ColoTensor pin_memory (#840 )	2022-04-22 17:07:46 +08:00
Frank Lee	9f6f656952	[setup] use env var instead of option for cuda ext (#839 )	2022-04-22 15:44:56 +08:00
Frank Lee	943982d29a	[unittest] refactored unit tests for change in dependency (#838 )	2022-04-22 15:39:07 +08:00
github-actions[bot]	f271f34716	Automated submodule synchronization (#827 ) Co-authored-by: github-actions <github-actions@github.com>	2022-04-22 15:24:58 +08:00
Frank Lee	01e9f834f5	[dependency] removed torchvision (#833 ) * [dependency] removed torchvision * fixed transforms	2022-04-22 15:24:35 +08:00
Jiarui Fang	cb5a4778e1	Revert "[WIP] Applying ColoTensor on TP-1D-row Linear. (#831 )" (#835 ) This reverts commit `ac88de6dfc`.	2022-04-22 14:45:57 +08:00
Frank Lee	5e00e6cf23	[setup] allow installation with python 3.6 (#834 )	2022-04-22 14:17:51 +08:00
Jiarui Fang	ac88de6dfc	[WIP] Applying ColoTensor on TP-1D-row Linear. (#831 ) * revert zero tensors back * [tensor] init row 1d linear	2022-04-22 14:03:26 +08:00
Jiarui Fang	595bedf767	revert zero tensors back (#829 )	2022-04-22 12:12:35 +08:00
Jiarui Fang	294a6060d0	[tensor] ZeRO use ColoTensor as the base class. (#828 ) * [refactor] moving InsertPostInitMethodToModuleSubClasses to utils. * [tensor] ZeRO use ColoTensor as the base class. * polish	2022-04-22 12:00:48 +08:00
Ziyue Jiang	8e6fdb4f29	[tensor]fix test_linear (#826 )	2022-04-21 17:18:56 +08:00
Ziyue Jiang	1a9e2c2dff	[tensor] fix kwargs in colo_tensor torch_funtion (#825 )	2022-04-21 16:47:35 +08:00
Jiarui Fang	eb1b89908c	[refactor] moving InsertPostInitMethodToModuleSubClasses to utils. (#824 )	2022-04-21 16:03:18 +08:00
Jiarui Fang	2ecc3d7a55	[tensor] lazy init (#823 )	2022-04-21 15:40:23 +08:00
Jiarui Fang	68dcd51d41	[Tensor] update ColoTensor torch_function (#822 ) * Revert "[zero] add ZeroTensorShardStrategy (#793)" This reverts commit `88759e289e`. * [gemini] set cpu memory capacity * [log] local throughput collecting * polish * polish * polish * polish code * polish * polish code * add a new tensor structure and override linear for it * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish * [tensor] renaming and reorganize directory structure. * rm useless dir * polish * polish * [tensor] hander the function not wrapped * polish	2022-04-21 14:25:27 +08:00
Jiarui Fang	660d2d1f1b	[Tensor] apply ColoTensor on Torch functions (#821 ) * Revert "[zero] add ZeroTensorShardStrategy (#793)" This reverts commit `88759e289e`. * [gemini] set cpu memory capacity * [log] local throughput collecting * polish * polish * polish * polish code * polish * polish code * add a new tensor structure and override linear for it * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish * [tensor] renaming and reorganize directory structure. * rm useless dir * polish * polish * [tensor] hander the function not wrapped	2022-04-21 14:21:10 +08:00
Jiarui Fang	0ce8924ceb	[tensor] reorganize files (#820 )	2022-04-21 14:15:48 +08:00
Jiarui Fang	ab962b9735	[gemini] a new tensor structure (#818 ) * Revert "[zero] add ZeroTensorShardStrategy (#793)" This reverts commit `88759e289e`. * [gemini] set cpu memory capacity * [log] local throughput collecting * polish * polish * polish * polish code * polish * polish code * add a new tensor structure and override linear for it * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish * polish	2022-04-21 11:42:37 +08:00
github-actions[bot]	413ce30c45	Automated submodule synchronization (#819 ) Co-authored-by: github-actions <github-actions@github.com>	2022-04-21 11:26:58 +08:00
github-actions[bot]	9aae4197bb	Automated submodule synchronization (#810 ) Co-authored-by: github-actions <github-actions@github.com>	2022-04-20 13:57:12 +08:00
YuliangLiu0306	e1b3899824	Merge pull request #815 from FrankLeeeee/feature/check-cli [cli] added check installation cli	2022-04-20 12:19:50 +08:00
FrankLeeeee	70ed11d07e	[cli] added check installation cli	2022-04-20 12:13:27 +08:00
YuliangLiu0306	c7eca40f51	Merge pull request #812 from FrankLeeeee/feature/cli [cli] fixed single-node process launching	2022-04-20 11:40:07 +08:00
Jiarui Fang	3ddbd1bce1	[gemini] collect cpu-gpu moving volume in each iteration (#813 )	2022-04-20 11:29:48 +08:00
FrankLeeeee	d522cb704e	[cli] fixed single-node process launching	2022-04-20 10:46:51 +08:00
Jiarui Fang	61c20b44bc	[log] local throughput metrics (#811 ) * Revert "[zero] add ZeroTensorShardStrategy (#793)" This reverts commit `88759e289e`. * [gemini] set cpu memory capacity * [log] local throughput collecting * polish * polish * polish * polish code * polish	2022-04-20 10:05:39 +08:00
ver217	dd92b90a68	[DO NOT MERGE] [zero] init fp16 params directly in ZeroInitContext (#808 ) * init fp16 param directly * polish code	2022-04-19 16:16:48 +08:00
Jiarui Fang	227d1cd4b3	[gemini] APIs to set cpu memory capacity (#809 )	2022-04-19 16:05:22 +08:00
YuliangLiu0306	f6dcd23fb9	Merge pull request #807 from FrankLeeeee/feature/cli [cli] fixed a bug in user args and refactored the module structure	2022-04-19 15:52:26 +08:00

... 9 10 11 12 13 ...

1022 Commits (7dc53237c3956d7696564f3e61f04d20ff0aef9a) All Branches Search

1022 Commits (7dc53237c3956d7696564f3e61f04d20ff0aef9a)

All Branches