ColossalAI

Commit Graph

Author	SHA1	Message	Date
YH	80aed29cd3	[zero] Refactor ZeroContextConfig class using dataclass (#3186 )	2023-03-21 12:36:47 +08:00
YH	9d644ff09f	Fix docstr for zero statedict (#3185 )	2023-03-21 11:48:21 +08:00
zbian	7bc0afc901	updated flash attention usage	2023-03-20 17:57:04 +08:00
Frank Lee	085e7f4eff	[test] fixed torchrec registration in model zoo (#3177 ) * [test] fixed torchrec registration in model zoo * polish code * polish code * polish code	2023-03-20 16:19:06 +08:00
NatalieC323	4e921cfbd6	[examples] Solving the diffusion issue of incompatibility issue#3169 (#3170 ) * Update requirements.txt * Update environment.yaml * Update README.md * Update environment.yaml	2023-03-20 14:19:05 +08:00
Frank Lee	a9b8402d93	[booster] added the accelerator implementation (#3159 )	2023-03-20 13:59:24 +08:00
Frank Lee	1ad3a636b1	[test] fixed torchrec model test (#3167 ) * [test] fixed torchrec model test * polish code * polish code * polish code * polish code * polish code * polish code	2023-03-20 11:40:25 +08:00
Saurav Maheshkar	20d1c99444	[refactor] update docs (#3174 ) * refactor: README-zh-Hans * refactor: REFERENCE * docs: update paths in README	2023-03-20 10:52:01 +08:00
BlueRum	7548ca5a54	[chatgpt]Reward Model Training Process update (#3133 ) * add normalize function to value_head in bloom rm * add normalization to value_function in gpt_rm * add normalization to value_head of opt_rm * add Anthropic/hh-rlhf dataset * Update __init__.py * Add LogExpLoss in RM training * Update __init__.py * update rm trainer to use acc as target * update example/train_rm * Update train_rm.sh * code style * Update README.md * Update README.md * add rm test to ci * fix tokenier * fix typo * change batchsize to avoid oom in ci * Update test_ci.sh	2023-03-20 09:59:06 +08:00
ver217	1e58d31bb7	[chatgpt] fix trainer generate kwargs (#3166 )	2023-03-17 17:31:22 +08:00
ver217	c474fda282	[chatgpt] fix ppo training hanging problem with gemini (#3162 ) * [chatgpt] fix generation early stopping * [chatgpt] fix train prompts example	2023-03-17 15:41:47 +08:00
ver217	6ae8ed0407	[lazyinit] add correctness verification (#3147 ) * [lazyinit] fix shared module * [tests] add lazy init test utils * [tests] add torchvision for lazy init * [lazyinit] fix pre op fn * [lazyinit] handle legacy constructor * [tests] refactor lazy init test models * [tests] refactor lazy init test utils * [lazyinit] fix ops don't support meta * [tests] lazy init test timm models * [lazyinit] fix set data * [lazyinit] handle apex layers * [tests] lazy init test transformers models * [tests] lazy init test torchaudio models * [lazyinit] fix import path * [tests] lazy init test torchrec models * [tests] update torch version in CI * [tests] revert torch version in CI * [tests] skip lazy init test	2023-03-17 13:49:04 +08:00
binmakeswell	3c01280a56	[doc] add community contribution guide (#3153 ) * [doc] update contribution guide * [doc] update contribution guide * [doc] add community contribution guide	2023-03-17 11:07:24 +08:00
Frank Lee	ed19290560	[booster] implemented mixed precision class (#3151 ) * [booster] implemented mixed precision class * polish code	2023-03-17 11:00:15 +08:00
YuliangLiu0306	ecd643f1e4	[test] add torchrec models to test model zoo (#3139 )	2023-03-15 05:46:04 +00:00
ver217	14a115000b	[tests] model zoo add torchaudio models (#3138 ) * [tests] model zoo add torchaudio models * [tests] refactor torchaudio wavernn * [tests] refactor fx torchaudio tests	2023-03-15 11:51:16 +08:00
Frank Lee	6d48eb0560	[test] added transformers models to test model zoo (#3135 )	2023-03-15 11:26:10 +08:00
Frank Lee	a674c63348	[test] added torchvision models to test model zoo (#3132 ) * [test] added torchvision models to test model zoo * polish code * polish code * polish code * polish code * polish code * polish code	2023-03-15 10:42:07 +08:00
HELSON	1216d1e7bd	[tests] diffuser models in model zoo (#3136 ) * [tests] diffuser models in model zoo * remove useless code * [tests] add diffusers to requirement-test	2023-03-14 17:20:28 +08:00
Saurav Maheshkar	1a46e71e07	[docker] Add opencontainers image-spec to `Dockerfile` (#3006 ) * feat(docker): Add opencontainers image-spec to `Dockerfile` This PR makes few changes to improve the overall quality of the docker image 🐳 . For reference more annotations can be found [here](https://github.com/opencontainers/image-spec/blob/main/annotations.md) * feat(docker): add inline version declaration * fix(docker): drop `org.opencontainers.image.version` LABEL	2023-03-14 16:28:06 +08:00
YuliangLiu0306	2eca4cd376	[DTensor] refactor dtensor with new components (#3089 ) * [DTensor] refactor dtensor with new components * polish	2023-03-14 16:25:47 +08:00
ver217	ed8f60b93b	[lazyinit] refactor lazy tensor and lazy init ctx (#3131 ) * [lazyinit] refactor lazy tensor and lazy init ctx * [lazyinit] polish docstr * [lazyinit] polish docstr	2023-03-14 15:37:12 +08:00
Frank Lee	86ac782d7c	[test] added timm models to test model zoo (#3129 ) * [test] added timm models to test model zoo * polish code * polish code * polish code * polish code * polish code	2023-03-14 14:29:18 +08:00
BlueRum	23cd5e2ccf	[chatgpt]update ci (#3087 ) * [chatgpt]update ci * Update test_ci.sh * Update test_ci.sh * Update test_ci.sh * test * Update train_prompts.py * Update train_dummy.py * add save_path * polish * add save path * polish * add save path * polish * delete bloom-560m test delete bloom-560m test because of oom * add ddp test	2023-03-14 11:01:17 +08:00
Frank Lee	169ed4d24e	[workflow] purged extension cache before GPT test (#3128 )	2023-03-14 10:11:32 +08:00
Xuanlei Zhao	30dd13c450	[autochunk] support complete benchmark (#3121 ) * refact memory code * dont log free var memory * add memory align * update chunk target * update setting for new memory * finish test * update tracer * update typo * update test * add unet test * add bench * update bench * update bench * init * support vit * move to cpu * add cpu benchmark	2023-03-13 17:42:37 +08:00
BlueRum	68577fbc43	[chatgpt]Fix examples (#3116 ) * fix train_dummy * fix train-prompts	2023-03-13 11:12:22 +08:00
BlueRum	0672b5afac	[chatgpt] fix lora support for gpt (#3113 ) * fix gpt-actor * fix gpt-critic * fix opt-critic	2023-03-13 10:37:41 +08:00
github-actions[bot]	0aa92c0409	Automated submodule synchronization (#3105 ) Co-authored-by: github-actions <github-actions@github.com>	2023-03-13 08:58:06 +08:00
Jeff Rasley	453f7ae5a0	prevent op_builder being installed in site-pkgs (#3104 )	2023-03-13 08:50:31 +08:00
hiko2MSP	191daf7411	[chatgpt] type miss of kwargs (#3107 )	2023-03-13 00:00:02 +08:00
binmakeswell	145ccfd7d1	[doc] add Intel cooperation for biomedicine (#3108 ) * [doc] add Intel cooperation for biomedicine	2023-03-11 15:21:45 +08:00
BlueRum	c9dd036592	[chatgpt] fix lora save bug (#3099 ) * fix colo-stratergy * polish * fix lora * fix ddp * polish * polish	2023-03-10 17:58:10 +08:00
binmakeswell	018936a3f3	[tutorial] update notes for TransformerEngine (#3098 )	2023-03-10 16:30:52 +08:00
Kirthi Shankar Sivamani	65a4dbda6c	[NVIDIA] Add FP8 example using TE (#3080 ) Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>	2023-03-10 16:24:08 +08:00
Frank Lee	26db1cb57b	[release] v0.2.7 (#3094 )	2023-03-10 14:54:37 +08:00
Fazzie-Maqianli	02ae80bf9c	[chatgpt]add flag of action mask in critic(#3086 )	2023-03-10 14:40:14 +08:00
Frank Lee	95a36eae63	[kernel] added kernel loader to softmax autograd function (#3093 ) * [kernel] added kernel loader to softmax autograd function * [release] v0.2.6	2023-03-10 14:27:09 +08:00
Super Daniel	fff98f06ed	[analyzer] a minimal implementation of static graph analyzer (#2852 ) * [hotfix] meta tensor default device. * [siu] add experimental submodules to main branch. * [siu] * [siu] * [analyzer] init. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [test] add test. * Update symbolic_trace.py * mark skip tests. * try except. * try except. * try except. * s * init * init * fix * skip * skip --------- Co-authored-by: Daniel Shao <superdainiu@MININT-PVARVID.fareast.corp.microsoft.com> Co-authored-by: Daniel Shao <superdainiu@Daniels-Mac.local>	2023-03-10 13:21:05 +08:00
Fazzie-Maqianli	5d5f475d75	[diffusers] fix ci and docker (#3085 )	2023-03-10 10:35:15 +08:00
Frank Lee	3213347b49	[doc] fixed typos in docs/README.md (#3082 )	2023-03-10 10:32:14 +08:00
Xuanlei Zhao	10c61de2f7	[autochunk] support vit (#3084 ) support vit for autochunk * support some new ops for vit * fix some bugs * add test for vit	2023-03-10 10:23:26 +08:00
Camille Zhong	e58a3c804c	Fix the version of lightning and colossalai in Stable Diffusion environment requirement (#3073 ) 1. Modify the README of stable diffusion 2. Fix the version of pytorch lightning&lightning and colossalai version to enable codes running successfully.	2023-03-10 09:55:58 +08:00
YuliangLiu0306	8e4e8601b7	[DTensor] implement layout converter (#3055 ) * [DTensor] refactor LayoutConverter for DTensor * polish code * polish docstring	2023-03-10 09:53:52 +08:00
Frank Lee	89aa7926ac	[release] v0.2.6 (#3057 )	2023-03-10 09:47:20 +08:00
Frank Lee	416a50dbd7	[doc] moved doc test command to bottom (#3075 )	2023-03-09 18:10:45 +08:00
Frank Lee	91ccf97514	[workflow] fixed doc build trigger condition (#3072 )	2023-03-09 17:31:41 +08:00
Frank Lee	f19b49e164	[booster] init module structure and definition (#3056 )	2023-03-09 11:27:46 +08:00
github-actions[bot]	faa8526b85	Automated submodule synchronization (#3062 ) Co-authored-by: github-actions <github-actions@github.com>	2023-03-09 11:22:56 +08:00
binmakeswell	360674283d	[example] fix redundant note (#3065 )	2023-03-09 10:59:28 +08:00

1 2 3 4 5 ...

2235 Commits (4e9989344d20e3f8af44767f0eadeaab5fff8c00) All Branches Search

2235 Commits (4e9989344d20e3f8af44767f0eadeaab5fff8c00)

All Branches