ColossalAI

Commit Graph

Author	SHA1	Message	Date
Frank Lee	a9b8402d93	[booster] added the accelerator implementation (#3159 )	2 years ago
Frank Lee	1ad3a636b1	[test] fixed torchrec model test (#3167 ) * [test] fixed torchrec model test * polish code * polish code * polish code * polish code * polish code * polish code	2 years ago
Saurav Maheshkar	20d1c99444	[refactor] update docs (#3174 ) * refactor: README-zh-Hans * refactor: REFERENCE * docs: update paths in README	2 years ago
BlueRum	7548ca5a54	[chatgpt]Reward Model Training Process update (#3133 ) * add normalize function to value_head in bloom rm * add normalization to value_function in gpt_rm * add normalization to value_head of opt_rm * add Anthropic/hh-rlhf dataset * Update __init__.py * Add LogExpLoss in RM training * Update __init__.py * update rm trainer to use acc as target * update example/train_rm * Update train_rm.sh * code style * Update README.md * Update README.md * add rm test to ci * fix tokenier * fix typo * change batchsize to avoid oom in ci * Update test_ci.sh	2 years ago
ver217	1e58d31bb7	[chatgpt] fix trainer generate kwargs (#3166 )	2 years ago
ver217	c474fda282	[chatgpt] fix ppo training hanging problem with gemini (#3162 ) * [chatgpt] fix generation early stopping * [chatgpt] fix train prompts example	2 years ago
ver217	6ae8ed0407	[lazyinit] add correctness verification (#3147 ) * [lazyinit] fix shared module * [tests] add lazy init test utils * [tests] add torchvision for lazy init * [lazyinit] fix pre op fn * [lazyinit] handle legacy constructor * [tests] refactor lazy init test models * [tests] refactor lazy init test utils * [lazyinit] fix ops don't support meta * [tests] lazy init test timm models * [lazyinit] fix set data * [lazyinit] handle apex layers * [tests] lazy init test transformers models * [tests] lazy init test torchaudio models * [lazyinit] fix import path * [tests] lazy init test torchrec models * [tests] update torch version in CI * [tests] revert torch version in CI * [tests] skip lazy init test	2 years ago
binmakeswell	3c01280a56	[doc] add community contribution guide (#3153 ) * [doc] update contribution guide * [doc] update contribution guide * [doc] add community contribution guide	2 years ago
Frank Lee	ed19290560	[booster] implemented mixed precision class (#3151 ) * [booster] implemented mixed precision class * polish code	2 years ago
YuliangLiu0306	ecd643f1e4	[test] add torchrec models to test model zoo (#3139 )	2 years ago
ver217	14a115000b	[tests] model zoo add torchaudio models (#3138 ) * [tests] model zoo add torchaudio models * [tests] refactor torchaudio wavernn * [tests] refactor fx torchaudio tests	2 years ago
Frank Lee	6d48eb0560	[test] added transformers models to test model zoo (#3135 )	2 years ago
Frank Lee	a674c63348	[test] added torchvision models to test model zoo (#3132 ) * [test] added torchvision models to test model zoo * polish code * polish code * polish code * polish code * polish code * polish code	2 years ago
HELSON	1216d1e7bd	[tests] diffuser models in model zoo (#3136 ) * [tests] diffuser models in model zoo * remove useless code * [tests] add diffusers to requirement-test	2 years ago
Saurav Maheshkar	1a46e71e07	[docker] Add opencontainers image-spec to `Dockerfile` (#3006 ) * feat(docker): Add opencontainers image-spec to `Dockerfile` This PR makes few changes to improve the overall quality of the docker image 🐳 . For reference more annotations can be found [here](https://github.com/opencontainers/image-spec/blob/main/annotations.md) * feat(docker): add inline version declaration * fix(docker): drop `org.opencontainers.image.version` LABEL	2 years ago
YuliangLiu0306	2eca4cd376	[DTensor] refactor dtensor with new components (#3089 ) * [DTensor] refactor dtensor with new components * polish	2 years ago
ver217	ed8f60b93b	[lazyinit] refactor lazy tensor and lazy init ctx (#3131 ) * [lazyinit] refactor lazy tensor and lazy init ctx * [lazyinit] polish docstr * [lazyinit] polish docstr	2 years ago
Frank Lee	86ac782d7c	[test] added timm models to test model zoo (#3129 ) * [test] added timm models to test model zoo * polish code * polish code * polish code * polish code * polish code	2 years ago
BlueRum	23cd5e2ccf	[chatgpt]update ci (#3087 ) * [chatgpt]update ci * Update test_ci.sh * Update test_ci.sh * Update test_ci.sh * test * Update train_prompts.py * Update train_dummy.py * add save_path * polish * add save path * polish * add save path * polish * delete bloom-560m test delete bloom-560m test because of oom * add ddp test	2 years ago
Frank Lee	169ed4d24e	[workflow] purged extension cache before GPT test (#3128 )	2 years ago
Xuanlei Zhao	30dd13c450	[autochunk] support complete benchmark (#3121 ) * refact memory code * dont log free var memory * add memory align * update chunk target * update setting for new memory * finish test * update tracer * update typo * update test * add unet test * add bench * update bench * update bench * init * support vit * move to cpu * add cpu benchmark	2 years ago
BlueRum	68577fbc43	[chatgpt]Fix examples (#3116 ) * fix train_dummy * fix train-prompts	2 years ago
BlueRum	0672b5afac	[chatgpt] fix lora support for gpt (#3113 ) * fix gpt-actor * fix gpt-critic * fix opt-critic	2 years ago
github-actions[bot]	0aa92c0409	Automated submodule synchronization (#3105 ) Co-authored-by: github-actions <github-actions@github.com>	2 years ago
Jeff Rasley	453f7ae5a0	prevent op_builder being installed in site-pkgs (#3104 )	2 years ago
hiko2MSP	191daf7411	[chatgpt] type miss of kwargs (#3107 )	2 years ago
binmakeswell	145ccfd7d1	[doc] add Intel cooperation for biomedicine (#3108 ) * [doc] add Intel cooperation for biomedicine	2 years ago
BlueRum	c9dd036592	[chatgpt] fix lora save bug (#3099 ) * fix colo-stratergy * polish * fix lora * fix ddp * polish * polish	2 years ago
binmakeswell	018936a3f3	[tutorial] update notes for TransformerEngine (#3098 )	2 years ago
Kirthi Shankar Sivamani	65a4dbda6c	[NVIDIA] Add FP8 example using TE (#3080 ) Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>	2 years ago
Frank Lee	26db1cb57b	[release] v0.2.7 (#3094 )	2 years ago
Fazzie-Maqianli	02ae80bf9c	[chatgpt]add flag of action mask in critic(#3086 )	2 years ago
Frank Lee	95a36eae63	[kernel] added kernel loader to softmax autograd function (#3093 ) * [kernel] added kernel loader to softmax autograd function * [release] v0.2.6	2 years ago
Super Daniel	fff98f06ed	[analyzer] a minimal implementation of static graph analyzer (#2852 ) * [hotfix] meta tensor default device. * [siu] add experimental submodules to main branch. * [siu] * [siu] * [analyzer] init. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [analyzer] readme. * [test] add test. * Update symbolic_trace.py * mark skip tests. * try except. * try except. * try except. * s * init * init * fix * skip * skip --------- Co-authored-by: Daniel Shao <superdainiu@MININT-PVARVID.fareast.corp.microsoft.com> Co-authored-by: Daniel Shao <superdainiu@Daniels-Mac.local>	2 years ago
Fazzie-Maqianli	5d5f475d75	[diffusers] fix ci and docker (#3085 )	2 years ago
Frank Lee	3213347b49	[doc] fixed typos in docs/README.md (#3082 )	2 years ago
Xuanlei Zhao	10c61de2f7	[autochunk] support vit (#3084 ) support vit for autochunk * support some new ops for vit * fix some bugs * add test for vit	2 years ago
Camille Zhong	e58a3c804c	Fix the version of lightning and colossalai in Stable Diffusion environment requirement (#3073 ) 1. Modify the README of stable diffusion 2. Fix the version of pytorch lightning&lightning and colossalai version to enable codes running successfully.	2 years ago
YuliangLiu0306	8e4e8601b7	[DTensor] implement layout converter (#3055 ) * [DTensor] refactor LayoutConverter for DTensor * polish code * polish docstring	2 years ago
Frank Lee	89aa7926ac	[release] v0.2.6 (#3057 )	2 years ago
Frank Lee	416a50dbd7	[doc] moved doc test command to bottom (#3075 )	2 years ago
Frank Lee	91ccf97514	[workflow] fixed doc build trigger condition (#3072 )	2 years ago
Frank Lee	f19b49e164	[booster] init module structure and definition (#3056 )	2 years ago
github-actions[bot]	faa8526b85	Automated submodule synchronization (#3062 ) Co-authored-by: github-actions <github-actions@github.com>	2 years ago
binmakeswell	360674283d	[example] fix redundant note (#3065 )	2 years ago
Tomek	af3888481d	[example] fixed opt model downloading from huggingface	2 years ago
Xuanlei Zhao	2ca9728cbb	[autochunk] refactor chunk memory estimation (#2762 ) * refact memory code * dont log free var memory * add memory align * update chunk target * update setting for new memory * finish test * update tracer * update typo * update test	2 years ago
wenjunyang	b51bfec357	[chatgpt] change critic input as state (#3042 ) * fix Critic * fix Critic * fix Critic * fix neglect of attention mask * fix neglect of attention mask * fix neglect of attention mask * add return --------- Co-authored-by: yangwenjun <yangwenjun@soyoung.com> Co-authored-by: yangwjd <yangwjd@chanjet.com>	2 years ago
ramos	2ef855c798	support shardinit option to avoid OPT OOM initializing problem (#3037 ) Co-authored-by: poe <poe@nemoramo>	2 years ago
YuliangLiu0306	29386a54e6	[DTensor] refactor CommSpec (#3034 )	2 years ago

1 2 3 4 5 ...

2180 Commits (31c78f2be3272a9a4062fe78eca34b3847a0c900) All Branches Search

2180 Commits (31c78f2be3272a9a4062fe78eca34b3847a0c900)

All Branches