2325 Commits (afb239bbf83737655bf6b6baef2e261768d5c60f)
 

Author SHA1 Message Date
ver217 c474fda282
[chatgpt] fix ppo training hanging problem with gemini (#3162) 2 years ago
ver217 6ae8ed0407
[lazyinit] add correctness verification (#3147) 2 years ago
binmakeswell 3c01280a56
[doc] add community contribution guide (#3153) 2 years ago
Frank Lee ed19290560
[booster] implemented mixed precision class (#3151) 2 years ago
YuliangLiu0306 ecd643f1e4
[test] add torchrec models to test model zoo (#3139) 2 years ago
ver217 14a115000b
[tests] model zoo add torchaudio models (#3138) 2 years ago
Frank Lee 6d48eb0560
[test] added transformers models to test model zoo (#3135) 2 years ago
Frank Lee a674c63348
[test] added torchvision models to test model zoo (#3132) 2 years ago
HELSON 1216d1e7bd
[tests] diffuser models in model zoo (#3136) 2 years ago
Saurav Maheshkar 1a46e71e07
[docker] Add opencontainers image-spec to `Dockerfile` (#3006) 2 years ago
YuliangLiu0306 2eca4cd376
[DTensor] refactor dtensor with new components (#3089) 2 years ago
ver217 ed8f60b93b
[lazyinit] refactor lazy tensor and lazy init ctx (#3131) 2 years ago
Frank Lee 86ac782d7c
[test] added timm models to test model zoo (#3129) 2 years ago
BlueRum 23cd5e2ccf
[chatgpt]update ci (#3087) 2 years ago
Frank Lee 169ed4d24e
[workflow] purged extension cache before GPT test (#3128) 2 years ago
Xuanlei Zhao 30dd13c450
[autochunk] support complete benchmark (#3121) 2 years ago
BlueRum 68577fbc43
[chatgpt]Fix examples (#3116) 2 years ago
BlueRum 0672b5afac
[chatgpt] fix lora support for gpt (#3113) 2 years ago
github-actions[bot] 0aa92c0409
Automated submodule synchronization (#3105) 2 years ago
Jeff Rasley 453f7ae5a0
prevent op_builder being installed in site-pkgs (#3104) 2 years ago
hiko2MSP 191daf7411
[chatgpt] type miss of kwargs (#3107) 2 years ago
binmakeswell 145ccfd7d1
[doc] add Intel cooperation for biomedicine (#3108) 2 years ago
BlueRum c9dd036592
[chatgpt] fix lora save bug (#3099) 2 years ago
binmakeswell 018936a3f3
[tutorial] update notes for TransformerEngine (#3098) 2 years ago
Kirthi Shankar Sivamani 65a4dbda6c
[NVIDIA] Add FP8 example using TE (#3080) 2 years ago
Frank Lee 26db1cb57b
[release] v0.2.7 (#3094) 2 years ago
Fazzie-Maqianli 02ae80bf9c
[chatgpt]add flag of action mask in critic(#3086) 2 years ago
Frank Lee 95a36eae63
[kernel] added kernel loader to softmax autograd function (#3093) 2 years ago
Super Daniel fff98f06ed
[analyzer] a minimal implementation of static graph analyzer (#2852) 2 years ago
Fazzie-Maqianli 5d5f475d75
[diffusers] fix ci and docker (#3085) 2 years ago
Frank Lee 3213347b49
[doc] fixed typos in docs/README.md (#3082) 2 years ago
Xuanlei Zhao 10c61de2f7
[autochunk] support vit (#3084) 2 years ago
Camille Zhong e58a3c804c
Fix the version of lightning and colossalai in Stable Diffusion environment requirement (#3073) 2 years ago
YuliangLiu0306 8e4e8601b7
[DTensor] implement layout converter (#3055) 2 years ago
Frank Lee 89aa7926ac
[release] v0.2.6 (#3057) 2 years ago
Frank Lee 416a50dbd7
[doc] moved doc test command to bottom (#3075) 2 years ago
Frank Lee 91ccf97514
[workflow] fixed doc build trigger condition (#3072) 2 years ago
Frank Lee f19b49e164
[booster] init module structure and definition (#3056) 2 years ago
github-actions[bot] faa8526b85
Automated submodule synchronization (#3062) 2 years ago
binmakeswell 360674283d
[example] fix redundant note (#3065) 2 years ago
Tomek af3888481d
[example] fixed opt model downloading from huggingface 2 years ago
Xuanlei Zhao 2ca9728cbb
[autochunk] refactor chunk memory estimation (#2762) 2 years ago
wenjunyang b51bfec357
[chatgpt] change critic input as state (#3042) 2 years ago
ramos 2ef855c798
support shardinit option to avoid OPT OOM initializing problem (#3037) 2 years ago
YuliangLiu0306 29386a54e6
[DTensor] refactor CommSpec (#3034) 2 years ago
Frank Lee ea0b52c12e
[doc] specified operating system requirement (#3019) 2 years ago
ver217 378d827c6b
[doc] update nvme offload doc (#3014) 2 years ago
Fazzie-Maqianli c21b11edce
change nn to models (#3032) 2 years ago
YuliangLiu0306 4269196c79
[hotfix] skip auto checkpointing tests (#3029) 2 years ago
Frank Lee 8fedc8766a
[workflow] supported conda package installation in doc test (#3028) 2 years ago