955 Commits (39f2582e987871c198f2f2526cd4435cbd569741)

Author SHA1 Message Date
Hongxin Liu f313babd11
[gemini] support save state dict in shards (#3581) 2 years ago
Hongxin Liu 152239bbfa
[gemini] gemini supports lazy init (#3379) 2 years ago
jiangmingyan 52a933e175
[checkpoint] support huggingface style sharded checkpoint (#3461) 2 years ago
Frank Lee 80eba05b0a
[test] refactor tests with spawn (#3452) 2 years ago
ver217 933048ad3e
[test] reorganize zero/gemini tests (#3445) 2 years ago
YuliangLiu0306 ffcdbf0f65
[autoparallel]integrate auto parallel feature with new tracer (#3408) 2 years ago
Frank Lee 1beb85cc25
[checkpoint] refactored the API and added safetensors support (#3427) 2 years ago
ver217 26b7aac0be
[zero] reorganize zero/gemini folder structure (#3424) 2 years ago
Frank Lee 638a07a7f9
[test] fixed gemini plugin test (#3411) 2 years ago
ver217 5f2e34e6c9
[booster] implement Gemini plugin (#3352) 2 years ago
HELSON 1a1d68b053
[moe] add checkpoint for moe models (#3354) 2 years ago
YuliangLiu0306 fee2af8610
[autoparallel] adapt autoparallel with new analyzer (#3261) 2 years ago
Frank Lee 73d3e4d309
[booster] implemented the torch ddd + resnet example (#3232) 2 years ago
YuliangLiu0306 4d5d8f98a4
[API] implement device mesh manager (#3221) 2 years ago
YuliangLiu0306 045afa3ea2
[hotfix] skip torchaudio tracing test (#3211) 2 years ago
Frank Lee cd142fbefa
[api] implemented the checkpoint io module (#3205) 2 years ago
ver217 f8289d4221
[lazyinit] combine lazy tensor with dtensor (#3204) 2 years ago
YuliangLiu0306 019a847432
[Analyzer] fix analyzer tests (#3197) 2 years ago
YuliangLiu0306 f57d34958b
[FX] refactor experimental tracer and adapt it with hf models (#3157) 2 years ago
Frank Lee e7f3bed2d3
[booster] added the plugin base and torch ddp plugin (#3180) 2 years ago
Zihao 18dbe76cae
[auto-parallel] add auto-offload feature (#3154) 2 years ago
zbian 7bc0afc901 updated flash attention usage 2 years ago
Frank Lee 085e7f4eff
[test] fixed torchrec registration in model zoo (#3177) 2 years ago
Frank Lee a9b8402d93
[booster] added the accelerator implementation (#3159) 2 years ago
Frank Lee 1ad3a636b1
[test] fixed torchrec model test (#3167) 2 years ago
ver217 6ae8ed0407
[lazyinit] add correctness verification (#3147) 2 years ago
Frank Lee ed19290560
[booster] implemented mixed precision class (#3151) 2 years ago
YuliangLiu0306 ecd643f1e4
[test] add torchrec models to test model zoo (#3139) 2 years ago
ver217 14a115000b
[tests] model zoo add torchaudio models (#3138) 2 years ago
Frank Lee 6d48eb0560
[test] added transformers models to test model zoo (#3135) 2 years ago
Frank Lee a674c63348
[test] added torchvision models to test model zoo (#3132) 2 years ago
HELSON 1216d1e7bd
[tests] diffuser models in model zoo (#3136) 2 years ago
YuliangLiu0306 2eca4cd376
[DTensor] refactor dtensor with new components (#3089) 2 years ago
Frank Lee 86ac782d7c
[test] added timm models to test model zoo (#3129) 2 years ago
Xuanlei Zhao 30dd13c450
[autochunk] support complete benchmark (#3121) 2 years ago
Super Daniel fff98f06ed
[analyzer] a minimal implementation of static graph analyzer (#2852) 2 years ago
Xuanlei Zhao 10c61de2f7
[autochunk] support vit (#3084) 2 years ago
YuliangLiu0306 8e4e8601b7
[DTensor] implement layout converter (#3055) 2 years ago
Xuanlei Zhao 2ca9728cbb
[autochunk] refactor chunk memory estimation (#2762) 2 years ago
YuliangLiu0306 29386a54e6
[DTensor] refactor CommSpec (#3034) 2 years ago
YuliangLiu0306 4269196c79
[hotfix] skip auto checkpointing tests (#3029) 2 years ago
YuliangLiu0306 cd2b0eaa8d
[DTensor] refactor sharding spec (#2987) 2 years ago
YuliangLiu0306 e414e4092b
[DTensor] implementation of dtensor (#2946) 2 years ago
YuliangLiu0306 197d0bf4ed
[autoparallel] apply repeat block to reduce solving time (#2912) 2 years ago
YuliangLiu0306 819e25d8b1
[hotfix] fix autoparallel compatibility test issues (#2754) 2 years ago
YuliangLiu0306 0f392d7403
[autoparallel] find repeat blocks (#2854) 2 years ago
Boyuan Yao c7764d3f22
[autoparallel] Patch meta information of `torch.where` (#2822) 2 years ago
Boyuan Yao fcc4097efa
[autoparallel] Patch meta information of `torch.tanh()` and `torch.nn.Dropout` (#2773) 2 years ago
Boyuan Yao 7ea6bc7f69
[autoparallel] Patch tensor related operations meta information (#2789) 2 years ago
HELSON 56ddc9ca7a
[hotfix] add correct device for fake_param (#2796) 2 years ago