ColossalAI/colossalai
Jiarui Fang 5a560a060a Feature/zero (#279)
* add zero1 (#209)

* add zero1

* add test zero1

* update zero stage 1 develop (#212)

* Implement naive zero3 (#240)

* naive zero3 works well

* add zero3 param manager

* add TODOs in comments

* add gather full param ctx

* fix sub module streams

* add offload

* fix bugs of hook and add unit tests

* fix bugs of hook and add unit tests (#252)

* add gather full param ctx

* fix sub module streams

* add offload

* fix bugs of hook and add unit tests

* polish code and add state dict hook

* fix bug

* update unit test

* refactor reconstructed zero code

* clip_grad support zero3 and add unit test

* add unit test for Zero3ParameterManager

* [WIP] initialize the shard param class

* [WIP] Yet another sharded model implementation (#274)

* [WIP] initialize the shard param class

* [WIP] Yes another implementation of shardModel. Using a better hook method.

* torch.concat -> torch.cat

* fix test_zero_level_1.py::test_zero_level_1 unitest

* remove deepspeed implementation and refactor for the reconstructed zero module

* polish zero dp unittests

Co-authored-by: ver217 <lhx0217@gmail.com>
Co-authored-by: Frank Lee <somerlee.9@gmail.com>
2022-03-11 15:50:28 +08:00
..
amp fixed apex import (#227) 2022-02-15 11:31:13 +08:00
builder add pytorch hooks (#179) 2022-01-25 22:20:54 +08:00
communication moved env variables to global variables; (#215) 2022-02-15 11:31:13 +08:00
context moved env variables to global variables; (#215) 2022-02-15 11:31:13 +08:00
engine Feature/zero (#279) 2022-03-11 15:50:28 +08:00
kernel Optimized MoE layer and fixed some bugs; 2022-03-11 15:50:28 +08:00
logging fixed mkdir conflict and align yapf config with flake (#220) 2022-02-15 11:31:13 +08:00
nn Added TPExpert for special situation 2022-03-11 15:50:28 +08:00
registry add pytorch hooks (#179) 2022-01-25 22:20:54 +08:00
trainer moved env variables to global variables; (#215) 2022-02-15 11:31:13 +08:00
utils Feature/zero (#279) 2022-03-11 15:50:28 +08:00
zero Feature/zero (#279) 2022-03-11 15:50:28 +08:00
__init__.py Develop/experiments (#59) 2021-12-09 15:08:29 +08:00
constants.py moved env variables to global variables; (#215) 2022-02-15 11:31:13 +08:00
core.py Develop/experiments (#59) 2021-12-09 15:08:29 +08:00
global_variables.py Optimized MoE layer and fixed some bugs; 2022-03-11 15:50:28 +08:00
initialize.py Feature/zero (#279) 2022-03-11 15:50:28 +08:00