203 Commits (fix-setup)

Author SHA1 Message Date
Ziyue Jiang c195d2814c
[Tensor] add from_pretrained support and bert pretrained test (#921) 3 years ago
Jiarui Fang ab95ec9aea
[Tensor] init ColoParameter (#914) 3 years ago
Jiarui Fang d16671da75
[Tensor] initialize the ColoOptimizer (#898) 3 years ago
Jiarui Fang 676f191532
[Tensor] activation is an attr of ColoTensor (#897) 3 years ago
Jiarui Fang 26c49639d8
[Tensor] overriding paramters() for Module using ColoTensor (#889) 3 years ago
ver217 4df6471f5d
fix import error (#880) 3 years ago
Jiarui Fang d01d3b8cb0
colo init context add device attr. (#866) 3 years ago
YuliangLiu0306 c6930d8ddf
[pipelinable]use ColoTensor to replace dummy tensor. (#853) 3 years ago
ver217 232142f402
[utils] refactor profiler (#837) 3 years ago
Jiarui Fang 62f059251b
[Tensor] init a tp network training unittest (#849) 3 years ago
ver217 0dea140760
[hotfix] add deconstructor for stateful tensor (#848) 3 years ago
YuliangLiu0306 35ea6e1023
[pipelinable]use pipelinable context to initialize non-pipeline model (#816) 3 years ago
Jiarui Fang 8789850eea
Init Conext supports lazy allocate model memory (#842) 3 years ago
Jiarui Fang eb1b89908c
[refactor] moving InsertPostInitMethodToModuleSubClasses to utils. (#824) 3 years ago
Jiarui Fang 227d1cd4b3
[gemini] APIs to set cpu memory capacity (#809) 3 years ago
Jiarui Fang 681addb512
[refactor] moving grad acc logic to engine (#804) 3 years ago
Jiarui Fang 4d9332b4c5
[refactor] moving memtracer to gemini (#801) 3 years ago
HELSON 84c6700b2a
[zero] refactor memstats_collector (#746) 3 years ago
HELSON 340e59f968
[utils] add synchronized cuda memory monitor (#740) 3 years ago
Jiarui Fang 53cb584808
[utils] correct cpu memory used and capacity in the context of multi-process (#726) 3 years ago
Frank Lee 2412429d54
[util] fixed activation checkpointing on torch 1.9 (#719) 3 years ago
Jiarui Fang 193dc8dacb
[refactor] refactor the memory utils (#715) 3 years ago
LuGY 140263a394
[hotfix]fixed bugs of assigning grad states to non leaf nodes (#711) 3 years ago
ver217 ab8c6b4a0e
[zero] refactor memstats collector (#706) 3 years ago
ver217 3c9cd5bb5e
[zero] stateful tensor manager (#687) 3 years ago
Jiarui Fang 59bf2dc590
[zero] initialize a stateful tensor manager (#614) 3 years ago
Jiarui Fang 0aab52301e
[hotfix] fix a bug in model data stats tracing (#655) 3 years ago
HELSON e5d615aeee
[hotfix] fix bugs in testing (#659) 3 years ago
LuGY 1e2557e801
[zero] fixed the activation offload (#647) 3 years ago
ver217 f5d3a9c2b0
polish checkpoint docstring (#637) 3 years ago
HELSON 055fbf5be6
[zero] adapt zero for unsharded paramters (Optimizer part) (#601) 3 years ago
アマデウス acae68eb04
[model checkpoint] updated checkpoint save/load utils (#592) 3 years ago
ver217 369a288bf3
polish utils docstring (#620) 3 years ago
LuGY 02b187c14f
[zero] add sampling time for memstats collector (#610) 3 years ago
アマデウス 54e688b623
moved ensure_path_exists to utils.common (#591) 3 years ago
Jiarui Fang e956d93ac2
[refactor] memory utils (#577) 3 years ago
HELSON e6d50ec107
[zero] adapt zero for unsharded parameters (#561) 3 years ago
ver217 7c6c427db1
[zero] trace states of fp16/32 grad and fp32 param (#571) 3 years ago
Jiarui Fang 7675366fce
[polish] rename col_attr -> colo_attr (#558) 3 years ago
Liang Bowen 2c45efc398
html refactor (#555) 3 years ago
Jiarui Fang d1211148a7
[utils] update colo tensor moving APIs (#553) 3 years ago
Jiarui Fang 107b99ddb1
[zero] dump memory stats for sharded model (#548) 3 years ago
Liang Bowen ec5086c49c Refactored docstring to google style 3 years ago
Jiarui Fang 53b1b6e340
[zero] non model data tracing (#545) 3 years ago
Jie Zhu 73d36618a6
[profiler] add MemProfiler (#356) 3 years ago
Jiarui Fang c11ff81b15
[zero] get memory usage of sharded optim v2. (#542) 3 years ago
Jiarui Fang 705f56107c
[zero] refactor model data tracing (#537) 3 years ago
Jiarui Fang 05e33b2578
[zero] fix grad offload (#528) 3 years ago
Jiarui Fang 8d8c5407c0
[zero] refactor model data tracing (#522) 3 years ago
Jiarui Fang 920c5889a7
[zero] add colo move inline (#521) 3 years ago