331 Commits (8241c0c054b38a109ed3ce7be1052a1e600b8471)

Author SHA1 Message Date
HELSON a65cbb7e4e
[zero] refactor shard and gather operation (#773) 3 years ago
ver217 6e553748a7
polish sharded optim docstr and warning (#770) 3 years ago
Jiarui Fang 10ef8afdd2
[gemini] init genimi individual directory (#754) 3 years ago
ver217 dcca614eee
[hotfix] fix test_stateful_tensor_mgr (#762) 3 years ago
ver217 a93a7d7364
[hotfix] fix reuse_fp16_shard of sharded model (#756) 3 years ago
ver217 8f7ce94b8e
[hotfix] fix auto tensor placement policy (#753) 3 years ago
HELSON 84c6700b2a
[zero] refactor memstats_collector (#746) 3 years ago
Jiarui Fang 3d7dc46d33
[zero] use factory pattern for tensor_placement_policy (#752) 3 years ago
ver217 4b048a8728
fix prepare grads in sharded optim (#749) 3 years ago
ver217 e396bb71f2
[zero] add tensor placement policies (#743) 3 years ago
HELSON 22c4b88d56
[zero] refactor ShardedParamV2 for convenience (#742) 3 years ago
ver217 e6212f56cd
[hotfix] fix memory leak in backward of sharded model (#741) 3 years ago
Jiarui Fang 7db3ccc79b
[hotfix] remove duplicated param register to stateful tensor manager (#728) 3 years ago
Jiarui Fang 4d90a7b513
[refactor] zero directory (#724) 3 years ago
Jiarui Fang 193dc8dacb
[refactor] refactor the memory utils (#715) 3 years ago
HELSON dbd96fe90a
[zero] check whether gradients have inf and nan in gpu (#712) 3 years ago
ver217 715b86eadd
[hotfix] fix stm cuda model data size (#710) 3 years ago
HELSON a9b8300d54
[zero] improve adaptability for not-shard parameters (#708) 3 years ago
ver217 ab8c6b4a0e
[zero] refactor memstats collector (#706) 3 years ago
HELSON ee112fe1da
[zero] adapt zero hooks for unsharded module (#699) 3 years ago
ver217 3c9cd5bb5e
[zero] stateful tensor manager (#687) 3 years ago
HELSON d7ecaf362b
[zero] fix init bugs in zero context (#686) 3 years ago
Jiarui Fang 59bf2dc590
[zero] initialize a stateful tensor manager (#614) 3 years ago
HELSON 17e73e62cc
[hotfix] fix bugs for unsharded parameters when restore data (#664) 3 years ago
Jiarui Fang 0aab52301e
[hotfix] fix a bug in model data stats tracing (#655) 3 years ago
Jiarui Fang 036404ca8a
Revert "[zero] polish init context (#645)" (#657) 3 years ago
Jiarui Fang 67b4928244
[zero] polish init context (#645) 3 years ago
HELSON 055fbf5be6
[zero] adapt zero for unsharded paramters (Optimizer part) (#601) 3 years ago
ver217 0ef8819c67
polish docstring of zero (#612) 3 years ago
ver217 9bee119104
[hotfix] fix sharded optim zero grad (#604) 3 years ago
Jiarui Fang e956d93ac2
[refactor] memory utils (#577) 3 years ago
HELSON e6d50ec107
[zero] adapt zero for unsharded parameters (#561) 3 years ago
ver217 7c6c427db1
[zero] trace states of fp16/32 grad and fp32 param (#571) 3 years ago
Jiarui Fang 7675366fce
[polish] rename col_attr -> colo_attr (#558) 3 years ago
ver217 014bac0c49
[zero] hijack p.grad in sharded model (#554) 3 years ago
Jiarui Fang f552b11294
[zero] label state for param fp16 and grad (#551) 3 years ago
Jiarui Fang 214da761d4
[zero] add stateful tensor (#549) 3 years ago
Jiarui Fang 107b99ddb1
[zero] dump memory stats for sharded model (#548) 3 years ago
HELSON 8c90d4df54
[zero] add zero context manager to change config during initialization (#546) 3 years ago
Jiarui Fang 53b1b6e340
[zero] non model data tracing (#545) 3 years ago
ver217 fb841dd5c5
[zero] optimize grad offload (#539) 3 years ago
ver217 1f90a3b129
[zero] polish ZeroInitContext (#540) 3 years ago
Jiarui Fang c11ff81b15
[zero] get memory usage of sharded optim v2. (#542) 3 years ago
HELSON a30e2b4c24
[zero] adapt for no-leaf module in zero (#535) 3 years ago
Jiarui Fang 705f56107c
[zero] refactor model data tracing (#537) 3 years ago
Jiarui Fang a590ed0ba3
[zero] improve the accuracy of get_memory_usage of sharded param (#538) 3 years ago
Jiarui Fang 37cb70feec
[zero] get memory usage for sharded param (#536) 3 years ago
Jiarui Fang 05e33b2578
[zero] fix grad offload (#528) 3 years ago
Jiarui Fang 8d8c5407c0
[zero] refactor model data tracing (#522) 3 years ago
Jiarui Fang 4d322b79da
[refactor] remove old zero code (#517) 3 years ago