YH
|
ae86a29e23
|
Refact method of grad store (#2687)
|
2023-02-15 22:27:58 +08:00 |
HELSON
|
df4f020ee3
|
[zero1&2] only append parameters with gradients (#2681)
|
2023-02-13 18:00:16 +08:00 |
HELSON
|
b528eea0f0
|
[zero] add zero wrappers (#2523)
* [zero] add zero wrappers
* change names
* add wrapper functions to init
|
2023-01-29 17:52:58 +08:00 |
HELSON
|
d565a24849
|
[zero] add unit testings for hybrid parallelism (#2486)
|
2023-01-18 10:36:10 +08:00 |
HELSON
|
a5dc4253c6
|
[zero] polish low level optimizer (#2473)
|
2023-01-13 14:56:17 +08:00 |
Jiarui Fang
|
867c8c2d3a
|
[zero] low level optim supports ProcessGroup (#2464)
|
2023-01-13 10:05:58 +08:00 |
HELSON
|
62c38e3330
|
[zero] polish low level zero optimizer (#2275)
|
2023-01-03 17:22:34 +08:00 |
HELSON
|
a7d95b7024
|
[example] add zero1, zero2 example in GPT examples (#2146)
* [example] add zero1 and zero2 for GPT
* update readme in gpt example
* polish code
* change init value
* update readme
|
2022-12-20 14:30:27 +08:00 |
HELSON
|
a1ce02d740
|
[zero] test gradient accumulation (#1964)
* [zero] fix memory leak for zero2
* [zero] test gradient accumulation
* [zero] remove grad clip test
|
2022-11-29 13:00:30 +08:00 |
HELSON
|
7066dfbf82
|
[zero] fix memory leak for zero2 (#1955)
|
2022-11-16 11:43:24 +08:00 |
HELSON
|
6e51d296f0
|
[zero] migrate zero1&2 (#1878)
* add zero1&2 optimizer
* rename test ditectory
* rename test files
* change tolerance in test
|
2022-11-11 09:26:40 +08:00 |