ver217
|
7faef93326
|
fix dist spec mgr (#1045)
|
3 years ago |
ver217
|
9492a561c3
|
[tensor] ColoTensor supports ZeRo (#1015)
* impl chunk manager
* impl param op hook
* add reduce_chunk
* add zero hook v2
* add zero dp
* fix TensorInfo
* impl load balancing when using zero without chunk
* fix zero hook
* polish chunk
* fix bugs
* ddp ok
* zero ok
* polish code
* fix bugs about load balancing
* polish code
* polish code
* add ene-to-end test
* polish code
* polish code
* polish code
* fix typo
* add test_chunk
* fix bugs
* fix bugs
* polish code
|
3 years ago |
YuliangLiu0306
|
9feff0f760
|
[titans]remove model zoo (#1042)
* [CLI] add CLI launcher
* Revert "[CLI] add CLI launcher"
This reverts commit df7e6506d4 .
* rm model zoo
|
3 years ago |
Ziyue Jiang
|
7c530b9de2
|
[Tensor] add Parameter inheritance for ColoParameter (#1041)
* add Parameter inheritance for ColoParameter
* remove tricks
* remove tricks
* polish
* polish
|
3 years ago |
Ziyue Jiang
|
6c5996a56e
|
[Tensor] add module check and bert test (#1031)
* add Embedding
* Add bert test
* polish
* add check module test
* polish
* polish
* polish
* polish
|
3 years ago |
YuliangLiu0306
|
7106bd671d
|
[p2p]add object list send/recv (#1024)
* [CLI] add CLI launcher
* Revert "[CLI] add CLI launcher"
This reverts commit df7e6506d4 .
* [p2p]add object list send recv
* refactor for code reusability
* polish
|
3 years ago |
Ziyue Jiang
|
32291dd73f
|
[Tensor] add module handler for linear (#1021)
* add module spec for linear
* polish
* polish
* polish
|
3 years ago |
ver217
|
cefc29ff06
|
[tensor] impl ColoDDP for ColoTensor (#1009)
* impl ColoDDP for ColoTensor
* polish code
|
3 years ago |
ver217
|
a3b66f6def
|
[tensor] refactor parallel action (#1007)
* refactor parallel action
* polish unit tests
|
3 years ago |
ver217
|
8e3d0ad8f1
|
[unit test] refactor test tensor (#1005)
* polish test_gpt
* update op unit tests
* update test model
|
3 years ago |
ver217
|
ad536e308e
|
[tensor] refactor colo-tensor (#992)
* refactor colo-tensor and update linear op
* polish code
* polish code
* update ops and unit tests
* update unit tests
* polish code
* rename dist_spec module
* polish code
* polish code
* remove unneeded import
* fix pipelinable
|
3 years ago |
ver217
|
c2fdc6a011
|
[tensor] derive compute pattern from dist spec (#971)
* derive compute pattern from dist spec
* polish code
|
3 years ago |
Ziyue Jiang
|
797a9dc5a9
|
add DistSpec for loss and test_model (#947)
|
3 years ago |
ver217
|
67c33f57eb
|
[tensor] design DistSpec and DistSpecManager for ColoTensor (#934)
* add dist spec
* update linear op
* polish code
* polish code
* update embedding op
* polish unit tests
* polish unit tests
* polish comments
* polish code
* add test_dist_spec_mgr
* polish code
* refactor folder structure
* polish unit tests
* add get_process_group() for TensorSpec
* polish code
|
3 years ago |
Ziyue Jiang
|
830d3bca26
|
[Tensor] add optimizer to bert test (#933)
* add optimizer to bert test
* polish
|
3 years ago |
Ziyue Jiang
|
d73c2b1d79
|
[Tensor] fix init context (#931)
* change torch.Parameter to ColoParameter
* fix post assignment for init context
* polish
* polish
|
3 years ago |
Ziyue Jiang
|
dfc88b85ea
|
[Tensor] simplify named param (#928)
* simplify ColoModulize
* simplify ColoModulize
* polish
* polish
|
3 years ago |
ver217
|
45b9124df4
|
[tensor] hijack addmm for colo tensor (#923)
* hijack addmm for colo tensor
* fix bugs
* polish unit test
* polish comments
|
3 years ago |
Jiarui Fang
|
534afb018a
|
test pretrain loading on multi-process (#922)
|
3 years ago |
Ziyue Jiang
|
c195d2814c
|
[Tensor] add from_pretrained support and bert pretrained test (#921)
* add from_pretrained support and test
* polish
* polish
* polish
* polish
|
3 years ago |
Jiarui Fang
|
845856ea29
|
[Graph] building computing graph with ColoTensor, Linear only (#917)
|
3 years ago |
Ziyue Jiang
|
75d221918a
|
[Tensor] add 1d vocab loss (#918)
* add 1d vocab loss
* polish
|
3 years ago |
Ziyue Jiang
|
dfaff4e243
|
[Tensor] fix test_model (#916)
* polish test_model
* polish
|
3 years ago |
Jiarui Fang
|
ed6426c300
|
[Tensor] polish model test (#915)
|
3 years ago |
Ziyue Jiang
|
0fab86b12a
|
[Tensor] add a basic bert. (#911)
* add base bert test
* Add bert test
* polish
* remove test_bert
* polish
|
3 years ago |
Jiarui Fang
|
ab95ec9aea
|
[Tensor] init ColoParameter (#914)
|
3 years ago |
Ziyue Jiang
|
193d629311
|
update pytest.mark.parametrize in tensor tests (#913)
|
3 years ago |
Ziyue Jiang
|
f593a5637e
|
[Tensor] add embedding tp1d row (#904)
|
3 years ago |
Ziyue Jiang
|
2c0d19d755
|
[Tensor] add ColoTensor TP1Dcol Embedding (#899)
|
3 years ago |
Jiarui Fang
|
d16671da75
|
[Tensor] initialize the ColoOptimizer (#898)
* [Tensor] activation is an attr of ColoTensor
* [Tensor] add optimizer
* only detach parameters in context
* polish code
|
3 years ago |
Jiarui Fang
|
e76f76c08b
|
[Tensor] test parameters() as member function (#896)
|
3 years ago |
Ziyue Jiang
|
cb182da7c5
|
[tensor] refine linear and add gather for laynorm (#893)
* refine linear and add function to ColoTensor
* add gather for layernorm
* polish
* polish
|
3 years ago |
Jiarui Fang
|
26c49639d8
|
[Tensor] overriding paramters() for Module using ColoTensor (#889)
|
3 years ago |
Ziyue Jiang
|
1d0aba4153
|
[tensor] add ColoTensor 1Dcol (#888)
|
3 years ago |
Jiarui Fang
|
a0e5971692
|
[Tensor] test model check results for a simple net (#887)
|
3 years ago |
Jiarui Fang
|
72cdc06875
|
[Tensor] make ColoTensor more robust for getattr (#886)
* [Tensor] make ColoTensor more robust for getattr
* polish
* polish
|
3 years ago |
Ziyue Jiang
|
9bc5a77c31
|
[tensor] wrap function in the torch_tensor to ColoTensor (#881)
|
3 years ago |
Jiarui Fang
|
7f76517a85
|
[Tensor] make a simple net works with 1D row TP (#879)
|
3 years ago |
ver217
|
c4d903e64a
|
[gemini] accelerate adjust_layout() (#878)
* add lru cache
* polish code
* update unit test
* fix sharded optim
|
3 years ago |
Jiarui Fang
|
909211453b
|
[Tensor] Add some attributes to ColoTensor (#877)
* [Tensor] add some function to ColoTensor
* torch.allclose
* rm torch.add
|
3 years ago |
Jiarui Fang
|
e43f83aa5c
|
[Tensor] get named parameters for model using ColoTensors (#874)
|
3 years ago |
Jiarui Fang
|
96211c2cc8
|
[tensor] customized op returns ColoTensor (#875)
* [tensor] customized op returns ColoTensor
* polish
* polish code
|
3 years ago |
Ziyue Jiang
|
26d4ab8b03
|
[Tensor] Add function to spec and update linear 1Drow and unit tests (#869)
|
3 years ago |
Jiarui Fang
|
1190b2c4a4
|
[tensor] add cross_entrophy_loss (#868)
|
3 years ago |
HELSON
|
3107817172
|
[gemini] add stateful tensor container (#867)
|
3 years ago |
Jiarui Fang
|
d01d3b8cb0
|
colo init context add device attr. (#866)
|
3 years ago |
Jiarui Fang
|
126ba573a8
|
[Tensor] add layer norm Op (#852)
|
3 years ago |
Frank Lee
|
1258af71cc
|
[ci] cache cuda extension (#860)
|
3 years ago |
Ziyue Jiang
|
bcc8655021
|
[Tensor ] Add 1Drow weight reshard by spec (#854)
|
3 years ago |
Jiarui Fang
|
62f059251b
|
[Tensor] init a tp network training unittest (#849)
|
3 years ago |