237 Commits (822241a99cca799e1fca250ff2fb7f54ea0f8dcd)

Author SHA1 Message Date
binmakeswell 822241a99c
[doc] sora release (#5425) 9 months ago
Camille Zhong 4b8312c08e
fix sft single turn inference example (#5416) 9 months ago
Tong Li a28c971516
update requirements (#5407) 9 months ago
CZYCW b833153fd5
[hotfix] fix variable type for top_p (#5313) 9 months ago
Hongxin Liu 7303801854
[llama] fix training and inference scripts (#5384) 9 months ago
Hongxin Liu 65e5d6baa5 [moe] fix mixtral optim checkpoint (#5344) 10 months ago
Hongxin Liu 956b561b54 [moe] fix mixtral forward default value (#5329) 10 months ago
Hongxin Liu b60be18dcc [moe] fix mixtral checkpoint io (#5314) 10 months ago
Hongxin Liu da39d21b71 [moe] support mixtral (#5309) 10 months ago
Hongxin Liu c904d2ae99 [moe] update capacity computing (#5253) 10 months ago
Xuanlei Zhao 7d8e0338a4 [moe] init mixtral impl 10 months ago
Hongxin Liu 084c91246c
[llama] fix memory issue (#5371) 10 months ago
Hongxin Liu eb4f2d90f9
[llama] polish training script and fix optim ckpt (#5368) 10 months ago
Camille Zhong a5756a8720
[eval] update llama npu eval (#5366) 10 months ago
Camille Zhong 44ca61a22b
[llama] fix neftune & pbar with start_step (#5364) 10 months ago
Hongxin Liu a4cec1715b
[llama] add flash attn patch for npu (#5362) 10 months ago
Hongxin Liu 73f9f23fc6
[llama] update training script (#5360) 10 months ago
Hongxin Liu 6c0fa7b9a8
[llama] fix dataloader for hybrid parallel (#5358) 10 months ago
YeAnbang c5239840e6
[Chat] fix sft loss nan (#5345) 10 months ago
李文军 ec912b1ba9
[NFC] polish applications/Colossal-LLaMA-2/colossal_llama2/tokenizer/init_tokenizer.py code style (#5228) 10 months ago
Desperado-Jia ddf879e2db
fix bug for mefture (#5299) 10 months ago
Michelle 32cb74493a
fix auto loading gpt2 tokenizer (#5279) 10 months ago
digger yu 756c400ad2
fix typo in applications/ColossalEval/README.md (#5250) 11 months ago
digger yu 41e52c1c6e
[doc] fix typo in Colossal-LLaMA-2/README.md (#5247) 11 months ago
Hongxin Liu d202cc28c0
[npu] change device to accelerator api (#5239) 11 months ago
binmakeswell 7bc6969ce6
[doc] SwiftInfer release (#5236) 11 months ago
github-actions[bot] 4fb4a22a72
[format] applied code formatting on changed files in pull request 5234 (#5235) 11 months ago
binmakeswell b9b32b15e6
[doc] add Colossal-LLaMA-2-13B (#5234) 11 months ago
Camille Zhong 915b4652f3
[doc] Update README.md of Colossal-LLAMA2 (#5233) 11 months ago
Tong Li d992b55968
[Colossal-LLaMA-2] Release Colossal-LLaMA-2-13b-base model (#5224) 11 months ago
Yuanchen eae01b6740
Improve logic for selecting metrics (#5196) 11 months ago
BlueRum af952673f7
polish readme in application/chat (#5194) 11 months ago
Yuanchen 3ff60d13b0
Fix ColossalEval (#5186) 11 months ago
Yuanchen cefdc32615
[ColossalEval] Support GSM, Data Leakage Evaluation and Tensor Parallel (#5169) 12 months ago
Michelle b07a6f4e27
[colossalqa] fix pangu api (#5170) 12 months ago
Yuanchen b397104438
[Colossal-Llama-2] Add finetuning Colossal-Llama-2 example (#4878) 12 months ago
Michelle 368b5e3d64
[doc] fix colossalqa document (#5146) 12 months ago
Michelle c7fd9a5213
[ColossalQA] refactor server and webui & add new feature (#5138) 12 months ago
github-actions[bot] f6731db67c
[format] applied code formatting on changed files in pull request 5115 (#5118) 12 months ago
digger yu 9110406a47
fix typo change JOSNL TO JSONL etc. (#5116) 12 months ago
Zian(Andy) Zheng 7b789f4dd2 [FEATURE] Add Safety Eval Datasets to ColossalEval (#5095) 1 year ago
digger yu d5661f0f25
[nfc] fix typo change directoty to directory (#5111) 1 year ago
YeAnbang e53e729d8e
[Feature] Add document retrieval QA (#5020) 1 year ago
Orion-Zheng 43ad0d9ef0 fix wrong EOS token in ColossalChat 1 year ago
Yuanchen 239cd92eff
Support mtbench (#5025) 1 year ago
Yuanchen abe071b663
fix ColossalEval (#4992) 1 year ago
github-actions[bot] a41cf88e9b
[format] applied code formatting on changed files in pull request 4908 (#4918) 1 year ago
Zian(Andy) Zheng 7768afbad0 Update flash_attention_patch.py 1 year ago
Camille Zhong 652adc2215 Update README.md 1 year ago
Camille Zhong afe10a85fd Update README.md 1 year ago