Commit Graph

3257 Commits (feat/online-serving)
 

Author SHA1 Message Date
xs_courtesy 01d289d8e5 Merge branch 'feature/colossal-infer' of https://github.com/hpcaitech/ColossalAI into add_gpu_launch_config
9 months ago
xs_courtesy a46598ac59 add reusable utils for cuda
9 months ago
傅剑寒 2b28b54ac6
Merge pull request #5433 from Courtesy-Xs/add_silu_and_mul
9 months ago
Runyu Lu cefaeb5fdd [feat] cuda graph support and refactor non-functional api
9 months ago
Hongxin Liu 8020f42630
[release] update version (#5411)
9 months ago
xs_courtesy 95c21498d4 add silu_and_mul for infer
9 months ago
Camille Zhong 743e7fad2f
[colossal-llama2] add stream chat examlple for chat version model (#5428)
9 months ago
Youngon 68f55a709c
[hotfix] fix stable diffusion inference bug. (#5289)
9 months ago
hugo-syn c8003d463b
[doc] Fix typo s/infered/inferred/ (#5288)
9 months ago
digger yu 5e1c93d732
[hotfix] fix typo change MoECheckpintIO to MoECheckpointIO (#5335)
9 months ago
Dongruixuan Li a7ae2b5b4c
[eval-hotfix] set few_shot_data to None when few shot is disabled (#5422)
9 months ago
digger yu 049121d19d
[hotfix] fix typo change enabel to enable under colossalai/shardformer/ (#5317)
9 months ago
digger yu 16c96d4d8c
[hotfix] fix typo change _descrption to _description (#5331)
9 months ago
digger yu 70cce5cbed
[doc] update some translations with README-zh-Hans.md (#5382)
9 months ago
Luo Yihang e239cf9060
[hotfix] fix typo of openmoe model source (#5403)
9 months ago
MickeyCHAN e304e4db35
[hotfix] fix sd vit import error (#5420)
9 months ago
Hongxin Liu 070df689e6
[devops] fix extention building (#5427)
9 months ago
binmakeswell 822241a99c
[doc] sora release (#5425)
9 months ago
flybird11111 29695cf70c
[example]add gpt2 benchmark example script. (#5295)
9 months ago
Frank Lee 593a72e4d5
Merge pull request #5424 from FrankLeeeee/sync/main
9 months ago
FrankLeeeee 0310b76e9d Merge branch 'main' into sync/main
9 months ago
Camille Zhong 4b8312c08e
fix sft single turn inference example (#5416)
9 months ago
binmakeswell a1c6cdb189 [doc] fix blog link
9 months ago
binmakeswell 5de940de32 [doc] fix blog link
9 months ago
Frank Lee 2461f37886
[workflow] added pypi channel (#5412)
9 months ago
Tong Li a28c971516
update requirements (#5407)
9 months ago
yuehuayingxueluo 0aa27f1961
[Inference]Move benchmark-related code to the example directory. (#5408)
9 months ago
yuehuayingxueluo 600881a8ea
[Inference]Add CUDA KVCache Kernel (#5406)
9 months ago
flybird11111 0a25e16e46
[shardformer]gather llama logits (#5398)
9 months ago
Frank Lee dcdd8a5ef7
[setup] fixed nightly release (#5388)
9 months ago
QinLuo bf34c6fef6
[fsdp] impl save/load shard model/optimizer (#5357)
9 months ago
Hongxin Liu d882d18c65
[example] reuse flash attn patch (#5400)
9 months ago
Hongxin Liu 95c21e3950
[extension] hotfix jit extension setup (#5402)
9 months ago
Yuanheng Zhao 19061188c3
[Infer/Fix] Fix Dependency in test - RMSNorm kernel (#5399)
9 months ago
yuehuayingxueluo bc1da87366
[Fix/Inference] Fix format of input prompts and input model in inference engine (#5395)
9 months ago
yuehuayingxueluo 2a718c8be8
Optimized the execution interval time between cuda kernels caused by view and memcopy (#5390)
9 months ago
Jianghai 730103819d
[Inference]Fused kv copy into rotary calculation (#5383)
9 months ago
Stephan Kölker 5d380a1a21
[hotfix] Fix wrong import in meta_registry (#5392)
9 months ago
CZYCW b833153fd5
[hotfix] fix variable type for top_p (#5313)
9 months ago
Yuanheng Zhao b21aac5bae
[Inference] Optimize and Refactor Inference Batching/Scheduling (#5367)
9 months ago
Frank Lee 705a62a565
[doc] updated installation command (#5389)
9 months ago
yixiaoer 69e3ad01ed
[doc] Fix typo (#5361)
9 months ago
Hongxin Liu 7303801854
[llama] fix training and inference scripts (#5384)
9 months ago
Hongxin Liu adae123df3
[release] update version (#5380)
10 months ago
Frank Lee efef43b53c
Merge pull request #5372 from hpcaitech/exp/mixtral
10 months ago
yuehuayingxueluo 8c69debdc7
[Inference]Support vllm testing in benchmark scripts (#5379)
10 months ago
Frank Lee 4c03347fc7
Merge pull request #5377 from hpcaitech/example/llama-npu
10 months ago
Frank Lee 9afa52061f
[inference] refactored config (#5376)
10 months ago
ver217 06db94fbc9 [moe] fix tests
10 months ago
Hongxin Liu 65e5d6baa5 [moe] fix mixtral optim checkpoint (#5344)
10 months ago