ColossalAI

Commit Graph

Author	SHA1	Message	Date
Runyu Lu	9fe61b4475	[fix]	2024-03-25 11:37:58 +08:00
Yuanheng Zhao	5fcd7795cd	[example] update Grok-1 inference (#5495 ) * revise grok-1 example * remove unused arg in scripts * prevent re-installing torch * update readme * revert modifying colossalai requirements * add perf * trivial * add tokenizer url	2024-03-24 20:24:11 +08:00
binmakeswell	6df844b8c4	[release] grok-1 314b inference (#5490 ) * [release] grok-1 inference * [release] grok-1 inference * [release] grok-1 inference	2024-03-22 15:48:12 +08:00
Hongxin Liu	848a574c26	[example] add grok-1 inference (#5485 ) * [misc] add submodule * remove submodule * [example] support grok-1 tp inference * [example] add grok-1 inference script * [example] refactor code * [example] add grok-1 readme * [exmaple] add test ci * [exmaple] update readme	2024-03-21 18:07:22 +08:00
Runyu Lu	5b017d6324	[fix]	2024-03-21 15:55:25 +08:00
Runyu Lu	606603bb88	Merge branch 'feature/colossal-infer' of https://github.com/hpcaitech/ColossalAI into colossal-infer-cuda-graph	2024-03-21 14:25:22 +08:00
Runyu Lu	4eafe0c814	[fix] unused option	2024-03-21 11:28:42 +08:00
binmakeswell	d158fc0e64	[doc] update open-sora demo (#5479 ) * [doc] update open-sora demo * [doc] update open-sora demo * [doc] update open-sora demo	2024-03-20 16:08:41 +08:00
傅剑寒	7ff42cc06d	add vec_type_trait implementation (#5473 )	2024-03-19 18:36:40 +08:00
傅剑寒	b96557b5e1	Merge pull request #5469 from Courtesy-Xs/add_vec_traits Refactor vector utils	2024-03-19 13:53:26 +08:00
Runyu Lu	aabc9fb6aa	[feat] add use_cuda_kernel option	2024-03-19 13:24:25 +08:00
xs_courtesy	48c4f29b27	refactor vector utils	2024-03-19 11:32:01 +08:00
binmakeswell	bd998ced03	[doc] release Open-Sora 1.0 with model weights (#5468 ) * [doc] release Open-Sora 1.0 with model weights * [doc] release Open-Sora 1.0 with model weights * [doc] release Open-Sora 1.0 with model weights	2024-03-18 18:31:18 +08:00
flybird11111	5e16bf7980	[shardformer] fix gathering output when using tensor parallelism (#5431 ) * fix * padding vocab_size when using pipeline parallellism padding vocab_size when using pipeline parallellism fix fix * fix * fix fix fix * fix gather output * fix * fix * fix fix resize embedding fix resize embedding * fix resize embedding fix * revert * revert * revert	2024-03-18 15:55:11 +08:00
傅剑寒	b6e9785885	Merge pull request #5457 from Courtesy-Xs/ly_add_implementation_for_launch_config add implementatino for GetGPULaunchConfig1D	2024-03-15 11:23:44 +08:00
xs_courtesy	5724b9e31e	add some comments	2024-03-15 11:18:57 +08:00
Runyu Lu	6e30248683	[fix] tmp for test	2024-03-14 16:13:00 +08:00
xs_courtesy	388e043930	add implementatino for GetGPULaunchConfig1D	2024-03-14 11:13:40 +08:00
Runyu Lu	d02e257abd	Merge branch 'feature/colossal-infer' into colossal-infer-cuda-graph	2024-03-14 10:37:05 +08:00
Runyu Lu	ae24b4f025	diverse tests	2024-03-14 10:35:08 +08:00
Runyu Lu	1821a6dab0	[fix] pytest and fix dyn grid bug	2024-03-13 17:28:32 +08:00
yuehuayingxueluo	f366a5ea1f	[Inference/kernel]Add Fused Rotary Embedding and KVCache Memcopy CUDA Kernel (#5418 ) * add rotary embedding kernel * add rotary_embedding_kernel * add fused rotary_emb and kvcache memcopy * add fused_rotary_emb_and_cache_kernel.cu * add fused_rotary_emb_and_memcopy * fix bugs in fused_rotary_emb_and_cache_kernel.cu * fix ci bugs * use vec memcopy and opt the gloabl memory access * fix code style * fix test_rotary_embdding_unpad.py * codes revised based on the review comments * fix bugs about include path * rm inline	2024-03-13 17:20:03 +08:00
Steve Luo	ed431de4e4	fix rmsnorm template function invocation problem(template function partial specialization is not allowed in Cpp) and luckily pass e2e precision test (#5454 )	2024-03-13 16:00:55 +08:00
Hongxin Liu	f2e8b9ef9f	[devops] fix compatibility (#5444 ) * [devops] fix compatibility * [hotfix] update compatibility test on pr * [devops] fix compatibility * [devops] record duration during comp test * [test] decrease test duration * fix falcon	2024-03-13 15:24:13 +08:00
傅剑寒	6fd355a5a6	Merge pull request #5452 from Courtesy-Xs/fix_include_path fix include path	2024-03-13 11:26:41 +08:00
xs_courtesy	c1c45e9d8e	fix include path	2024-03-13 11:21:06 +08:00
Steve Luo	b699f54007	optimize rmsnorm: add vectorized elementwise op, feat loop unrolling (#5441 )	2024-03-12 17:48:02 +08:00
傅剑寒	368a2aa543	Merge pull request #5445 from Courtesy-Xs/refactor_infer_compilation Refactor colossal-infer code arch	2024-03-12 14:14:37 +08:00
digger yu	385e85afd4	[hotfix] fix typo s/keywrods/keywords etc. (#5429 )	2024-03-12 11:25:16 +08:00
xs_courtesy	095c070a6e	refactor code	2024-03-11 17:06:57 +08:00
Camille Zhong	da885ed540	fix tensor data update for gemini loss caluculation (#5442 )	2024-03-11 13:49:58 +08:00
傅剑寒	21e1e3645c	Merge pull request #5435 from Courtesy-Xs/add_gpu_launch_config Add query and other components	2024-03-11 11:15:29 +08:00
Runyu Lu	633e95b301	[doc] add doc	2024-03-11 10:56:51 +08:00
Runyu Lu	9dec66fad6	[fix] multi graphs capture error	2024-03-11 10:51:16 +08:00
Runyu Lu	b2c0d9ff2b	[fix] multi graphs capture error	2024-03-11 10:49:31 +08:00
Steve Luo	f7aecc0c6b	feat rmsnorm cuda kernel and add unittest, benchmark script (#5417 )	2024-03-08 16:21:12 +08:00
xs_courtesy	5eb5ff1464	refactor code	2024-03-08 15:41:14 +08:00
xs_courtesy	01d289d8e5	Merge branch 'feature/colossal-infer' of https://github.com/hpcaitech/ColossalAI into add_gpu_launch_config	2024-03-08 15:04:55 +08:00
xs_courtesy	a46598ac59	add reusable utils for cuda	2024-03-08 14:53:29 +08:00
傅剑寒	2b28b54ac6	Merge pull request #5433 from Courtesy-Xs/add_silu_and_mul 【Inference】Add silu_and_mul for infer	2024-03-08 14:44:37 +08:00
Runyu Lu	cefaeb5fdd	[feat] cuda graph support and refactor non-functional api	2024-03-08 14:19:35 +08:00
Hongxin Liu	8020f42630	[release] update version (#5411 )	2024-03-07 23:36:07 +08:00
xs_courtesy	95c21498d4	add silu_and_mul for infer	2024-03-07 16:57:49 +08:00
Camille Zhong	743e7fad2f	[colossal-llama2] add stream chat examlple for chat version model (#5428 ) * add stream chat for chat version * remove os.system clear * modify function name	2024-03-07 14:58:56 +08:00
Youngon	68f55a709c	[hotfix] fix stable diffusion inference bug. (#5289 ) * Update train_ddp.yaml delete "strategy" to fix DDP config loading bug in "main.py" * Update train_ddp.yaml fix inference with scripts/txt2img.py config file load bug. * Update README.md add pretrain model test code.	2024-03-05 22:03:40 +08:00
hugo-syn	c8003d463b	[doc] Fix typo s/infered/inferred/ (#5288 ) Signed-off-by: hugo-syn <hugo.vincent@synacktiv.com>	2024-03-05 22:02:08 +08:00
digger yu	5e1c93d732	[hotfix] fix typo change MoECheckpintIO to MoECheckpointIO (#5335 ) Co-authored-by: binmakeswell <binmakeswell@gmail.com>	2024-03-05 21:52:30 +08:00
Dongruixuan Li	a7ae2b5b4c	[eval-hotfix] set few_shot_data to None when few shot is disabled (#5422 )	2024-03-05 21:48:55 +08:00
digger yu	049121d19d	[hotfix] fix typo change enabel to enable under colossalai/shardformer/ (#5317 )	2024-03-05 21:48:46 +08:00
digger yu	16c96d4d8c	[hotfix] fix typo change _descrption to _description (#5331 )	2024-03-05 21:47:48 +08:00

... 2 3 4 5 6 ...

3294 Commits (c06208e72c35d74e150b6a83e72375f5021d10b1) All Branches Search

3294 Commits (c06208e72c35d74e150b6a83e72375f5021d10b1)

All Branches