ColossalAI

History

Jianghai 730103819d [Inference]Fused kv copy into rotary calculation (#5383 ) * revise rotary embedding * remove useless print * adapt * fix * add * fix * modeling * fix * fix * fix * fused kv copy * fused copy * colossalai/kernel/triton/no_pad_rotary_embedding.py * del padding llama * del		2024-02-21 11:31:48 +08:00
..
test_models	[Infer] Optimize Blocked KVCache And Kernels Using It (#5325 )	2024-01-30 16:06:09 +08:00
test_ops/triton	[Inference]Fused kv copy into rotary calculation (#5383 )	2024-02-21 11:31:48 +08:00
_utils.py	[Inference] Add the logic of the inference engine (#5173 )	2024-01-11 13:39:56 +00:00
test_batch_bucket.py	[Inference] Optimize and Refactor Inference Batching/Scheduling (#5367 )	2024-02-19 17:18:20 +08:00
test_config_and_struct.py	[Inference] Optimize and Refactor Inference Batching/Scheduling (#5367 )	2024-02-19 17:18:20 +08:00
test_inference_engine.py	[Inference] User Experience: update the logic of default tokenizer and generation config. (#5337 )	2024-02-07 17:55:48 +08:00
test_kvcache_manager.py	[Inference] Optimize and Refactor Inference Batching/Scheduling (#5367 )	2024-02-19 17:18:20 +08:00
test_request_handler.py	[Inference] Optimize and Refactor Inference Batching/Scheduling (#5367 )	2024-02-19 17:18:20 +08:00