ColossalAI/README.md

# Colossal-AI

An integrated large-scale model training system with efficient parallelization techniques.

Paper: [Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training](https://arxiv.org/abs/2110.14883)

Blog: [Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training](https://www.hpcaitech.com/blog)

## Installation

### PyPI

```bash
pip install colossalai
```

### Install From Source

```shell
git clone git@github.com:hpcaitech/ColossalAI.git
cd ColossalAI
# install dependency
pip install -r requirements/requirements.txt

# install colossalai
pip install .
```

Install and enable CUDA kernel fusion (compulsory installation when using fused optimizer)

```shell
pip install -v --no-cache-dir --global-option="--cuda_ext" .
```

## Documentation

- [Documentation](https://www.colossalai.org/)

## Quick View

### Start Distributed Training in Lines

```python
import colossalai
from colossalai.trainer import Trainer
from colossalai.core import global_context as gpc

engine, train_dataloader, test_dataloader = colossalai.initialize()

trainer = Trainer(engine=engine,
                  verbose=True)
trainer.fit(
    train_dataloader=train_dataloader,
    test_dataloader=test_dataloader,
    epochs=gpc.config.num_epochs,
    hooks_cfg=gpc.config.hooks,
    display_progress=True,
    test_interval=5
)
```

### Write a Simple 2D Parallel Model

Let's say we have a huge MLP model and its very large hidden size makes it difficult to fit into a single GPU. We can
then distribute the model weights across GPUs in a 2D mesh while you still write your model in a familiar way.

```python
from colossalai.nn import Linear2D
import torch.nn as nn


class MLP_2D(nn.Module):

    def __init__(self):
        super().__init__()
        self.linear_1 = Linear2D(in_features=1024, out_features=16384)
        self.linear_2 = Linear2D(in_features=16384, out_features=1024)

    def forward(self, x):
        x = self.linear_1(x)
        x = self.linear_2(x)
        return x

```

## Features

Colossal-AI provides a collection of parallel training components for you. We aim to support you to write your
distributed deep learning models just like how you write your single-GPU model. We provide friendly tools to kickstart
distributed training in a few lines.

- [Data Parallelism](./docs/parallelization.md)
- [Pipeline Parallelism](./docs/parallelization.md)
- [1D, 2D, 2.5D, 3D and sequence parallelism](./docs/parallelization.md)
- [Friendly trainer and engine](./docs/trainer_engine.md)
- [Extensible for new parallelism](./docs/add_your_parallel.md)
- [Mixed Precision Training](./docs/amp.md)
- [Zero Redundancy Optimizer (ZeRO)](./docs/zero.md)

## Cite Us

```
@article{bian2021colossal,
  title={Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training},
  author={Bian, Zhengda and Liu, Hongxin and Wang, Boxiang and Huang, Haichen and Li, Yongbin and Wang, Chuanrui and Cui, Fan and You, Yang},
  journal={arXiv preprint arXiv:2110.14883},
  year={2021}
}
```
fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			`# Colossal-AI`
Migrated project 2021-10-28 16:21:23 +00:00
update documentation 2021-10-29 01:29:20 +00:00			`An integrated large-scale model training system with efficient parallelization techniques.`

fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			`Paper: [Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training](https://arxiv.org/abs/2110.14883)`

			`Blog: [Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training](https://www.hpcaitech.com/blog)`
Migrated project 2021-10-28 16:21:23 +00:00
			`## Installation`

			`### PyPI`

			```bash
			`pip install colossalai`
			```

			`### Install From Source`

			```shell
			`git clone git@github.com:hpcaitech/ColossalAI.git`
			`cd ColossalAI`
			`# install dependency`
			`pip install -r requirements/requirements.txt`

			`# install colossalai`
			`pip install .`
			```

			`Install and enable CUDA kernel fusion (compulsory installation when using fused optimizer)`

			```shell
			`pip install -v --no-cache-dir --global-option="--cuda_ext" .`
			```

			`## Documentation`

			`- [Documentation](https://www.colossalai.org/)`

			`## Quick View`

			`### Start Distributed Training in Lines`

			```python
			`import colossalai`
			`from colossalai.trainer import Trainer`
			`from colossalai.core import global_context as gpc`

Support TP-compatible Torch AMP and Update trainer API (#27) * Add gradient accumulation, fix lr scheduler * fix FP16 optimizer and adapted torch amp with tensor parallel (#18) * fixed bugs in compatibility between torch amp and tensor parallel and performed some minor fixes * fixed trainer * Revert "fixed trainer" This reverts commit 2e0b0b76990e8d4e337add483d878c0f61cf5097. * improved consistency between trainer, engine and schedule (#23) Co-authored-by: 1SAA <c2h214748@gmail.com> Co-authored-by: 1SAA <c2h214748@gmail.com> Co-authored-by: ver217 <lhx0217@gmail.com> 2021-11-18 11:45:06 +00:00			`engine, train_dataloader, test_dataloader = colossalai.initialize()`
Migrated project 2021-10-28 16:21:23 +00:00
			`trainer = Trainer(engine=engine,`
			`verbose=True)`
			`trainer.fit(`
			`train_dataloader=train_dataloader,`
			`test_dataloader=test_dataloader,`
Support TP-compatible Torch AMP and Update trainer API (#27) * Add gradient accumulation, fix lr scheduler * fix FP16 optimizer and adapted torch amp with tensor parallel (#18) * fixed bugs in compatibility between torch amp and tensor parallel and performed some minor fixes * fixed trainer * Revert "fixed trainer" This reverts commit 2e0b0b76990e8d4e337add483d878c0f61cf5097. * improved consistency between trainer, engine and schedule (#23) Co-authored-by: 1SAA <c2h214748@gmail.com> Co-authored-by: 1SAA <c2h214748@gmail.com> Co-authored-by: ver217 <lhx0217@gmail.com> 2021-11-18 11:45:06 +00:00			`epochs=gpc.config.num_epochs,`
			`hooks_cfg=gpc.config.hooks,`
Migrated project 2021-10-28 16:21:23 +00:00			`display_progress=True,`
			`test_interval=5`
			`)`
			```

			`### Write a Simple 2D Parallel Model`

			`Let's say we have a huge MLP model and its very large hidden size makes it difficult to fit into a single GPU. We can`
			`then distribute the model weights across GPUs in a 2D mesh while you still write your model in a familiar way.`

			```python
			`from colossalai.nn import Linear2D`
			`import torch.nn as nn`


			`class MLP_2D(nn.Module):`

			`def __init__(self):`
			`super().__init__()`
			`self.linear_1 = Linear2D(in_features=1024, out_features=16384)`
			`self.linear_2 = Linear2D(in_features=16384, out_features=1024)`

			`def forward(self, x):`
			`x = self.linear_1(x)`
			`x = self.linear_2(x)`
			`return x`

			```

			`## Features`

fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			`Colossal-AI provides a collection of parallel training components for you. We aim to support you to write your`
Migrated project 2021-10-28 16:21:23 +00:00			`distributed deep learning models just like how you write your single-GPU model. We provide friendly tools to kickstart`
			`distributed training in a few lines.`

			`- [Data Parallelism](./docs/parallelization.md)`
			`- [Pipeline Parallelism](./docs/parallelization.md)`
			`- [1D, 2D, 2.5D, 3D and sequence parallelism](./docs/parallelization.md)`
fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			`- [Friendly trainer and engine](./docs/trainer_engine.md)`
Migrated project 2021-10-28 16:21:23 +00:00			`- [Extensible for new parallelism](./docs/add_your_parallel.md)`
			`- [Mixed Precision Training](./docs/amp.md)`
			`- [Zero Redundancy Optimizer (ZeRO)](./docs/zero.md)`

fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			`## Cite Us`
Migrated project 2021-10-28 16:21:23 +00:00
fixed some typos in the documents, added blog link and paper author information in README 2021-11-03 08:07:28 +00:00			```
			`@article{bian2021colossal,`
			`title={Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training},`
			`author={Bian, Zhengda and Liu, Hongxin and Wang, Boxiang and Huang, Haichen and Li, Yongbin and Wang, Chuanrui and Cui, Fan and You, Yang},`
			`journal={arXiv preprint arXiv:2110.14883},`
			`year={2021}`
			`}`
			```