Commit Graph

2 Commits (2cddeac7174c5617b7a35cd83925a161173afe1b)

Author SHA1 Message Date
hxwang 74eccac0db [moe] test deepseek 2024-08-01 10:06:59 +08:00
botbw 9b9b76bdcd [moe] add mixtral dp grad scaling when not all experts are activated 2024-08-01 10:06:59 +08:00