2 posts found
A deep read of the Kimi K3 technical report: a 2.8T parameter open MoE with 104B activated …
Kimi's Attention Residuals paper proposes replacing fixed residual connections with learned softmax …