2 posts found
A deep read of the Kimi K3 technical report: a 2.8T parameter open MoE with 104B activated …
A deep dive comparing standard softmax attention, linear attention, and Flash Attention: their math, …