Tags
2 个页面
Transformer
为什么Scale-Dot Attention的分母是根号d
论文阅读:Rethinking Positional Encoding In Language Pre-Training