A novel implementation of fusing ViT with Mamba into a fast, agile, and high performance Multi-Modal Model. Powered by Z...
Minimalist NMT for educational purposes
[TNSRE 23] EEG Transformer 2.0. i. Convolutional Transformer for EEG Decoding. ii. Novel visualization - Class Activatio...
Next-generation Video instance recognition framework on top of Detectron2 which supports InstMove (CVPR 2023), SeqFormer...
CTR prediction models based on deep learning(基于深度学习的广告推荐CTR预估模型)
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
Graph Transformer Architecture. Source code for "A Generalization of Transformer Networks to Graphs", DLG-AAAI'21.
Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting (NeurIPS 2019)
list of efficient attention modules
第 1-9 条,共 9 条