Machine learning

Scaling transformer to 1m tokens and beyond with RMT

M. Burtsev, A. Bulatov, Y. Kuratov

Proceedings of the AAAI Conference 38, 16 (2025)

The quadratic complexity of attention in transformers is tackled by combining token-based memory and segment-level recurrence, using RMT.

Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
LCP
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"
Image for the paper "Scaling transformer to 1m tokens and beyond with RMT"