While Linear and Sparse Attention promise efficiency, real-world implementation reveals hidden costs and performance trade-offs that make Full Attention the safer choice for now.
Evolutionary algorithms discover a novel attention mechanism that outperforms standard transformers by 4% on WikiText-2, using sparsemax and output gating instead of softmax.