Dharma HallVolume IIIThe Yogācāra Positioning of Large Language Models

"Attention Mechanism" Is Not Attention

Vaswani et al.'s 2017 paper was titled "Attention Is All You Need." The naming caused immense conceptual confusion. Transformer attention is mathematics: weighting tokens in the input sequence to determine output influence. It doesn't "want" to attend to anything. It distributes weights per parameters. Yogācāra's manaskāra is the mind's "active directedness" — with intentionality and orientation. One is passive computation; the other is active cognition. Same name, entirely different nature.