
On the Necessity of Attention–FFN Split in Vision Transformers
AttenFeed unifies Attention and FFN into a single module. Building a ViT entirely from it (uViT) reveals that the rigid Attention–FFN split of standard ViTs is an inductive bias that hinders learning at smaller model scales.








