1.
Delete 72% of a transformer's parameters — the feed-forward network — and the model loses 0.47 nats. Give those parameters back as attention depth, and the gap shrinks to 0.006 nats, agreeing to 10⁻⁴ across clean seeds. The FFN is not a computational primitive. It is a write path.