1.
Eight Google researchers trained a model on 8 GPUs for 12 hours and beat every neural translation system ever built — including teams that trained for weeks. The architecture they invented, the Transformer, now powers ChatGPT, Claude, and Gemini. Its secret: throw out everything the field knew about sequence processing and replace it with a single mechanism called attention.