Skip to main content
tech futures.Introduction to AI

Lesson 5: Under the Hood

Transformer Architecture

Before the Transformer Architecture was introduced in a 2017 Google research paper: "Attention Is All You Need", machine learning models read words one at a time and could not remember relationships with words very far away in the sentence.

Extract from Google Research

They had limited context.

The breakthrough was that Transformers could look at ALL the words at once and work out how each word relates to the others, even distant words.

The particular structure or architecture in the Transformers allowed for many different operations to happen all at the same time or "in parallel".