Lesson 5: Under the Hood
Transformer Architecture
Before the Transformer Architecture was introduced in a 2017 Google research paper: "Attention Is All You Need", machine learning models read words one at a time and could not remember relationships with words very far away in the sentence.
They had limited context.
The breakthrough was that Transformers could look at ALL the words at once and work out how each word relates to the others, even distant words.
The particular structure or architecture in the Transformers allowed for many different operations to happen all at the same time or "in parallel".