Lesson 2: Different Types of AI
Multimodal AI
We have looked at some different modes of AI in this lesson: text, images, video, audio and code. However, most popular AI tools these days are multi-modal. That means they support more than just a single mode.
For example, we could use Google's Antigravity code editor to generate a website for us and it could produce:
- Text for the website
- Images for the website
- Code to make the website work
- Video and audio for the website
The AI tools that you use (including Lumen) are usually multimodal. This means that we can use AI in a variety of ways to help us solve problems or make improvements to things we are working on.
For example, imagine an app with multimodal AI features designed to help a basketball player improve their game. The table below summarises different features the app could have and the AI mode that feature uses.
| Mode | Feature |
|---|---|
| Video | The player films a shot and the AI analyses arm angle, jump height, and release point. |
| Images | It compares key frames to photos of professional players. |
| Audio | The player asks, “How can I fix my shooting form?” and the AI understands. |
| Code | It generates a small graph showing accuracy over time. |
| Output video | It creates an overlay video with arrows showing mistakes and a side-by-side comparison with a pro. |
Some example screenshots of what this app could look like (generated with Gemini) are shown in the image below.