Skip to main content
tech futures.Introduction to AI

Lesson 2: Different Types of AI

Multimodal AI

We have looked at some different modes of AI in this lesson: text, images, video, audio and code. However, most popular AI tools these days are multi-modal. That means they support more than just a single mode.

For example, we could use Google's Antigravity code editor to generate a website for us and it could produce:

  • Text for the website
  • Images for the website
  • Code to make the website work
  • Video and audio for the website

The AI tools that you use (including Lumen) are usually multimodal. This means that we can use AI in a variety of ways to help us solve problems or make improvements to things we are working on.

For example, imagine an app with multimodal AI features designed to help a basketball player improve their game. The table below summarises different features the app could have and the AI mode that feature uses.

ModeFeature
VideoThe player films a shot and the AI analyses arm angle, jump height, and release point.
ImagesIt compares key frames to photos of professional players.
AudioThe player asks, “How can I fix my shooting form?” and the AI understands.
CodeIt generates a small graph showing accuracy over time.
Output videoIt creates an overlay video with arrows showing mistakes and a side-by-side comparison with a pro.

Some example screenshots of what this app could look like (generated with Gemini) are shown in the image below.

Example screenshots of a basketball game improvement app, generated with Gemini