2024 is the year of multimodal AI
Retroactive Grade (February 2025): Correct.
2023: a year of triumph for natural language interfaces
Last year, I predicted that 2023 would be the year of natural language interfaces, and in retrospect it looks like a crazy conservative guess. The tech world has been abuzz with Copilots, “ChatGPT for X”, and the announcement of ChatGPT’s GPT Store. It’s been a rollercoaster since ChatGPT came onto the scene about 13 months ago.

Monthly web visits to ChatGPT and its competitors through October 2023. SimilarWeb data, chart by Shane Burke.
The dawn of multimodal AI in 2024
Imagine designing a building, composing a symphony, or planning a health regimen, all from a few photos or spoken words. That’s the imminent reality of multimodal AI. My prediction for 2024 is that multimodal generative AI will go mainstream. In a nutshell, multimodal AI means moving beyond text in generative AI to inputs and outputs in other formats: images, video, and audio.
Recent multimodal explorations
In recent months, ChatGPT, Bard, and Bing have all added multimodal features to their chatbots, and I’ve been playing with them a lot since then. Some of my favorite use cases so far:
- App development: While helping a friend with their app, I input UX wireframes and a brief description, and out came a UML diagram for the required data model.
- Creative assistance: Working on a holiday card, I fed the AI a color palette and design brief, and it suggested compatible colors and palette alterations.
- Real-world perception: Unsure of what to buy for a new recipe, I uploaded a recipe screenshot and a fridge photo, and the AI listed the ingredients I needed.
- Visual understanding: At a French restaurant, I snapped a photo of the menu along with my dietary preferences, and the AI recommended the perfect dishes.
The unleashing of multimodal foundation models
The real shift will be the widespread availability of multimodal foundation models via APIs in early 2024, with fine-tuning to follow by mid-year. As a product builder and software developer, I anticipate 4 widespread use cases for these APIs:
- Image classification: Input an image corpus with labels, get text classifications for each image.
- Design advice: Submit a photo of a bedroom and remodeling instructions, receive interior design suggestions.
- Creative augmentation: Feed a low-fidelity CAD file with a brief to envision a building, and get multiple floor plan variations.
- Real-world intelligence: Input a work site photo with a request to identify safety hazards, and receive a list of potential risks.
Beyond the text box: a renaissance in human-computer interaction
Interactions with software are getting more conversational and more domain-specific. We’re moving from plain “natural language” to conversations that look a lot more like how professionals actually interact with each other.
A Cambrian explosion of creativity
Text-to-text generative AI amplified creativity across numerous fields. I expect multimodal generative AI to spark a Cambrian explosion of creative applications across an even broader set of industries and disciplines.