Creative compression
Multimodal language models have fundamentally compressed the software development lifecycle. The traditional sequence (product manager writes requirements, designer creates wireframes, engineer builds prototype, team gathers feedback, repeat) has collapsed. One startup product exec told their PMs they were done with specs and PRDs: prototypes only. This hasn’t eliminated jobs but dramatically accelerated development from months to hours.
ChatGPT’s unified image generation model extends this compression beyond software. A strip mall investor visualized a remodel by uploading a photo of a property he’d just bought and asking for cosmetic concepts, cutting months of architect back-and-forth down to days. Someone else assembled a CPG brand’s omnichannel marketing campaign with no graphic designer involved. Google’s Gemini had already shown the same trick for interior decoration: one user asked it to clear the furniture out of a room photo and redecorate it, and got on the first try what an interior designer would charge thousands for.
What we’re witnessing is a Cambrian explosion of use cases through native multimodal models that generate language and images in one system. The outer boundaries remain unclear, but they suggest a profound shift.
The strategy of building single, larger unified models rather than domain-specific ones aligns with Richard Sutton’s “Bitter Lesson”, and the results so far back it up. This trajectory points toward models that will natively produce video, 3D models, audio, complete applications, CAD drawings, and virtually any digital media, where each new capability compounds on existing world knowledge captured in the models’ latent space.
The full implications remain elusive, but we’re clearly experiencing another ChatGPT-level inflection point in creative production.

ChatGPT adds 1 million users in 1 hour (31 March 2025)
This inflection point is particularly thrilling for those of us building at the application layer in visual communication. Companies that create intuitive interfaces between these unified models and everyday users stand to transform how we communicate visually. The acceleration isn’t only about production speed. It also democratizes sophisticated visual creation. By abstracting away technical complexity while expanding creative possibilities, we’re entering an era where the bottleneck shifts from technical ability to imagination. For anyone passionate about amplifying human creativity, there has never been a more exciting time to build.