Foundation Models · since 2023
Multimodal
Single models that jointly process text, images, audio, and video under one attention-based backbone — collapsing what used to be separate model families into one. The capability that turned LLMs into general perception+generation systems.
17
events traced
17
source records
30 Jun 2026
first signal
9 Jul 2026
last activity
Who drove it
Key movements
4+↑
modalities under one backbone
2024+↑
native (not bolted-on) multimodal
The story, event by event
Every point below is traced to a real source — nothing on this page is invented.
5 Jan 2021impact 68
CLIP — connecting text and images
value: CLIP; metric: precursor
25 Sept 2023impact 72
GPT-4V brings vision to the frontier
value: vision; metric: modality
6 Dec 2023impact 74
Gemini — natively multimodal from the start
value: native; metric: design
13 May 2024impact 68
Real-time voice + video (GPT-4o)
value: real-time; metric: latency
26 Jun 2025impact 60
Release of Gemma 3n Model
features: Per-Layer Embeddings, KV Cache Sharing, MobileNet-V5-300M; architecture: MatFormer; target audience: developers
18 Nov 2025impact 70
Launch of Gemini 3 with advanced reasoning and multimodal capabilities
user engagement: 2 billion monthly users; leaderboard score: 1501 Elo on LMArena; safety evaluations: most comprehensive set of safety evaluations; reasoning capabilities: state-of-the-art reasoning and multimodal capabilities
20 Nov 2025impact 60
Introduction of Nano Banana Pro
model: Nano Banana Pro; features: advanced reasoning, real-time information integration, improved text rendering; base model: Gemini 3 Pro; applications: image generation, editing, content creation
20 Nov 2025impact 60
Release of Nano Banana Pro (Gemini 3 Pro Image)
features: studio-quality image generation, improved text rendering, integration with Google Search; model name: Gemini 3 Pro Image; benchmark performance: excels on Text to Image AI benchmarks
12 Dec 2025impact 60
Enhanced Gemini 2.5 Flash Native Audio Models Released
features: live speech-to-speech translation, enhanced voice interaction capabilities; performance metrics: {'ComplexFuncBench_Audio': '71.5%', 'adherence_to_instructions': {'current': '90%', 'previous': '84%'}}
5 Mar 2026impact 60
Deployment of VLA Models on Embedded Platforms
strategies: Architectural decomposition, Latency-aware scheduling, Hardware-aligned execution; key argument: Bringing VLA models to embedded platforms requires addressing complex systems engineering challenges rather than just model compression.
9 Mar 2026impact 60
Introduction of Ulysses Sequence Parallelism
efficiency: SP=4 processes 13,396 tokens/second at 64K tokens, outperforming the baseline by 3.7x.
25 Mar 2026impact 60
Lyria 3 Pro Public Preview Launch
model: Lyria 3 Pro; features: public preview on Vertex AI, available in AI Studio, collaboration enhancement with ProducerAI, AI integration for creative workflows, outputs embedded with SynthID watermark
9 Apr 2026impact 60
Advancements in Multimodal Embedding and Reranker Models
model: Qwen3-VL-2B; description: The article discusses the implementation of multimodal embedding and reranker models using Sentence Transformers, enhancing cross-modal capabilities.; gpu requirement: 8 GB VRAM
7 May 2026impact 65
OpenAI introduces new real-time audio models (GPT-Realtime-2, Translate, Whisper)
OpenAI introduces three new real-time audio models, GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper, significantly enhancing voice application development.
GPT-Realtime-2 offers GPT-5-class reasoning, enabling it to handle complex requests and maintain natural conversations.
GPT-Realtime-Translate provides live translation from over 70 input languages into 13 output languages.
GPT-Realtime-Whisper delivers live speech-to-text transcription.
GPT-Realtime-2 demonstrates a 15.2% improvement in audio intelligence compared to its predecessor.
These advancements push the frontier of real-time multimodal interaction, making voice a more integrated and capable interface for software.
7 Jul 2026impact 50
Launch of Muse Image Model
features: @ mention integration; developer: Meta; model name: Muse Image; integration platforms: Instagram, WhatsApp
8 Jul 2026impact 60
Integration of Transformers into vLLM
details: {'integration': 'Transformers as a modeling backend in vLLM', 'performance': 'Native vLLM speeds without additional coding', 'optimizations': ['Dynamic layer fusions', 'Static analysis via torch.fx']}; description: Enhancements in inference speed and usability for model authors.
8 Jul 2026impact 60
Launch of GPT-Live Voice Model
architecture: full-duplex; backend model: GPT-5.5; model versions: GPT-Live-1, GPT-Live-1 mini
Lineage
Descended from
This is one node. The map holds the whole field.
Watch Multimodal — and everything it connects to — grow as real news threads onto the map every day.