Foundation Models · since 2023

Multimodal

Single models that jointly process text, images, audio, and video under one attention-based backbone — collapsing what used to be separate model families into one. The capability that turned LLMs into general perception+generation systems.

17

events traced

17

source records

30 Jun 2026

first signal

9 Jul 2026

last activity

Who drove it

Google DeepMind32%
OpenAI30%
Meta AI20%
open ecosystem18%

Key movements

4+

modalities under one backbone

2024+

native (not bolted-on) multimodal

The story, event by event

Every point below is traced to a real source — nothing on this page is invented.

  1. 5 Jan 2021impact 68

    CLIP — connecting text and images

    value: CLIP; metric: precursor

  2. 25 Sept 2023impact 72

    GPT-4V brings vision to the frontier

    value: vision; metric: modality

  3. 6 Dec 2023impact 74

    Gemini — natively multimodal from the start

    value: native; metric: design

  4. 13 May 2024impact 68

    Real-time voice + video (GPT-4o)

    value: real-time; metric: latency

  5. 26 Jun 2025impact 60

    Release of Gemma 3n Model

    features: Per-Layer Embeddings, KV Cache Sharing, MobileNet-V5-300M; architecture: MatFormer; target audience: developers

  6. 18 Nov 2025impact 70

    Launch of Gemini 3 with advanced reasoning and multimodal capabilities

    user engagement: 2 billion monthly users; leaderboard score: 1501 Elo on LMArena; safety evaluations: most comprehensive set of safety evaluations; reasoning capabilities: state-of-the-art reasoning and multimodal capabilities

  7. 20 Nov 2025impact 60

    Introduction of Nano Banana Pro

    model: Nano Banana Pro; features: advanced reasoning, real-time information integration, improved text rendering; base model: Gemini 3 Pro; applications: image generation, editing, content creation

  8. 20 Nov 2025impact 60

    Release of Nano Banana Pro (Gemini 3 Pro Image)

    features: studio-quality image generation, improved text rendering, integration with Google Search; model name: Gemini 3 Pro Image; benchmark performance: excels on Text to Image AI benchmarks

  9. 12 Dec 2025impact 60

    Enhanced Gemini 2.5 Flash Native Audio Models Released

    features: live speech-to-speech translation, enhanced voice interaction capabilities; performance metrics: {'ComplexFuncBench_Audio': '71.5%', 'adherence_to_instructions': {'current': '90%', 'previous': '84%'}}

  10. 5 Mar 2026impact 60

    Deployment of VLA Models on Embedded Platforms

    strategies: Architectural decomposition, Latency-aware scheduling, Hardware-aligned execution; key argument: Bringing VLA models to embedded platforms requires addressing complex systems engineering challenges rather than just model compression.

  11. 9 Mar 2026impact 60

    Introduction of Ulysses Sequence Parallelism

    efficiency: SP=4 processes 13,396 tokens/second at 64K tokens, outperforming the baseline by 3.7x.

  12. 25 Mar 2026impact 60

    Lyria 3 Pro Public Preview Launch

    model: Lyria 3 Pro; features: public preview on Vertex AI, available in AI Studio, collaboration enhancement with ProducerAI, AI integration for creative workflows, outputs embedded with SynthID watermark

  13. 9 Apr 2026impact 60

    Advancements in Multimodal Embedding and Reranker Models

    model: Qwen3-VL-2B; description: The article discusses the implementation of multimodal embedding and reranker models using Sentence Transformers, enhancing cross-modal capabilities.; gpu requirement: 8 GB VRAM

  14. 7 May 2026impact 65

    OpenAI introduces new real-time audio models (GPT-Realtime-2, Translate, Whisper)

    OpenAI introduces three new real-time audio models, GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper, significantly enhancing voice application development.

    GPT-Realtime-2 offers GPT-5-class reasoning, enabling it to handle complex requests and maintain natural conversations.

    GPT-Realtime-Translate provides live translation from over 70 input languages into 13 output languages.

    GPT-Realtime-Whisper delivers live speech-to-text transcription.

    GPT-Realtime-2 demonstrates a 15.2% improvement in audio intelligence compared to its predecessor.

    These advancements push the frontier of real-time multimodal interaction, making voice a more integrated and capable interface for software.

  15. 7 Jul 2026impact 50

    Launch of Muse Image Model

    features: @ mention integration; developer: Meta; model name: Muse Image; integration platforms: Instagram, WhatsApp

  16. 8 Jul 2026impact 60

    Integration of Transformers into vLLM

    details: {'integration': 'Transformers as a modeling backend in vLLM', 'performance': 'Native vLLM speeds without additional coding', 'optimizations': ['Dynamic layer fusions', 'Static analysis via torch.fx']}; description: Enhancements in inference speed and usability for model authors.

  17. 8 Jul 2026impact 60

    Launch of GPT-Live Voice Model

    architecture: full-duplex; backend model: GPT-5.5; model versions: GPT-Live-1, GPT-Live-1 mini

Lineage

Descended from

This is one node. The map holds the whole field.

Watch Multimodal — and everything it connects to — grow as real news threads onto the map every day.