Multimodal AI
7 pieces on Multimodal AI.
News & Analysis
Gemini Omni 1.1 Flash: 40s Extension, First/Last Frame, and 4K Upscaling
Google ships gemini-omni-1.1-flash with 10s context reads, first/last frame interpolation, 360p draft mode at ⅓ cost, and upscaled 4K output.
Google DeepMind SL2T Brings Sign Language Dictation to Pixel 11
SL2T scores 70 BLEURT zero-shot on FLEURS-ASL, ships free in Gboard and Live Transcribe on Pixel 11 across 50+ sign languages.

Google DeepMind SL2T: Sign Language to Text at Consumer Scale
SL2T scores 70 BLEURT zero-shot on FLEURS-ASL, powering Gboard and Live Transcribe on Pixel 11 with ASL-to-English translation.

SeedRealtime: ByteDance's Native Audio-Visual Full-Duplex LLM
ByteDance's SeedRealtime fuses audio, video, and text in one end-to-end model, replacing cascaded ASR→VLM→LLM→TTS pipelines with parallel perception and speech.
Alibaba Launches Qwen3.7-Plus With Vision, Tool Use, and Agentic Iteration
Qwen3.7-Plus brings image and video understanding, deep reasoning, tool invocation, verification loops, and autonomous iteration to Alibaba's Bailian platform.
The Rise of Multi-Modal Reasoning in Next-Gen LLMs
How native multi-modality trained from the ground up is revolutionizing application development and changing the way we interact with AI.
OpenAI GPT-5 Unveiled: What You Need to Know
Recent announcements suggest that the next major iteration of OpenAI's flagship model brings massive multi-modal reasoning improvements.