Multimodal
5 pieces on Multimodal.
News & Analysis
Qwen-Image-2.1: 7B Model Beats 32B FLUX 2 Max on Qwen's Benchmark
Alibaba's Qwen-Image-2.1 unifies image generation and editing in a 7B diffusion transformer, scoring 60.28 vs FLUX 2 Max's 55.33 at a quarter the parameters.
NeoMME 260M Matches ColQwen2.5 3.75B on ViDoRe v3
H Company's NeoMME encodes text and raw image patches through one bidirectional Transformer, matching a 3.75B model at 260M parameters.
Qwen3.8-Flash-Next: 125B MoE Runs at 6B Active Params, Previews Qwen4
Alibaba's Qwen team releases a 180B-on-disk multimodal MoE that activates only 6B parameters per token, trained at one-ninth the cost of Qwen3.7-Plus.
MiniMax-H3 Video Pipeline via ComfyUI APIs: A Reference Implementation
A headless Python pipeline drives MiniMax-H3 video and audio generation through ComfyUI HTTP and WebSocket APIs, with VRAM-tiered model selection.

Shieldstral 1.0 3B: Mistral's Policy-Adaptive Multimodal Safety Classifier
Mistral's 3B Apache 2.0 safety classifier matches a 20B model on text F1 and leads all evaluated baselines on multimodal safety, with policy set at inference time.