Multimodal

5 pieces on Multimodal.

News & Analysis

news
Shan2026-09-21
Open WeightsImage GenerationAlibabaDiffusion ModelsMultimodal

Qwen-Image-2.1: 7B Model Beats 32B FLUX 2 Max on Qwen's Benchmark

Alibaba's Qwen-Image-2.1 unifies image generation and editing in a 7B diffusion transformer, scoring 60.28 vs FLUX 2 Max's 55.33 at a quarter the parameters.

Read more
news
Shan2026-09-06
MultimodalDocument RetrievalOpen WeightsEncodersH Company

NeoMME 260M Matches ColQwen2.5 3.75B on ViDoRe v3

H Company's NeoMME encodes text and raw image patches through one bidirectional Transformer, matching a 3.75B model at 260M parameters.

Read more
news
Shan2026-08-26
Open WeightsMixture of ExpertsAlibabaMultimodalInference Efficiency

Qwen3.8-Flash-Next: 125B MoE Runs at 6B Active Params, Previews Qwen4

Alibaba's Qwen team releases a 180B-on-disk multimodal MoE that activates only 6B parameters per token, trained at one-ninth the cost of Qwen3.7-Plus.

Read more
news
Shan2026-08-12
Open WeightsMultimodalComfyUIVideo GenerationTutorials

MiniMax-H3 Video Pipeline via ComfyUI APIs: A Reference Implementation

A headless Python pipeline drives MiniMax-H3 video and audio generation through ComfyUI HTTP and WebSocket APIs, with VRAM-tiered model selection.

Read more
Shieldstral 1.0 3B: Mistral's Policy-Adaptive Multimodal Safety Classifier
news
Shan2026-08-09
Open WeightsMistral AIContent ModerationMultimodalSafety

Shieldstral 1.0 3B: Mistral's Policy-Adaptive Multimodal Safety Classifier

Mistral's 3B Apache 2.0 safety classifier matches a 20B model on text F1 and leads all evaluated baselines on multimodal safety, with policy set at inference time.

Read more