Multimodal Models
Models that work with multiple data types (text, image, audio, video)
2 models in this category
MiniMax M3
MiniMax AI
MiniMax M3 is a frontier-tier open-weight multimodal agent model with 428B total parameters (23B activated) and a 1M token context window. Released May 31, 2026, it features MiniMax Sparse Attention (MSA) for 9x prefill and 15x decode speedup over its predecessor, and is specialized for coding, agentic tasks, and complex reasoning.
NVIDIA Cosmos 3
NVIDIA
NVIDIA Cosmos 3 is an open frontier foundation model for physical AI, built on a Mixture-of-Transformers (MoT) architecture. It unifies text, image, video, ambient sound, and action generation within a single omnimodal system designed for robotics, autonomous vehicles, and physical environment simulation. Available in Nano (16B) and Super (64B) variants.