AI Tools Directory
HomeAI ModelsAI ToolsCompareLandscapeCategoriesAboutHealth

Multimodal Models

Models that work with multiple data types (text, image, audio, video)

2 models in this category

MiniMax M3

MiniMax AI

🟢
Multimodal Models

MiniMax M3 is a frontier-tier open-weight multimodal agent model with 428B total parameters (23B activated) and a 1M token context window. Released May 31, 2026, it features MiniMax Sparse Attention (MSA) for 9x prefill and 15x decode speedup over its predecessor, and is specialized for coding, agentic tasks, and complex reasoning.

NVIDIA Cosmos 3

NVIDIA

🟢
Multimodal Models

NVIDIA Cosmos 3 is an open frontier foundation model for physical AI, built on a Mixture-of-Transformers (MoT) architecture. It unifies text, image, video, ambient sound, and action generation within a single omnimodal system designed for robotics, autonomous vehicles, and physical environment simulation. Available in Nano (16B) and Super (64B) variants.

AI Tools Directory
Browse ModelsCompareLandscapeAbout

© 2026 AI Tools Directory. All rights reserved.