MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Specifications
- Provider
- Xiaomi
- Type
- Open-source / open-weight
- Modality
- text+image+audio+video->text
- Context window
- 1050K
- Released
- April 22, 2026
Capabilities
Input: textInput: audioInput: imageInput: videoOutput: text
Strengths
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
More from Xiaomi
