Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Specifications
- Provider
- Qwen (Alibaba)
- Type
- Open-source / open-weight
- Modality
- text+image->text
- Context window
- 262,144
- Knowledge cutoff
- 2025-03-31
- Released
- October 6, 2025
Capabilities
Input: textInput: imageOutput: text
Strengths
frequency_penaltylogit_biaslogprobsmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
