Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Specifications
- Provider
- Qwen (Alibaba)
- Type
- Open-source / open-weight
- Modality
- text+image->text
- Context window
- 262,144
- Released
- October 14, 2025
Capabilities
Input: imageInput: textOutput: text
Strengths
frequency_penaltylogit_biaslogprobsmax_tokenspresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
