GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Specifications
- Provider
- Z.AI
- Type
- Open-source / open-weight
- Modality
- text+image+video->text
- Context window
- 131,072
- Released
- December 8, 2025
Capabilities
Input: imageInput: textInput: videoOutput: text
Strengths
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p
