Back to Models
qwen logo
qwen/qwen2-5-vl-32b-instruct
Not Available

Qwen2.5 VL 32B Instruct

Qwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities. It excels at visual analysis tasks, including object recognition, textual interpretation within images, and precise event localization in extended videos. Qwen2.5-VL-32B demonstrates state-of-the-art performance across multimodal benchmarks such as MMMU, MathVista, and VideoMME, while maintaining strong reasoning and clarity in text-based tasks like MMLU, mathematical problem-solving, and code generation.

3/24/2025
128,000 tokens

Specifications

Modalities

Input
text
image
Output
text

Supported Parameters

frequency_penalty
max_tokens
min_p
presence_penalty
repetition_penalty
response_format
seed
stop
temperature
top_k
top_p

Frequently asked questions

Is Qwen2.5 VL 32B Instruct available on FastMetal?
Not at the moment. Qwen3.8 Max (0902), from the same lab, is available on the FastMetal API today.
What is the context window of Qwen2.5 VL 32B Instruct?
128,000 tokens, shared between the prompt and the response.