Qwen3 VL 235B A22B Thinking
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math. The series emphasizes robust perception (recognition of diverse real-world and synthetic categories), spatial understanding (2D/3D grounding), and long-form visual comprehension, with competitive results on public multimodal benchmarks for both perception and reasoning. Beyond analysis, Qwen3-VL supports agentic interaction and tool use: it can follow complex instructions over multi-image, multi-turn dialogues; align text to video timelines for precise temporal queries; and operate GUI elements for automation tasks. The models also enable visual coding workflows, turning sketches or mockups into code and assisting with UI debugging, while maintaining strong text-only performance comparable to the flagship Qwen3 language models. This makes Qwen3-VL suitable for production scenarios spanning document AI, multilingual OCR, software/UI assistance, spatial/embodied tasks, and research on vision-language agents.
Specifications
Modalities
Supported Parameters
Max Output Tokens
32,768Reasoning Configuration
- Default
- Thinking on
- Turning thinking off
- Not possible (always thinks)
Frequently asked questions
- Is Qwen3 VL 235B A22B Thinking available on FastMetal?
- Not at the moment. Qwen3.8 Max (0902), from the same lab, is available on the FastMetal API today.
- What is the context window of Qwen3 VL 235B A22B Thinking?
- 131,072 tokens, shared between the prompt and the response.
- How does Qwen3 VL 235B A22B Thinking rank?
- #154 on the public arena's Overall board (ELO 1,395). Ranks move as the leaderboard is updated.