Ask anything about a picture with Qwen3-VL for free, without an account: Alibaba's open vision-language model looks at the image you upload and answers in plain language. Describe a photo, read the text in a screenshot or a receipt, explain a chart, identify an object, or simply chat without an image. The demo below runs the 4-billion-parameter instruct version.
Qwen3-VL is the vision-language family of the Qwen3 models released by Alibaba in 2025 under an open licence. It reads images and videos as well as text, with strong results on document reading, charts and multilingual OCR.
A hosted demo of a 4B instruct variant with vision, tuned by a community contributor for fewer refusals. Site rules still apply: adult content and violence are not allowed anywhere on the site.
The image and the question are sent to a shared hosted demo to produce the answer and are not stored by this site. Do not upload personal or confidential documents. This demo is hosted by a third party (Hugging Face Space) and embedded here; it can be slow, queued or temporarily unavailable when the host is busy. Nothing to install, no account needed.
Read the guide: How to write a good prompt