PinpointQA: Spatial Grounding in Indoor Videos
Ask questions about the spatial location of small objects in indoor video scenes. The model identifies objects and describes their position relative to nearby references.
Built on Qwen3-VL-8B-Instruct fine-tuned with PinpointQA LoRA from the PinpointQA benchmark.
Examples