Officical repository for the paper“ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering”(EMNLP'24)
-
Updated
Nov 16, 2024 - Python
Officical repository for the paper“ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering”(EMNLP'24)
Vision-language AI for chart question answering using Qwen3-VL with SFT and GRPO training
VisDoTQA is a chart visual reasoning benchmark and dataset for Position, Length, Pattern, and Extract tasks.
This is the official repository for our paper 📄 “In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding”
Reproducible framework for document-centric VLM robustness and claim-faithfulness evaluation.
Multimodal tool-use RL on ChartQA: Qwen2.5-VL learns visual tool calls (highlight/mask/box) with two-stage rollout and three generations of reward design (rule / shaped / G-RA + Dr.GRPO) on verl/EasyR1
On-policy distillation for a vision-language model (Qwen3-VL 2B student, 8B teacher, ChartQA): OPD matches full-data SFT with 10% of the questions, plus a token-level analysis of where teacher feedback lands.
Factorized candidate sharing for ChartQA and Raven-style visual reasoning.
Open multimodal evaluation harness for Xiaomi MiMo
Directional scale errors in ChartQA and interference-aware LoRA merging with MIRAGE.
Fine-tuning UniChart (vision encoder-decoder) for chart question answering on ChartQA — parameter-efficient training, relaxed-accuracy evaluation, 3x exact-match gain over base
To associate your repository with the chartqa topic, visit your repo's landing page and select "manage topics."