Autonomous Qwen3-VL training-code research on the official DocVQA benchmark. main: NVIDIA multi-GPU, mlx: Apple Silicon/MPS.
-
Updated
Jun 14, 2026 - Python
Autonomous Qwen3-VL training-code research on the official DocVQA benchmark. main: NVIDIA multi-GPU, mlx: Apple Silicon/MPS.
99.156% Accuracy from Agentic Document Extraction DPT-2 model on DocVQA val split
Reproducible framework for document-centric VLM robustness and claim-faithfulness evaluation.
Production-grade Multimodal Document Intelligence & RAG Platform featuring layout-aware OCR (Docling/RapidOCR), hybrid retrieval (BM25 + pgvector + Reranking), grounded QA with provenance citations, and a real-time React dashboard.
Quasar PoC, Multitenant PoC.
Open multimodal evaluation harness for Xiaomi MiMo
Config-driven long-context benchmark toolkit for vision-language models
Document XAI Model for DocVQA. Official implementation of "Towards Self-Explainable DocVQA with Chain-of-Explanation Predictions". Submitted to NeurIPS 2026
The repository host codes, link to datasets and models for our research paper. In this paper we have developed a novel approach that can perform DocVQA, RCVQA and MathVQA tasks.
To associate your repository with the docvqa topic, visit your repo's landing page and select "manage topics."