DUBLIN: Visual Document Understanding By Language-Image Network

Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Khan Mohammed, Qiang Liu, Monojit Choudhury, Hardik Chauhan, Subhojit Som, Vishrav Chaudhary, Saurabh Tiwary


Abstract
In this paper, we present DUBLIN, a pixel-based model for visual document understanding that does not rely on OCR. DUBLIN can process both images and texts in documents just by the pixels and handle diverse document types and tasks. DUBLIN is pretrained on a large corpus of document images with novel tasks that enhance its visual and linguistic abilities. We evaluate DUBLIN on various benchmarks and show that it achieves state-of-the-art performance on extractive tasks such as DocVQA, InfoVQA, AI2D, OCR-VQA, RefExp, and CORD, as well as strong performance on abstraction datasets such as VisualMRC and text captioning. Our model demonstrates the potential of OCR-free document processing and opens new avenues for applications and research.
Anthology ID:
2023.emnlp-industry.65
Volume:
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track
Month:
December
Year:
2023
Address:
Singapore
Editors:
Mingxuan Wang, Imed Zitouni
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
693–706
Language:
URL:
https://aclanthology.org/2023.emnlp-industry.65
DOI:
10.18653/v1/2023.emnlp-industry.65
Bibkey:
Cite (ACL):
Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Khan Mohammed, Qiang Liu, Monojit Choudhury, Hardik Chauhan, Subhojit Som, Vishrav Chaudhary, and Saurabh Tiwary. 2023. DUBLIN: Visual Document Understanding By Language-Image Network. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 693–706, Singapore. Association for Computational Linguistics.
Cite (Informal):
DUBLIN: Visual Document Understanding By Language-Image Network (Aggarwal et al., EMNLP 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.emnlp-industry.65.pdf
Video:
 https://aclanthology.org/2023.emnlp-industry.65.mp4