DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA Oral
European Conference on Computer Vision (ECCV), 2026
Grounded document question answering with step-wise reasoning over textual and graphical elements.
Research Engineer, BharatGen · M.S. by Research, Computer Science, IIIT Hyderabad
My research lies at the intersection of computer vision and natural language processing, with a focus on document intelligence: teaching models to read, ground, and reason over complex documents, from modern PDFs to centuries-old palm-leaf manuscripts. I am broadly interested in how vision-language models can move beyond pattern matching toward grounded reasoning, where every prediction traces back to specific visual evidence. This thread runs through M3Grounder (CVPR 2026), DoCoG (ECCV 2026 Oral), and DocLayout-VL (ECCV 2026).
In parallel, I work on OCR for historical Indic manuscripts, where script diversity, physical degradation, and data scarcity collide. UniLipi (ICDAR 2026) unifies OCR across 13 Indic scripts, and CURIO (WACV 2026) tackles curvature and warping in degraded manuscripts. I am currently developing Patram-7B-Instruct at BharatGen, India's National AI Mission, a multilingual vision-language model trained across 256 H100 GPUs.
More broadly, I want to keep asking how models can see, understand, and reason about complex documents in a way that stays faithful to the evidence on the page, especially for the languages, scripts, and archives that today's foundation models still leave behind.
During my M.S., I was advised by Prof. Ravi Kiran Sarvadevabhatla at CVIT.
Email / CV / GitHub / LinkedIn / Google Scholar
European Conference on Computer Vision (ECCV), 2026
Grounded document question answering with step-wise reasoning over textual and graphical elements.
European Conference on Computer Vision (ECCV), 2026
A promptable foundational document layout model achieving SOTA results across document layout benchmarks.
17th IAPR International Workshop on Document Analysis Systems (DAS), Vienna, 2026
A live demonstration of a multi-script OCR system tailored for handwritten Indic manuscripts across multiple scripts.
International Conference on Document Analysis and Recognition (ICDAR), 2026
A unified OCR model across 13 Indic scripts with script-aware synthetic data, achieving 6.9% CER and strong cross-script generalization.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
A mask-based grounding framework that localizes multi-span evidence at phrase, line, and block levels for document question answering.
Winter Conference on Applications of Computer Vision (WACV), 2026
A curvature-aware OCR model that handles warped and degraded historical Indic manuscripts, improving recognition accuracy for low-resource scripts.
BharatGen · National AI Mission, IIIT Hyderabad
BharatGen · National AI Mission, IIIT Hyderabad
Center for Visual Information Technology (CVIT), IIIT Hyderabad
Advisor: Prof. Ravi Kiran Sarvadevabhatla
Language Technologies Research Centre (LTRC), IIIT Hyderabad
Advisor: Prof. Radhika Mamidi
M.S. by Research in Computer Science and Engineering · CGPA 9.0/10
B.Tech in Computer Science and Engineering · CGPA 8.4/10