Sai Madhusudan Gunda

Research Engineer, BharatGen  ·  M.S. by Research, Computer Science, IIIT Hyderabad

My research lies at the intersection of computer vision and natural language processing, with a focus on document intelligence: teaching models to read, ground, and reason over complex documents, from modern PDFs to centuries-old palm-leaf manuscripts. I am broadly interested in how vision-language models can move beyond pattern matching toward grounded reasoning, where every prediction traces back to specific visual evidence. This thread runs through M3Grounder (CVPR 2026), DoCoG (ECCV 2026 Oral), and DocLayout-VL (ECCV 2026).

In parallel, I work on OCR for historical Indic manuscripts, where script diversity, physical degradation, and data scarcity collide. UniLipi (ICDAR 2026) unifies OCR across 13 Indic scripts, and CURIO (WACV 2026) tackles curvature and warping in degraded manuscripts. I am currently developing Patram-7B-Instruct at BharatGen, India's National AI Mission, a multilingual vision-language model trained across 256 H100 GPUs.

More broadly, I want to keep asking how models can see, understand, and reason about complex documents in a way that stays faithful to the evidence on the page, especially for the languages, scripts, and archives that today's foundation models still leave behind.

During my M.S., I was advised by Prof. Ravi Kiran Sarvadevabhatla at CVIT.

News

Publications

DoCoG

DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA Oral

European Conference on Computer Vision (ECCV), 2026

Grounded document question answering with step-wise reasoning over textual and graphical elements.

DocLayout
VL

DocLayout-VL: Hierarchical, Open-set, and Promptable Document Layout Segmentation

European Conference on Computer Vision (ECCV), 2026

A promptable foundational document layout model achieving SOTA results across document layout benchmarks.

Indic
OCR

Multi-Script OCR System for Handwritten Indic Manuscripts Demo

17th IAPR International Workshop on Document Analysis Systems (DAS), Vienna, 2026

A live demonstration of a multi-script OCR system tailored for handwritten Indic manuscripts across multiple scripts.

Uni
Lipi

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

International Conference on Document Analysis and Recognition (ICDAR), 2026

A unified OCR model across 13 Indic scripts with script-aware synthetic data, achieving 6.9% CER and strong cross-script generalization.

M3
Grounder

M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

A mask-based grounding framework that localizes multi-span evidence at phrase, line, and block levels for document question answering.

CURIO

CURIO: Curvature-Aligned and Efficient OCR for Low-Resource Historical Manuscripts

Winter Conference on Applications of Computer Vision (WACV), 2026

A curvature-aware OCR model that handles warped and degraded historical Indic manuscripts, improving recognition accuracy for low-resource scripts.

Experience

May 2026 – Present

Research Engineer

BharatGen · National AI Mission, IIIT Hyderabad

  • Converted to full-time after leading the initial development of Patram-7B-Instruct as an intern.
  • Continuing to advance multilingual vision-language modeling for document understanding across Indian languages, and scaling training across large H100 clusters.
Jan 2025 – Apr 2026

Generative AI Research Intern

BharatGen · National AI Mission, IIIT Hyderabad

  • Developed Patram-7B-Instruct, a multilingual vision-language model for document understanding, OCR, layout parsing, QA, classification, and summarization across Indian languages.
  • Built scalable data pipelines for synthetic QA generation, key-value extraction, and summarization for instruction tuning and supervised fine-tuning.
  • Implemented efficient multi-node distributed training using DeepSpeed ZeRO-3 on 256 H100 GPUs.
May 2023 – May 2026

Research Student, Computer Vision

Center for Visual Information Technology (CVIT), IIIT Hyderabad

Advisor: Prof. Ravi Kiran Sarvadevabhatla

  • Built ML systems for grounded document QA, historical manuscript OCR, and multimodal document understanding, with a focus on scalable training, evaluation, and inference.
  • Published research papers at top computer-vision venues (CVPR, WACV, ICDAR, ECCV), spanning document grounding, OCR, multilingual tasks, and document layout segmentation.
Jan 2024 – Apr 2024

Undergraduate Researcher, NLP

Language Technologies Research Centre (LTRC), IIIT Hyderabad

Advisor: Prof. Radhika Mamidi

  • Developed classification models for Indian languages, achieving 92% accuracy on multilingual datasets by fine-tuning.
  • Conducted experiments on cross-lingual transfer and low-resource language classification for Indic NLP.

Education

2025 – 2026

International Institute of Information Technology, Hyderabad

M.S. by Research in Computer Science and Engineering · CGPA 9.0/10

2021 – 2025

International Institute of Information Technology, Hyderabad

B.Tech in Computer Science and Engineering · CGPA 8.4/10

Awards & Honors