-->

SHAGUN DWIVEDI

Shagun Dwivedi S. DWIVEDI
↳ /ʃəˈɡʊn diˈveːdi/

I am an AI researcher working on NLP and language modelling for morphologically complex and low-resource languages using computational linguistics. I aim to build systems that work equitably across languages.

At the Centre for Interdisciplinary AI, where I am advised by Dr. Kaushik Gopalan, my research focuses on how script-specific and morphological properties of tokenizers can affect model design and performance for Indic languages. Parallel to this, I have worked on linguistically informed, data-efficient methods for handwritten text recognition for historical (manu)scripts.

In my free time, I enjoy reading, painting, and listening to podcasts. I have also been getting into puzzles and clay art lately.

Low-resource Languages Tokenization Multilingual NLP Handwritten text recognition Computational Linguistics
Dispatches

Updates

AUG2026
paper"Performance of Grapheme-Based Tokenizers on Word-level and Natural Language Understanding tasks for Indic Languages" was accepted to EMNLP 2026 (Findings).
AUG2026
paper"Synthetic Data Generation using Grapheme-level Visual Units for Handwritten Text Recognition of Grantha Manuscripts" was accepted to Second Workshop on Curated Data for Efficient Learning (CDEL), ECCV 2026.
JUL2026
paper"Comparative Analysis of the Intrinsic Metrics for Tokenizers and their effect on Downstream Tasks for Hindi and Marathi" has been published in ACL Main 2026.
JUN2026
PAPEROur paper "Synthetic Line Image Generation from Cropped Grapheme Images for Handwritten Marathi Text Recognition" has been accepted to ICDAR 2026 Workshop on Document Analysis of Low-resource Languages.
The archive

Publications

ACL 2026 — Main Conference

Comparative Analysis of the Intrinsic Metrics for Tokenizers and their effect on Downstream Tasks for Hindi and Marathi

Dwivedi, S. & Gopalan, K.

Various studies have pointed out that the performance of language models is poor in non-English or non-European languages. One of the factors affecting this performance is the effectiveness and suitability of the tokenization scheme used in the model. Indic scripts require multiple Unicode codepoints to represent a single visual unit to be encoded in the standard UTF-8 scheme. This paper investigates the effect of multiple tokenizers that use UTF-8 text input on the downstream performance of pretrained language models for Hindi and Marathi, languages written in Devanāgari script. We present the intrinsic performance of the tokenizers using Fertility, Rényi Efficiency and Percentile Frequency, and report the extrinsic performance of monolingual and multilingual models on question-answering tasks, using an automated parts-of-speech and sentence similarity based evaluation framework, and on word-level tasks such as grapheme-to-phoneme conversion and transliteration. We propose a grapheme cluster tokenizer for the script which shows performance better than or competitive with other popular tokenizers. We also find that the Rényi Efficiency metric is highly correlated to downstream performance on question answering.

EMNLP 2026 - Findings

Performance of Grapheme-Based Tokenizers on Word-level and Natural Language Understanding tasks for Indic Languages

Dwivedi, S., Saquib, M., Gopalan, K.

Tokenization is one of the factors which influences the multilingual downstream performance of a language model. This paper demonstrates that vocabulary constrained grapheme cluster based tokenization leads to improved performance on Natural Language Understanding and Word-level tasks. We assess three different grapheme cluster based tokenizers, which are the Grapheme Cluster tokenizer, the Grapheme Pair Encoding, and a vocabulary constrained version of Grapheme Pair Encoding, across eleven Indic languages, namely Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We present the intrinsic performance of three grapheme cluster based tokenizers using Fertility, Rényi Efficiency, and Percentile Frequency. We also compare the performance of these grapheme based tokenizers with seven popular multilingual and Indic-focused tokenizers on downstream tasks. We find that the vocabulary constrained Grapheme Pair Encoding tokenizer exhibits competitive and sometimes superior downstream performance in comparison to other grapheme cluster based tokenizers.

CDEL — ECCV Workshop 2026

Synthetic Data Generation using Grapheme-level Visual Units for Handwritten Text Recognition of Grantha Manuscripts

Dwivedi, S. & Gopalan, K.

Handwritten Text Recognition of the Grantha script can help digitise several thousands of manuscripts and aid scholars in their research. However, deep learning algorithms require vast amounts of data to train recognition models that are robust to real world handwritten data. We propose a manuscript-native synthetic data generation pipeline, which extracts grapheme-level visual units from the manuscript leaves to create a glyph bank and recombines the glyphs to form line images using words from a Sanskrit or Tamil lexicon in the Grantha script. We assess the effectiveness of the dataset for HTR using a CNN-BiLSTM-CTC model for six manuscripts consisting of both Tamil and Sanskrit texts. The model trained with manuscript-based synthetic dataset consistently outperforms the model trained solely on font-rendered synthetic data. We also find that further improvements can be obtained by combining a smaller fraction of the synthetic data with a small amount of augmented real manuscript images, providing an effective and data-efficient approach for developing HTR systems for low-resource Grantha manuscripts.

DALL — ICDAR 2026 Workshop

Synthetic Line Image Generation from Cropped Grapheme Images for Handwritten Marathi Text Recognition

Vaishampayan, J.; Gopalan, K.; Dwivedi, S.

We study line-level handwritten text recognition for historical Marathi manuscripts written in the Devanagari script. Existing synthetic data pipelines for Indic scripts render characters from printed fonts, producing ink distributions that diverge from handwritten strokes. Our main contribution is a manuscript-native synthetic line-image generation pipeline for this low-resource setting. Synthetic lines are assembled from a collection of real grapheme-cluster (akshara) crops extracted from manuscript pages in collection EAP248 of the Endangered Archives Programme at the British Library. Aksharas are combined into words drawn from a Marathi dictionary, and the images containing the aksharas are concatenated into line images by aligning the shirorekhas horizontally. The pipeline yields a corpus of roughly one million synthetic lines. The evaluation data is drawn from eight EAP248 manuscripts, where four contribute training lines and four are held out entirely. This results in a 525-line test set with an explicit seen/unseen handwriting split. CRNN and ViT+CTC recognisers trained on the synthetic corpus together with 657 real lines attain a character error rate of 24.0% (ViT+CTC) and 25.1% (CRNN) under a space-agnostic metric, against 45.4–86.4% for off-the-shelf OCR engines, fine-tuned baselines, and off-the-shelf vision–language models.

Computational Sanskrit and Digital Humanities - World Sanskrit Conference 2025

A Case Study of Handwritten Text Recognition from Pre-Colonial era Sanskrit Manuscripts

Chincholikar, K.*, Dwivedi, S.*, Gopalan, K., Awasthi, T.

* Equal contribution

Extracting text from handwritten documents and converting it to a digital form, commonly called handwritten Text Recognition (HTR), opens up new ways to access and study scanned texts. In this case study, we perform HTR on Sanskrit manuscripts from the Early Modern period, namely Vādakautūhala of Svāmiśāstrin and Bhāskararāya (early eighteenth century), and Mahāvākyārtha and Dvādaśamahāvākyārthavicāra of unknown authorship. The first is a Pūrva Mīmāṃsā text from the 18th century, while the other two are texts on the tradition of Vedānta. The proposed HTR method consists of three steps: line segmentation, line recognition, and post-correction. For segmenting line images from page images, we propose an approach that uses the scene text detection method Character Region Awareness for Text Detection (CRAFT) to detect individual characters and applies customised logic to segment them into discrete lines of text. To recognise text content from the segmented line images, we fine-tune a pre-trained recognition model for the Devanāgarī script provided by the EasyOCR library. The post-correction model, which makes language-aware corrections to the recognition model outputs, is a fine-tuned version of ByT5-Sanskrit, a pre-trained language model for Sanskrit. We observe that the proposed method performs better than out-of-the-box HTR solutions by addressing problems that are unique to older Sanskrit manuscripts, such as cramped and irregular line spacing, changes in the script over centuries, and stylistic differences in writing due to location, period, and scribe.

Intl. Conf. on Human-Computer Interaction, 2025

A Semi-Automatic Text Recognition Tool for Pre-Colonial Handwritten Manuscripts in Devanāgari Script

Valaboju, B., Dwivedi, S., Chincholikar, K., Gopalan, K., Vidwans, V.

Manual annotation of undigitized manuscripts is a resource and labor intensive endeavour. On the other hand, using out-of-the-box Optical Character Recognition (OCR) tools to digitise text from handwritten manuscripts is challenging, particularly in languages like Sanskrit, where the scripts have evolved significantly over time. This poster presents an annotation tool which allows the user to extract text from undigitized manuscripts using OCR, following which users can make corrections to the OCR-detected text. Users can then request fine tuning on a few pages corrected by them, making the annotation process easier and more efficient for the subsequent pages by improving OCR performance. Additionally, an issue faced in the annotation of manuscripts is that the Devanāgari script is sometimes laborious to edit due to the behaviour of conjunct clusters and the halant, necessitating many keystrokes for simple edits. To mitigate this issue, the application uses an additional text box that represents the text in the Harvard-Kyoto transliteration scheme, which is easier to edit given that most keyboards use the Roman alphabet. The tool also supports rare Sanskrit characters, which are not usually supported by standard typing tools. The frontend, which is designed to reduce effort on the annotator’s part by placing the recognised text in editable text boxes right below the line images, is implemented in Vue.js. The backend is a Flask server employing PyTorch for inference and fine-tuning. The application aims to facilitate efficient annotation by having users upload manuscript images, which are then processed by the backend, which segments text lines from the leaves and recognises the text in the line images. Initial evaluations from beta testers suggest that the application significantly enhances the efficiency of the annotation process, making it more accessible and reducing the time required for its completion.

Intl. Conf. on Computer Vision & Image Processing, 2024

Converting Gujarati Text in Custom-Embedded Subsetted Non-Unicode Fonts to Searchable Formats: A Case Study Using Jain Religious Texts

Jain, R., Dwivedi, S., Gopalan, K.

Unicode font is a globally recognized coding system that assigns a unique number to every character, regardless of the platform, program, or language. However, a significant fraction of Gujarati text documents available online follow legacy fonts which are non-Unicode standard. These documents are rendered unsearchable due to the use of custom-embedded non-Unicode font subsets. This study presents a novel approach for extracting Gujarati text from non-Unicode standard PDFs with embedded fonts, addressing the challenges posed by these legacy fonts and providing a pathway to convert and preserve such documents in a searchable, Unicode-compliant format. Using VGG16, features are extracted from images of glyphs taken from non-Unicode font files. Cosine similarity is used to compare these glyph images with a consolidated reference set, assigning text characters to each Unicode character in the non-Unicode font files based on the best matching image in the reference set. The reconstructed text from the proposed method was compared with State-of-the-Art OCR technologies, including Google Cloud Vision OCR and Tesseract OCR. The results demonstrated a substantial improvement, with the proposed method achieving error rates between 0–2% for the majority of pages, compared to 4–6% with Tesseract OCR and 6–13% with Google Vision OCR.