
Princeton Journal of Interdisciplinary Research, Volume 1, Issue 3
— Bridging Horizons (March 2026) - ISSN 3069-8200
Evaluating Sentiment Analysis Models on Telugu-English Code-Mixed and Monolingual Texts
Author: Pranav Nimmagadda
Affiliation: Carnegie Mellon University
Abstract: The increasing prevalence of code-mixing in non-English languages presents a significant challenge for traditional sentiment analysis tools, which are designed to handle monolingual text. The complexities of grammar, syntax, and cultural context introduce challenges that are not sufficiently addressed by existing English-centric models. In this paper, the performance of a trained sentiment analysis model was evaluated on code-mixed Telugu-English text (CMTET) and monolingual English datasets. The results show that the trained model has a better weighted precision than the untrained model, while English benefited slightly from the pretrained model’s bias. The findings highlight the limitations of code mixing training for generalizing to monolingual datasets and the need for improvements in text data preprocessing. This analysis could help refine sentiment classification models by incorporating linguistic features linked with specific parts of speech.
Keywords: