About IndiaSocialBench

IndiaSocialBench measures whether language models understand the emotional and social texture of Indian life. It evaluates conversations, not trivia about India.

Frontier models top every English emotional-intelligence benchmark. Meanwhile the fastest-growing population of new AI users types in Hinglish, and the conversations they bring include a boss who can't be contradicted directly, a friend's loan that can't be refused outright, a rishta the family is pushing, a condolence message that must strike exactly the right register. These are precisely the conversations no benchmark measures.

“Set firm boundaries with your mother-in-law” is a coherent English sentence and a culturally impossible action. A model that gives that advice hasn't failed at empathy; it has failed at India. That difference is measurable, and this project measures it.

Why it matters

Indic model builders train culturally grounded models but have no instrument that proves the cultural advantage. Enterprises deploying conversational AI to hundreds of millions of Indian users select vendors on latency and ASR accuracy because nobody can tell them which model will mishandle a grieving customer. A benchmark is the smallest product that moves both: it converts “our model understands Indian users” from a marketing claim into a number, a per-dimension diagnostic, and a set of receipts.

The author

Built by Naresh Silla as a product and research portfolio project: an exercise in finding the eval a market actually needs, then building it with the rigor the claim requires. The dataset, harness, scoring code, and site are open at github.com/sillanaresh/IndiaSocialBench.

The current technical report covers the method and first results. It remains provisional until human agreement is measured.