Does your model understand India?

IndiaSocialBench measures how language models respond to emotionally difficult Indian conversations. It evaluates indirect refusals, family negotiations, honor and shame, grief etiquette, and money between friends in English, Hinglish, and Hindi. Every score links to the transcript and judge explanation behind it.

28
models
18
pilot scenarios
50
items per model
3
language modes
8
dimensions
$30
API budget
Toggle the language to see how the ranking changes.
ModelOverallGap en→hiRefusalsIndirectHierarchyFamilyHonorMixingRitualsMoneySupport
1Claude Fable 5
Anthropic · reasoning: capped low
9.47−0.040%
2Claude Opus 5
Anthropic · reasoning: capped low
9.42−0.310%
3Kimi K3
Moonshot · reasoning: capped low
9.34−0.020%
4GPT-5.6 Sol
OpenAI · reasoning: capped low
8.78+0.740%
5Qwen3.7 Max
Alibaba · reasoning: capped low
8.74+0.240%
6Claude Sonnet 5
Anthropic · reasoning: capped low
8.64+0.110%
7Grok 4.5
xAI · reasoning: capped low
8.35+0.470%
8MiniMax M3partial · 42/50
MiniMax · reasoning: capped low
8.27−0.585%
9Gemini 3.5 Flash
Google · reasoning: capped low
7.97−0.510%
10MiMo V2.5 Pro
Xiaomi · reasoning: capped low
7.84−2.270%
11GPT-5.6 Luna
OpenAI · reasoning: capped low
7.78+0.270%
12Gemini 3.1 Flash Lite
Google · reasoning: capped low
7.61−0.110%
13Inkling
Thinking Machines · reasoning: capped low
7.56+0.280%
14Nex N2 Propartial · 46/50
Nex AGI · reasoning: capped low
7.47+0.120%
15GLM-5.2
Zhipu · reasoning: capped low
7.31−1.334%
16DeepSeek V4 Pro
DeepSeek · reasoning: capped low
7.27−3.230%
17DeepSeek V4 Flash
DeepSeek · no hidden reasoning
7.25−1.274%
18Hy3partial · 46/50
Tencent · reasoning: capped low
7.21+0.950%
19Qwen3.6 Flash
Alibaba · reasoning: capped low
7.07−0.540%
20Sarvam 105B (high reasoning)
Sarvam AI · reasoning: high
7.04−0.490%
21Mistral Large 2512
Mistral · no hidden reasoning
6.80−0.570%
22Claude Haiku 4.5
Anthropic · reasoning: capped low
6.72−1.370%
23Grok 4.3
xAI · reasoning: capped low
6.50−0.210%
24Sarvam 30Bpartial · 42/50
Sarvam AI · reasoning: capped low
5.89+2.410%
25Sarvam 30B (high reasoning)
Sarvam AI · reasoning: high
5.86−1.490%
26Sarvam 105B
Sarvam AI · reasoning: capped low
5.54−0.330%
27GPT-5 Mini
OpenAI · no hidden reasoning
5.48+0.490%
28Llama 4 Maverick
Meta · no hidden reasoning
3.23−0.470%

Scores range from 0 to 10. The whisker on the Overall bar shows the bootstrap 95% CI. “Gap en→hi” is how much the model loses when the same conversations arrive in Hindi (−) or gains (+). Refusals are excluded from scores and reported separately. Click any row for per-dimension detail and full transcripts.
Reasoning policy: to keep runs comparable and affordable, reasoning-capable models are run with thinking effort capped at “low” (marked per model above); rows labeled “high reasoning” are explicit higher-effort variants, shown separately so no model gets a hidden advantage.

What the numbers say

The pilot shows a language gap

18 of 28 models score lower when the identical situations arrive in Hindi instead of English. The average drop is 0.32 points. Same problems, same rubric, same judges; only the language changed.

Culture is harder than empathy

The field averages 8.55 on the culture-neutral support control, but 7.20 across the seven cultural dimensions and is weakest on indirect speech & face-saving (5.96). Models know how to feel; they don't yet know how India works.

Refusals on family topics

MiniMax M3 refused ordinary family-life scenarios at 4% or more. Over-refusal is reported as its own column and is never hidden in the average.

Excluded from this board: GLM-4.7 (only 19/50 items completed, insufficient coverage (provider errors)).

Dataset vd7bb3900ca057116 · generated 2026-07-26 · 18 base scenarios · single uniform blinded judge (see methodology & limits)