Does your model understand India?

IndiaSocialBench measures how language models respond to emotionally difficult Indian conversations. It evaluates indirect refusals, family negotiations, honor and shame, grief etiquette, and money between friends in English, Hinglish, and Hindi. Every score links to the transcript and judge explanation behind it.

31
models
18
pilot scenarios
50
items per model
3
language modes
8
dimensions
$30
API budget
Toggle the language to see how the ranking changes.
ModelOverallGap en→hiRefusalsIndirectHierarchyFamilyHonorMixingRitualsMoneySupport
1Claude Fable 5
Anthropic · reasoning: capped low
9.47−0.040%
2Claude Opus 5
Anthropic · reasoning: capped low
9.42−0.310%
3GLM-5.3-Flash (Ox Alpha)
Z.ai · reasoning: capped low
9.36−0.250%
4Kimi K3
Moonshot · reasoning: capped low
9.34−0.020%
5GPT-5.6 Sol
OpenAI · reasoning: capped low
8.78+0.740%
6Qwen3.7 Max
Alibaba · reasoning: capped low
8.74+0.240%
7Claude Sonnet 5
Anthropic · reasoning: capped low
8.64+0.110%
8Grok 4.5
xAI · reasoning: capped low
8.35+0.470%
9MiniMax M342/50 judged
MiniMax · reasoning: capped low
8.27−0.585%
10Gemini 3.6 Flash
Google · reasoning: capped low
8.13−0.200%
11Grok 4.6
xAI · reasoning: capped low
8.08−0.010%
12Gemini 3.5 Flash
Google · reasoning: capped low
7.97−0.510%
13MiMo V2.5 Pro
Xiaomi · reasoning: capped low
7.84−2.270%
14GPT-5.6 Luna
OpenAI · reasoning: capped low
7.78+0.270%
15Gemini 3.1 Flash Lite48/50 judged
Google · reasoning: capped low
7.61−0.110%
16Inkling
Thinking Machines · reasoning: capped low
7.56+0.280%
17Nex N2 Pro46/50 judged
Nex AGI · reasoning: capped low
7.47+0.120%
18GLM-5.2
Zhipu · reasoning: capped low
7.31−1.334%
19DeepSeek V4 Pro
DeepSeek · reasoning: capped low
7.27−3.230%
20DeepSeek V4 Flash
DeepSeek · no hidden reasoning
7.25−1.274%
21Hy346/50 judged
Tencent · reasoning: capped low
7.21+0.950%
22Qwen3.6 Flash
Alibaba · reasoning: capped low
7.07−0.540%
23Sarvam 105B (high reasoning)
Sarvam AI · reasoning: high
7.04−0.490%
24Mistral Large 2512
Mistral · no hidden reasoning
6.80−0.570%
25Claude Haiku 4.5
Anthropic · reasoning: capped low
6.72−1.370%
26Grok 4.3
xAI · reasoning: capped low
6.50−0.210%
27Sarvam 30B42/50 judged
Sarvam AI · reasoning: capped low
5.89+2.410%
28Sarvam 30B (high reasoning)
Sarvam AI · reasoning: high
5.86−1.490%
29Sarvam 105B49/50 judged
Sarvam AI · reasoning: capped low
5.54−0.330%
30GPT-5 Mini
OpenAI · no hidden reasoning
5.48+0.490%
31Llama 4 Maverick
Meta · no hidden reasoning
3.23−0.470%

Scores range from 0 to 10. The whisker on the Overall bar shows the bootstrap 95% CI. “Gap en→hi” is how much the model loses when the same conversations arrive in Hindi (−) or gains (+). Refusals are excluded from scores and reported separately. Click any row for per-dimension detail and full transcripts.
Reasoning policy: to keep runs comparable and affordable, reasoning-capable models are run with thinking effort capped at “low” (marked per model above); rows labeled “high reasoning” are explicit higher-effort variants, shown separately so no model gets a hidden advantage.

What the numbers say

The pilot shows a language gap

21 of 31 models score lower when the identical situations arrive in Hindi instead of English. The average drop is 0.31 points. Same problems, same rubric, same judges; only the language changed.

Culture is harder than empathy

The field averages 8.67 on the culture-neutral support control, but 7.31 across the seven cultural dimensions and is weakest on indirect speech & face-saving (6.11). Models know how to feel; they don't yet know how India works.

Refusals on family topics

MiniMax M3 refused ordinary family-life scenarios at 4% or more. Over-refusal is reported as its own column and is never hidden in the average.

Excluded from this board: GLM-4.7 (only 19/50 items completed, insufficient coverage (provider errors)).

Dataset vd7bb3900ca057116 · generated 2026-08-26 · 18 base scenarios · single uniform blinded judge (see methodology & limits)