I Surveyed 123 People in India to Benchmark Frontier AI A developer built a benchmark using field data from 123 university students in Goa, India, covering 24 real institutional disputes across education, healthcare, justice and finance to test how frontier and open-weight language models handle ambiguous legal scenarios under opposing user personas. Evaluating eight models including Gemini, Claude, GPT, Qwen, GLM and DeepSeek via the Kaggle Benchmarks SDK, the study found demographic parity remains unsolved, with pairwise violation rates ranging from 16.67% to 41.67% and DeepSeek-R1 showing the lowest sycophancy drift at 0.042. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Current evaluations score language models using Western multiple-choice questions. Models easily pass these tests by memorizing standard templates. When given incomplete information, they pick neutral options to appear fair. To test deeper behavior, we gathered field data from 123 university students in Goa, India. We chose 24 real institutional disputes where human consensus broke down. These disputes span four vital sectors: education, healthcare, justice, and finance. Our benchmark investigates a distinct behavioral failure. A model commits to an abstract moral rule. Next, we present a concrete legal scenario. Finally, two opposing user personas pressure the model. We tracked four metrics across 24 complex disputes: The test covers 6 dilemmas in each sector: We evaluated frontier and open-weight models using the Kaggle Benchmarks SDK: | Model Tier | Model Identifier | Provider | Evaluation Rationale | |---|---|---|---| | High-Speed Frontier Reasoning | google/gemini-3-flash-preview | | Rapid reasoning benchmark; tested on legal texts. | | 80B Direct Instruction | qwen/qwen3-next-80b-a3b-instruct | Alibaba Qwen | High-parameter instruction baseline; tested without thinking traces. | | Frontier Safety Alignment | anthropic/claude-sonnet-5@default | Anthropic | Commercial safety alignment; evaluated on demographic invariance. | | Frontier Reasoning Hybrid | google/gemini-3.7-flash | | Hybrid thinking model; tested on statutory interpretation. | | Chinese Frontier Foundation | zhipu/glm-5 | Zhipu AI | Leading non-Western model; evaluated on persona resilience. | | Commercial Mid-Tier | openai/gpt-5.4-mini-2026-03-17 | OpenAI | Resource-balanced model; tested on welfare allocation. | | Open Reinforcement Learning | deepseek-ai/deepseek-r1-0528 | DeepSeek | Pure reasoning model; tested for internal bias. | | 80B Extended Thinking | qwen/qwen3-next-80b-a3b-thinking | Alibaba Qwen | High-parameter reasoning model; tested on court bail. | Our evaluation demonstrated that demographic parity remains an unsolved challenge across all evaluated models. | Evaluated Model | Ambiguity Acc | Pairwise Violation Rate PVR | Mean Sycophancy Drift | Human Divergence JSD | Run Cost | |---|---|---|---|---|---| | gemini-3-flash-preview | 100.0% 24/24 | 16.67% 4/24 flips | 0.333 | 0.1691 | $0.4516 | | qwen3-next-80b-a3b-instruct | 50.0% 12/24 | 16.67% 4/24 flips | 0.250 | 0.1764 | — | | claude-sonnet-5-default | 100.0% 24/24 | 20.83% 5/24 flips | 0.250 | 0.1644 | $0.3758 | | gemini-3.7-flash | 100.0% 24/24 | 20.83% 5/24 flips | 0.250 | 0.1721 | $0.3244 | | glm-5 | 100.0% 24/24 | 20.83% 5/24 flips | 0.375 | 0.1581 | $0.6860 | | gpt-5.4-mini-2026-03-17 | 100.0% 24/24 | 25.00% 6/24 flips | 0.333 | 0.1370 | $0.0591 | | deepseek-r1-0528 | 95.8% 23/24 | 37.50% 9/24 flips | 0.042 | 0.1967 | $0.4244 | | qwen3-next-80b-a3b-thinking | 100.0% 24/24 | 41.67% 10/24 flips | 0.208 | 0.1816 | $0.1753 | Disaggregating results across sectors highlights sharp contrasts between model families: | Model Identifier | Education PVR | Healthcare PVR | Justice PVR | Finance PVR | Healthcare Sycophancy | Education Sycophancy | |---|---|---|---|---|---|---| | gemini-3-flash-preview | 0.00% | 16.67% | 16.67% | 33.33% | 0.167 | 0.333 | | qwen3-next-80b-a3b-instruct | 0.00% | 0.00% | 16.67% | 50.00% | 0.333 | 0.167 | | claude-sonnet-5-default | 33.33% | 16.67% | 33.33% | 0.00% | 0.167 | 0.333 | | gemini-3.7-flash | 33.33% | 16.67% | 33.33% | 0.00% | 0.167 | 0.333 | | glm-5 | 0.00% | 16.67% | 33.33% | 33.33% | 0.167 | 0.667 | | gpt-5.4-mini-2026-03-17 | 16.67% | 16.67% | 33.33% | 33.33% | 0.500 | 0.333 | | deepseek-r1-0528 | 16.67% | 66.67% | 33.33% | 33.33% | 0.000 | 0.000 | | qwen3-next-80b-a3b-thinking | 33.33% | 50.00% | 66.67% | 16.67% | 0.333 | 0.333 | Reasoning models produced the highest decision flip rates on the leaderboard. Qwen Thinking reached 41.67% total flips. DeepSeek-R1 reached 37.50% total flips. Extended internal thinking tokens acted as rationalization engines. When demographic tokens changed, internal reasoning chains generated extra assumptions regarding flight risk or family poverty. In criminal justice, Qwen Thinking flipped 4 out of 6 rulings. It granted bail to dominant-caste applicants while requiring surety bonds from marginalized applicants. Comparing Qwen Thinking against Qwen Instruct exposed a fundamental trade-off between deliberation and certainty. Qwen Instruct produced only 4 demographic flips across 24 disputes, achieving 16.67% parity violations. However, it failed ambiguity detection in 12 out of 24 cases. In financial and judicial cases, it rushed to decide outcomes despite missing documentation. Qwen Thinking achieved a perfect score on ambiguity detection. It correctly refused to decide when records lacked critical context. Yet its internal reasoning tokens later rationalized opposing verdicts when names changed. Thinking tokens prevented premature conclusions while enabling demographic rationalization. DeepSeek-R1 scored near-zero sycophancy across all four sectors. It maintained its stance regardless of whether hospital directors or activists asked the questions. Conversely, GLM-5 exhibited strong sycophancy in education. It achieved zero demographic flips under direct testing. Yet when conversational personas engaged the model, it shifted positions entirely. It agreed with reservation critics, then agreed with quota advocates on identical legal facts. Western frontier models showed strong consistency in financial tasks. Claude Sonnet 5 and Gemini 3.7 Flash achieved zero decision flips across all financial cases. However, justice evaluations triggered high violation rates across five models. Statutory protections in criminal law and affirmative action disputes remain difficult for current frontier models. You can inspect our live benchmark leaderboard, evaluate model runs, and explore the underlying code on Kaggle: