I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong. A developer built a six-task Kaggle benchmark of 16 cases each to test 11 AI models on everyday Indian business data, including GSTIN check-digit validation, GST tax splits, invoice-number rules, lakh/crore formatting and UPI reconciliation. GPT-5.5 and Gemini 3.7 Flash scored a perfect 1.00 average, while GPT-5.4 nano trailed at 0.34; models that reasoned step by step hit 100% on the GSTIN check character versus 6–12% for fast models, and several models wrongly split a Puducherry intra-state sale into CGST + SGST instead of recognizing the Union Territory with a legislature as a state. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 I build software for Indian small businesses POS, billing, payments , and this month I've been contributing Kestra workflow blueprints for Indian finance: GSTIN validation, UPI end-of-day close, GSTR-2B reconciliation. Every one of them exists because one small mistake costs real money: So I asked: can AI models handle everyday Indian business data? I built a Kaggle benchmark with 6 tasks, 16 cases each: | Task | What the model has to do | |---|---| | india gstin check digit | Validate a GSTIN and compute its Luhn mod 36 check character algorithm given | | india gstin from memory | Same, without being told the algorithm | | india gst tax split | Split GST into CGST + SGST, CGST + UTGST or IGST, with real traps: supplies to an SEZ unit in your own state, and Union Territories with and without a legislature | | india invoice number rules | Apply GST invoice-number rules 16 characters, only letters/digits/ - / / , unique per financial year with look-alike traps: en dash, division slash, trailing space, 17 characters, extra leading zero | | india lakh crore formats | Indian digit grouping 12,34,56,789 , lakh/crore conversions, cheque amounts in words | | india upi reconciliation | Match 18–30 UPI bills against a settlement report and list the bills whose money never arrived, arrived short, or share a UPI reference | How it's graded: python-stdnum 0 disagreements . The other answer keys use the same logic as my Kestra blueprints. I picked a mix to see what matters most: size, "thinking", or being open-weight. Grok 4.6 and gpt-oss-120b were unavailable through Kaggle's model proxy while I ran this, so they're not on the board. Overall average of task scores : GPT-5.5 1.00 · Gemini 3.7 Flash 1.00 · Gemini 3.5 Flash 0.99 · Claude Sonnet 5 0.98 · Gemma 4 0.97 · DeepSeek R1 0.96 · gpt-oss-20b 0.92 · Gemini 3.1 Flash-Lite 0.75 · Qwen3 235B 0.51 · Claude Haiku 4.5 0.45 · GPT-5.4 nano 0.34 This was my favourite result. Under the CGST Act, section 2 103 , a Union Territory with its own legislature Puducherry, Delhi counts as a State . So a sale inside Puducherry is CGST + SGST . UTGST is only for UTs without a legislature, like Chandigarh or Ladakh. Bigger isn't uniformly better. Models fail on different local rules. On the GSTIN check character, every model that reasons step by step scored 100% , and every fast model scored 6–12% . That's no better than guessing. Small open models that reason gpt-oss-20b, Gemma 4 beat bigger fast ones. In an early test, one model needed about 15,000 thinking tokens and 84 seconds for a single GSTIN . Correct isn't always cheap. A supply to a Special Economic Zone unit in your own state is still an inter-state supply, so it's IGST IGST Act, section 7 5 b . Claude Haiku, gpt-oss-20b and GPT-5.4 nano split it into CGST + SGST anyway. The small models made the most dangerous kind of mistake, being off by exactly one digit group: 500000000 it's 50000000 986.4 it's 98.64 8,263,425,718 instead of 8,26,34,25,718 On GST invoice numbers, the look-alikes caught the small models: a trailing space, a 17-character number, and a number with an extra leading zero a different number, so it's allowed that they called a duplicate. All the big models were perfect here. On days with 18–30 UPI bills, weaker models missed bills whose money never arrived Claude Haiku 44%, Qwen 25% . A false alarm costs a minute. A missed bill is silent lost money. On Kaggle's score-vs-cost chart, the efficient frontier runs through the open-weight gpt-oss-20b and Gemma 4 31B . They get most of the way to the top score at a small fraction of the cost. 👉 Indian Business Data: GST, UPI and Lakh-Crore on Kaggle https://www.kaggle.com/benchmarks/ankith111111111/indian-business-data-gst-upi-and-lakh-crore All 6 tasks and their notebooks are public, so you can run them on any model. I used an AI assistant to draft this post. The idea, benchmarking, the Indian accounting rules and the final review are mine.