{"slug": "i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got", "title": "I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.", "summary": "A developer built a six-task Kaggle benchmark of 16 cases each to test 11 AI models on everyday Indian business data, including GSTIN check-digit validation, GST tax splits, invoice-number rules, lakh/crore formatting and UPI reconciliation. GPT-5.5 and Gemini 3.7 Flash scored a perfect 1.00 average, while GPT-5.4 nano trailed at 0.34; models that reasoned step by step hit 100% on the GSTIN check character versus 6–12% for fast models, and several models wrongly split a Puducherry intra-state sale into CGST + SGST instead of recognizing the Union Territory with a legislature as a state.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*\n\nI build software for Indian small businesses (POS, billing, payments), and this month I've been contributing Kestra workflow blueprints for Indian finance: GSTIN validation, UPI end-of-day close, GSTR-2B reconciliation. Every one of them exists because one small mistake costs real money:\n\nSo I asked: **can AI models handle everyday Indian business data?** I built a Kaggle benchmark with 6 tasks, 16 cases each:\n\n| Task | What the model has to do | \n|---|---|\n| `india_gstin_check_digit` | Validate a GSTIN and compute its Luhn mod 36 check character (algorithm given) | \n| `india_gstin_from_memory` | Same, without being told the algorithm | \n| `india_gst_tax_split` | Split GST into CGST + SGST, CGST + UTGST or IGST, with real traps: supplies to an SEZ unit in your own state, and Union Territories with and without a legislature | \n| `india_invoice_number_rules` | Apply GST invoice-number rules (16 characters, only letters/digits/ `-` /`/` , unique per financial year) with look-alike traps: en dash, division slash, trailing space, 17 characters, extra leading zero | \n| `india_lakh_crore_formats` | Indian digit grouping (12,34,56,789), lakh/crore conversions, cheque amounts in words | \n| `india_upi_reconciliation` | Match 18–30 UPI bills against a settlement report and list the bills whose money never arrived, arrived short, or share a UPI reference | \n\nHow it's graded:\n\n`python-stdnum` (0 disagreements). The other answer keys use the same logic as my Kestra blueprints.\nI picked a mix to see what matters most: size, \"thinking\", or being open-weight.\n\n(Grok 4.6 and gpt-oss-120b were unavailable through Kaggle's model proxy while I ran this, so they're not on the board.)\n\nOverall (average of task scores): **GPT-5.5 1.00 · Gemini 3.7 Flash 1.00 · Gemini 3.5 Flash 0.99 · Claude Sonnet 5 0.98 · Gemma 4 0.97 · DeepSeek R1 0.96 · gpt-oss-20b 0.92 · Gemini 3.1 Flash-Lite 0.75 · Qwen3 235B 0.51 · Claude Haiku 4.5 0.45 · GPT-5.4 nano 0.34**\n\nThis was my favourite result. Under the CGST Act, section 2(103), a Union Territory **with its own legislature** (Puducherry, Delhi) counts as a *State*. So a sale inside Puducherry is **CGST + SGST**. UTGST is only for UTs without a legislature, like Chandigarh or Ladakh.\n\nBigger isn't uniformly better. Models fail on *different* local rules.\n\nOn the GSTIN check character, every model that reasons step by step scored **100%**, and every fast model scored **6–12%**. That's no better than guessing. Small open models that reason (gpt-oss-20b, Gemma 4) beat bigger fast ones. In an early test, one model needed about **15,000 thinking tokens and 84 seconds for a single GSTIN**. Correct isn't always cheap.\n\nA supply to a Special Economic Zone unit in your *own* state is still an inter-state supply, so it's IGST (IGST Act, section 7(5)(b)). Claude Haiku, gpt-oss-20b and GPT-5.4 nano split it into CGST + SGST anyway.\n\nThe small models made the most dangerous kind of mistake, being off by exactly one digit group:\n\n`500000000` (it's `50000000`)` 986.4` (it's `98.64`)` 8,263,425,718` instead of `8,26,34,25,718`\nOn GST invoice numbers, the look-alikes caught the small models: a trailing space, a 17-character number, and a number with an extra leading zero (a different number, so it's allowed) that they called a duplicate. All the big models were perfect here.\n\nOn days with 18–30 UPI bills, weaker models **missed** bills whose money never arrived (Claude Haiku 44%, Qwen 25%). A false alarm costs a minute. A missed bill is silent lost money.\n\nOn Kaggle's score-vs-cost chart, the efficient frontier runs through the open-weight **gpt-oss-20b** and **Gemma 4 31B**. They get most of the way to the top score at a small fraction of the cost.\n\n👉 [Indian Business Data: GST, UPI and Lakh-Crore on Kaggle](https://www.kaggle.com/benchmarks/ankith111111111/indian-business-data-gst-upi-and-lakh-crore)\n\nAll 6 tasks and their notebooks are public, so you can run them on any model.\n\n*I used an AI assistant to draft this post. The idea, benchmarking,  the Indian accounting rules and the final review are mine.*", "url": "https://wpnews.pro/news/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got", "canonical_source": "https://dev.to/ankithm1006/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got-puducherry-wrong-53e9", "published_at": "2026-10-06 10:34:02+00:00", "updated_at": "2026-10-06 10:48:10.454887+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools"], "entities": ["Kaggle", "GPT-5.5", "Gemini 3.7 Flash", "Claude Sonnet 5", "Gemma 4", "DeepSeek R1", "gpt-oss-20b", "Qwen3 235B"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got", "markdown": "https://wpnews.pro/news/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got.md", "text": "https://wpnews.pro/news/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got.txt", "jsonld": "https://wpnews.pro/news/i-tested-11-ai-models-on-indian-gst-upi-and-lakh-crore-three-famous-ones-got.jsonld"}}