{"slug": "tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to", "title": "tl;dr - A local LLM that reads your college group chat so you don't have to.", "summary": "A developer built tl;dr, a local LLM tool that reads exported WhatsApp college group chats and extracts deadlines, events and plan changes into a markdown summary and importable calendar file. The pipeline runs Gemma 4 (gemma4:e4b) via Ollama entirely on-device, handling Hinglish date expressions like \"kal 5 baje\" and \"parso\" while reconciling conflicting deadline updates into a single item with message-level provenance. The developer said the project was built for a friend who struggled to keep up with assignment deadlines buried in a 1,000+ message chat.", "body_md": "*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\nMy friend Atharva has a hard time keeping up with college assignments and deadlines. The info he needs is almost always sitting in our class group chat, which is 1,000+ messages of memes, \"bhai, kl clg jana h?\" and a rickroll every few days. Somewhere in there is one line saying the deadline moved to Friday. If you miss that line, you miss the deadline.\n\nSo I built **tl;dr**. You export the WhatsApp chat, a model running on your own laptop reads it, and you get back just the deadlines, events and changes of plan. It writes a short markdown summary and a calendar file you can import.\n\nThe hard part is that people in the chat don't write clean dates. They write \"kal 5 baje\" and \"parso\" and \"ab Friday tak hai\", and the plan keeps changing. A message says the assignment is due Thursday, a later one says it got postponed to Friday, and later still someone says 1:30 tak. tl;dr should end up with one deadline and its history, not three that contradict each other.\n\nHonestly, I'm proud of this one. I get to use what I know to help someone, and that someone is my friend. That feels better than any project I've built just to have something on my GitHub.\n\nThese are from a synthetic sample chat, since the real one stays private. A chat line like\n\n```\nguys DBMS assignment 2 ka deadline Thursday hai, 1 Oct\n...\nUpdate: assignment deadline postpone ho gaya, ab Friday 2 Oct tak hai. Sir ne portal pe daala\n```\n\n(plus a couple more changes later) ends up as one card:\n\n```\n## Fri 02 Oct\n- **13:30** · DBMS assignment 2 deadline — due · msgs 3, 22, 35, 37, 50\n```\n\nEvery card lists the messages it came from, so you can always check it against the chat.\n\nWhen I showed him, he was properly surprised. He didn't know something like this was even possible, and I think he was even more surprised that I was the one who built it. That reaction is the best part of this whole project for me.\n\n**A local LLM that reads your college group chat so you don't have to.**\n\nGroup chats bury the one real deadline under hundreds of messages, memes and \"bhai sab aa gaye?\". tl;dr reads a WhatsApp export, extracts only the deadlines, events and changes of plan, and gives you a clean summary plus a calendar file you can import.\n\nIt runs **entirely on your machine** (Ollama + an open model). The chats it was\nbuilt for are group chats of friends, so nothing is ever sent to a cloud API\nIt handles **Hinglish**: \"kal 5 baje\", \"parso\", \"ab Friday tak hai\".\n\nBuilt for the DEV Hacktoberfest 2026 Weekend Challenge, *Build for a Friend*.\n\nInput (abridged, from `data/sample_chat.txt`):\n\n```\n[3]  Aditi: guys DBMS assignment 2 ka deadline Thursday hai, 1 Oct\n[22] Aditi: Update: assignment deadline postpone ho gaya, ab Friday 2 Oct tak hai\n[35] Aditi: kal\n```\n\n…\nIt's tagged `v0.1.0`, and the README has setup plus a one-command run on the sample chat.\n\nThe model is Gemma 4 (`gemma4:e4b`), running locally through Ollama with temperature 0, seed 0 and output forced into a JSON schema. The pipeline goes parse, split into windows, extract with the model, reconcile, resolve dates, render.\n\nThe chat gets split into windows wherever the conversation pauses, and the windows overlap a little so a change right on a boundary doesn't get lost. Each window sees the items found so far and says whether each new thing is `new`, an `update` of something known, or a `cancel`. Dates are tied to the message that mentioned them, so \"kal\" means tomorrow relative to when it was said, not relative to the day I run the tool.\n\nThe bit that taught me the most was not trusting the model's bookkeeping. A 4B model is great at reading Hinglish and pretty bad at keeping track of which item is which. It made up ids. It once put a title where an id should go. Worst of all, it took the messages about a makeup class and filed them under the DBMS paper, with the paper's title copied on top. The id matched and the title matched. Only the content was wrong, so no simple check could catch it.\n\nSo now there's a reconcile step that decides for itself whether something is already known, mostly by comparing titles, and treats the model's id as a hint. If it isn't sure, it makes a duplicate instead of merging, because a duplicate is easy to spot and a swallowed event isn't. I learned that the hard way: my first version let a plain \"Lab\" absorb a separate lab-record deadline, and I only noticed because the deadline disappeared. There's a trace flag now that prints every decision. This is one where the model wanted to attach a study session to the wrong thing and the matcher refused:\n\n```\n[match] NEW   'Study session for DBMS joins' (best 0.36 vs K2 'Lab session')\n```\n\nTo measure it I hand-labelled 7 items from the sample chat and wrote a small scorer. With the prompt frozen:\n\n| mode | items found | precision | date correct | time correct | per run | \n|---|---|---|---|---|---|\n| thinking off | 5 / 7 | 100% | 57% | 57% | ~58 s | \n| thinking on | 7 / 7 | 78% | 71% (range 71-100%) | 71% (range 29-86%) | ~330 s | \n\nThinking off gave identical results in 9 runs, but it misses events. Thinking on finds all 7, takes about 6x longer, adds a few extra items, and its dates change between runs. Pinning the seed mattered too: one unseeded run scored 6/7 with 86% on everything, and it was just luck. I also tried two prompt tweaks that I thought would help, a slimmer known-items list and an extra day/time rule. Both made things worse (time accuracy 57% to 43%, and 5/7 down to 4/7), so I reverted them. With only 7 items, one item is worth 14 points, and I tuned the prompt on this chat, so treat all of this as a demo and not a benchmark.\n\nThen I ran it on a real chat: 1,178 messages, 61 windows. It found 15 dated events and flagged 18 more as \"date unclear\". Clear, one-off events came out well. What broke:\n\n`01.12.2025` weren't parsed.\nA group chat is full of other people's names, plans and in-jokes. Running Gemma locally means none of my classmates' messages go to a server I don't control, and once the model is downloaded it all works on my own machine without internet.\n\nIt also made debugging possible. Because I could pin the seed, switch thinking on and off and read every raw output, I could see that the model was copying titles onto the wrong items, and that my first good-looking run was luck. And it costs nothing per run, which matters when you run the pipeline dozens of times while debugging.\n\nI didn't benchmark a closed model, so I can't say open is more accurate here. For this project the wins were privacy, cost and control.\n\nI used Claude as a pair-programmer. It wrote the title-matching and reconcile logic, the eval harness, several date-parsing fixes and the README with me, and it helped me draft this write-up. The idea, the extraction pipeline, the prompt, the experiments, running everything on real chats and deciding what to ship were mine.", "url": "https://wpnews.pro/news/tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to", "canonical_source": "https://dev.to/karansingh-in/tldr-a-local-llm-that-reads-your-college-group-chat-so-you-dont-have-to-3cnf", "published_at": "2026-10-03 21:25:15+00:00", "updated_at": "2026-10-03 21:37:49.498489+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "natural-language-processing", "generative-ai"], "entities": ["Ollama", "Gemma 4", "WhatsApp", "tl;dr", "DEV Hacktoberfest 2026", "Atharva"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to", "markdown": "https://wpnews.pro/news/tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to.md", "text": "https://wpnews.pro/news/tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to.txt", "jsonld": "https://wpnews.pro/news/tl-dr-a-local-llm-that-reads-your-college-group-chat-so-you-don-t-have-to.jsonld"}}