Buddy Base:Built an LLM Powered Group Chat for My Friends to Make Plans This Weekend A developer built BuddyBase, a private, offline Django web application that gives a friend group a shared AI assistant running entirely on a laptop over a local network. The app uses a locally loaded Qwen model via HuggingFace with lru_cache and ThreadPoolExecutor timeout handling, and offers four features: group chat with an @buddy AI summon, an AI-assisted decision maker, smart shared lists, and per-user personal assistants with persistent preference memory. The developer says the project is not production-ready but plans to enhance it. This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend https://dev.to/challenges/hacktoberfest-weekend-2026-10-01 BuddyBase — a private, offline web application that gives your friend group a shared AI assistant. It runs entirely on your laptop. Friends connect through your local network WiFi, Bluetooth tethering, or mobile hotspot . No data ever leaves your machine. Four Features, All AI-Powered Group Chat with AIA real-time chatroom where anyone can summon the AI by typing @buddy https://dev.to/buddy . It knows your group name, all members, and the last 20 messages for context. It responds like a friend who's always around — casual, funny, and actually helpful. Decision MakerCreate a question with options. AI immediately gives a recommendation with reasoning. Everyone votes. After each vote, AI updates its recommendation considering the majority. No more 40-minute indecision loops. Smart Shared ListsCollaborative lists where anyone can add items. Ask the AI to "suggest 5 more movies like these" or "what are we forgetting for the road trip?" It reads the list and gives real suggestions. Personal BuddyEach friend gets a private AI assistant that learns their preferences. Tell it "I love spicy food" and next time someone asks about dinner, your buddy already knows. Memories persist in the database forever. cannot make it in production but will be surely enhanced it and make it to production github link : https://github.com/sanjaynep/django-websocket https://github.com/sanjaynep/django-websocket This is the core — loads Qwen once using lru cache , generates responses using the same apply chat template pattern from HuggingFace, with timeout handling via ThreadPoolExecutor : python @lru cache maxsize=1 def load llm : tokenizer = AutoTokenizer.from pretrained MODEL NAME model = AutoModelForCausalLM.from pretrained MODEL NAME, torch dtype=torch.float32, device map={"": "cpu"}, model.eval return tokenizer, model def generate messages : tokenizer, model = load llm text = tokenizer.apply chat template messages, tokenize=False, add generation prompt=True model inputs = tokenizer text , return tensors="pt" .to "cpu" generated ids = model.generate model inputs, max new tokens=256, do sample=True output ids = generated ids 0 len model inputs.input ids 0 : .tolist return tokenizer.decode output ids, skip special tokens=True .strip def chat messages : with ThreadPoolExecutor max workers=1 as executor: future = executor.submit generate, messages return future.result timeout=60 Key Code: WebSocket Consumer python class GroupChatConsumer AsyncWebsocketConsumer : async def receive self, text data : message = json.loads text data .get 'message', '' .strip user = self.scope 'user' chat msg = await self. save message user, message, is ai=False await self.channel layer.group send self.room group name, { 'type': 'chat message', 'message': message, 'username': user.username, 'is ai': False, 'time': chat msg.created at.strftime '%H:%M' , } if '@buddy' in message.lower : await self.channel layer.group send self.room group name, { 'type': 'typing indicator', } ai response = await self. get ai response ai msg = await self. save message None, ai response, is ai=True await self.channel layer.group send self.room group name, { 'type': 'chat message', 'message': ai response, 'username': 'Buddy', 'is ai': True, 'time': ai msg.created at.strftime '%H:%M' , } Key Code: Preference Memory python def try save memory user, message : prefixes = { 'i love': 'loves', 'i hate': 'hates', 'i like': 'likes', "i don't like": 'dislikes', 'my favorite': 'favorite', } for prefix, key prefix in prefixes.items : if prefix in message.lower : value = message message.lower .index prefix + len prefix : .strip BuddyMemory.objects.update or create user=user, key=f"{key prefix} {value :20 }", defaults={'value': value} Tech Stack Layer Technology Why Backend Django 4.2 Clean MVC, auth, ORM, admin panel out of the box Real-Time Django Channels + WebSocket Messages appear instantly, no page reload AI Model Qwen2.5-1.5B-Instruct via HuggingFace Transformers Runs locally on CPU, 3GB RAM, good chat quality Frontend Vanilla JS + CSS Zero dependencies, works offline, no CDN needed Database SQLite Local file, zero setup, perfect for small group Server Daphne ASGI Supports both HTTP and WebSocket Open-Source AI I Used Component What LLM Qwen2.5-1.5B-Instruct by Alibaba Inference HuggingFace Transformers PyTorch PyTorch CPU build Framework Django WebSocket Django Channels Privacy is not a feature, it's a fundamental right. If I used OpenAI's API or Google's Gemini API for this project, every message my friends send — their food preferences, their votes, their inside jokes, their travel plans — would be sent to a third-party server. That data would be logged, possibly used for training, and stored in someone else's database forever. With open-weight models like Qwen, that doesn't happen. My friends' data stays on my laptop. In a SQLite file that I control. I can delete it, back it up, or move it. Nobody else has access. No API keys to leak. There are no API keys in the code. No .env file with secrets. No billing account to worry about. No internet dependency. We used this at a cabin with no WiFi. I turned on mobile hotspot, everyone connected, and the AI still worked because Qwen runs locally. No usage limits. No "you've exceeded your quota" errors at 3 AM when we're still planning. The model runs as long as my laptop is on. Full customization. I can change the AI's personality, what it remembers, how it makes decisions. Try getting ChatGPT to always pick a restaurant based on your group's dietary restrictions. With open models, I just edit the system prompt. Zero cost. Running Qwen2.5-1.5B on my laptop costs exactly $0. Forever. For unlimited messages, unlimited friends, unlimited decisions. Open innovation made this possible because: Qwen's Apache 2.0 license lets me use it commercially or personally without restrictions HuggingFace Transformers provides the inference engine for free Django's BSD license gives me a production-grade web framework for free Django Channels adds real-time capabilities for free A closed API would make this project either impossible no offline mode , expensive paying per message for 5 friends chatting all weekend , or privacy-violating sending personal preferences to the cloud . Open models make it free, private, and offline.