Building PasteDB: A Bulletproof Code-Sharing Platform Powered by FastAPI and OpenAI's gpt-oss-20b A developer built PasteDB, a lightweight code-sharing platform with an "Explain This Code" AI assistant layer, using FastAPI on the backend and streaming requests to Groq's cloud infrastructure to run an open-weight model. The project works around Render's 512MB free-tier RAM limit by keeping the model off the server, and adds a custom in-memory IP rate limiter capping users at 3 requests per minute per IP to protect the public API quota. This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend https://dev.to/challenges/hacktoberfest-weekend-2026-10-01 I built PasteDB , a lightweight code-sharing platform designed to make sharing code blocks seamless. To take it a step further for this challenge, I integrated an "Explain This Code" AI assistant layer. Who I built it for: I built this feature specifically for my developer friends and peers who are learning to code. Often, when we share raw code snippets over chat platforms, beginners struggle to understand the core logic without context. PasteDB now automatically analyzes any shared snippet at the click of a button, acting as an on-demand mentor. Live Link: https://pastedb.netlify.app https://pastedb.netlify.app Backend API: https://pastedb-rw62.onrender.com https://pastedb-rw62.onrender.com Here is the official open-source repository for PasteDB: https://github.com/sorathiya903/pastedb https://github.com/sorathiya903/pastedb PasteDB is engineered using a robust, free-tier distributed stack: The biggest engineering hurdle was Render's strict 512MB RAM constraint on the free tier. Running even a small 350MB model locally on the backend would trigger an Out-of-Memory OOM crash once the Python dependencies and KV caches loaded. To solve this, I decoupled the compute by making outbound streaming requests to Groq's cloud infrastructure to tap into the official open-weight OpenAI Model model. This leaves a 0MB memory footprint on Render while generating lightning-fast, structured Markdown explanations for the user. To protect my public API quota from malicious spam or heavy judging traffic, I engineered a zero-RAM, custom in-memory IP Rate Limiter directly inside the FastAPI routing layer. It limits users to 3 requests per minute per IP, protecting the app against HTTP 429 exhaustion while keeping the experience completely smooth and available for the judges. A snippet of my custom rate-limiter guarding the open-weight pipeline Open innovation was the entire foundation of this project. Relying on premium, closed-source corporate APIs forces developers into commercial paywalls, rigid token meters, and strict monetization loops from day one. By utilizing open-weight ecosystem components like Llama 3.1 , I was able to build a completely free, highly scalable utility tool for my friends. Open innovation democratizes AI execution, proving that independent developers can ship fully secured AI tools without a massive corporate budget. I am entering PasteDB into the following categories: gpt-oss-20b across a modern full-stack ecosystem consisting of