Batch LLM Jobs Without Breaking the Bank: A Queue-First Architecture for Free Tiers An engineer detailed a queue-first architecture that turns free-tier LLM quotas into batch engines, using a SQLite-backed task queue and Python workers to achieve ten times the throughput of synchronous calls. The approach inverts control flow by writing tasks to a queue instead of waiting per request, avoiding timeout waste and retry feedback loops. Free model quotas have a hidden enemy: synchronous calls. Every request blocks on the network. Timeouts get wasted. Retries pile up. Core conclusion: an asynchronous queue turns a free model quota from a demo tool into a batch engine. A SQLite-backed queue, a few Python workers, and a dead-letter table can extract ten times the throughput of synchronous code. Synchronous calls look simple. Send a request, wait for a response, process the result. The problem is the waiting. A 30-second timeout means 30 seconds of idle. Ten concurrent requests mean ten threads waiting. Free endpoints get slow, so your threads wait longer. Soon your application becomes a queueing system without the reliability of a queue. Retries make it worse. A timeout triggers a retry. The retry adds load. The load causes more timeouts. It is a feedback loop. The fix is to invert the control flow. Instead of waiting for a response per task, write tasks to a queue and let workers process them. The architecture is deliberately simple: SQLite is a deliberate choice. It is everywhere, zero-config, and handles thousands of tasks without complaint. You do not need Redis for a batch pipeline. python queue.py — SQLite-backed task queue for LLM batch jobs import sqlite3, time, uuid from pathlib import Path DB PATH = Path "tasks.db" def init db : conn = sqlite3.connect DB PATH conn.execute """ CREATE TABLE IF NOT EXISTS tasks id TEXT PRIMARY KEY, prompt TEXT NOT NULL, status TEXT DEFAULT 'pending', attempts INTEGER DEFAULT 0, result TEXT, error TEXT, created at REAL, updated at REAL """ conn.commit conn.close def enqueue prompt : conn = sqlite3.connect DB PATH task id = str uuid.uuid4 conn.execute "INSERT INTO tasks id, prompt, status, created at, updated at VALUES ?, ?, 'pending', ?, ? ", task id, prompt, time.time , time.time conn.commit conn.close return task id def claim next : conn = sqlite3.connect DB PATH row = conn.execute "SELECT id, prompt FROM tasks WHERE status = 'pending' ORDER BY created at LIMIT 1" .fetchone if row: conn.execute "UPDATE tasks SET status = 'running', updated at = ? WHERE id = ?", time.time , row 0 conn.commit conn.close return row The worker loop: python worker.py — process tasks from the queue import os, time, sqlite3, httpx from queue import init db, claim next def complete task id, result : conn = sqlite3.connect "tasks.db" conn.execute "UPDATE tasks SET status = 'done', result = ?, updated at = ? WHERE id = ?", result, time.time , task id conn.commit conn.close def fail task id, error : conn = sqlite3.connect "tasks.db" conn.execute "UPDATE tasks SET status = 'failed', error = ?, updated at = ? WHERE id = ?", error, time.time , task id conn.commit conn.close def process prompt : url = os.environ "LLM URL" key = os.environ "LLM KEY" model = os.environ "LLM MODEL" with httpx.Client timeout=60 as client: r = client.post url, headers={"Authorization": f"Bearer {key}"}, json={"model": model, "messages": {"role": "user", "content": prompt} , "max tokens": 200}, r.raise for status return r.json "choices" 0 "message" "content" def main : init db while True: row = claim next if not row: time.sleep 1 continue task id, prompt = row try: result = process prompt complete task id, result except Exception as exc: fail task id, type exc . name if name == " main ": main Run three workers against one database: python worker.py & python worker.py & python worker.py & Three workers, one database. Each worker handles tasks independently. If one crashes, the tasks stay in the database. Failed tasks need a second chance. Unlimited retries become an infinite loop. Add a retry counter. After three attempts, move the task to a dead-letter state. python def main : init db while True: row = claim next if not row: time.sleep 1 continue task id, prompt = row try: result = process prompt complete task id, result except Exception as exc: conn = sqlite3.connect "tasks.db" conn.execute "UPDATE tasks SET attempts = attempts + 1, status = 'pending' WHERE id = ?", task id, conn.commit conn.close Simple but effective. A task fails, the counter increments, the task returns to the queue. After three failures, inspect it manually. The queue's value is throughput. Synchronous code wastes idle time waiting for responses. Queue code is always working. Load the queue with a batch: python enqueue batch.py — load the queue with test tasks from queue import init db, enqueue init db for i in range 100 : enqueue f"Write a one-line Python function that returns {i}" print "queued 100 tasks" time python enqueue batch.py time python worker.py run until the queue is empty Compare that wall-clock time against a synchronous loop that sends 100 requests and waits for each. The gap is your free quota's real capacity. MonkeyCode's free server option is a natural home for these workers. It stays online, it is free, and it can reach the same network. Deploy the worker as a background process and let it drain the queue. The free model access supplies the tokens for processing. The advertised 10 million token figure is a starting point, not a contract. Disclosure: This article was prepared as part of MonkeyCode's product outreach. SQLite is not a high-throughput queue. It performs well at thousands of tasks and degrades at millions. For production-scale workloads, use a real message broker. The free server is shared infrastructure. Quotas shift. Latency fluctuates. Your results will differ from mine. Who should not use this pattern? Teams that need real-time responses. Teams that need transactional guarantees. Teams that need throughput beyond what a free tier can deliver. Who should use it? Teams processing batch jobs. People building prototypes. Anyone who wants actual work out of a free quota. Synchronous calls are the enemy of free. Queues are the friend of free. Move work out of waiting, into a database, and let workers chew through it. The pattern is not new. It is just rarely applied to LLM calls. Now you have the code. Run it, measure your throughput, and share the numbers.