cd /news/large-language-models/batch-llm-jobs-without-breaking-the-… · home topics large-language-models article
[ARTICLE · art-107182] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Batch LLM Jobs Without Breaking the Bank: A Queue-First Architecture for Free Tiers

An engineer detailed a queue-first architecture that turns free-tier LLM quotas into batch engines, using a SQLite-backed task queue and Python workers to achieve ten times the throughput of synchronous calls. The approach inverts control flow by writing tasks to a queue instead of waiting per request, avoiding timeout waste and retry feedback loops.

read5 min views1 publishedAug 22, 2026

Free model quotas have a hidden enemy: synchronous calls. Every request blocks on the network. Timeouts get wasted. Retries pile up.

Core conclusion: an asynchronous queue turns a free model quota from a demo tool into a batch engine. A SQLite-backed queue, a few Python workers, and a dead-letter table can extract ten times the throughput of synchronous code.

Synchronous calls look simple. Send a request, wait for a response, process the result. The problem is the waiting.

A 30-second timeout means 30 seconds of idle. Ten concurrent requests mean ten threads waiting. Free endpoints get slow, so your threads wait longer. Soon your application becomes a queueing system without the reliability of a queue.

Retries make it worse. A timeout triggers a retry. The retry adds load. The load causes more timeouts. It is a feedback loop.

The fix is to invert the control flow. Instead of waiting for a response per task, write tasks to a queue and let workers process them.

The architecture is deliberately simple:

SQLite is a deliberate choice. It is everywhere, zero-config, and handles thousands of tasks without complaint. You do not need Redis for a batch pipeline.

import sqlite3, time, uuid
from pathlib import Path

DB_PATH = Path("tasks.db")

def init_db():
    conn = sqlite3.connect(DB_PATH)
    conn.execute("""
        CREATE TABLE IF NOT EXISTS tasks (
            id TEXT PRIMARY KEY,
            prompt TEXT NOT NULL,
            status TEXT DEFAULT 'pending',
            attempts INTEGER DEFAULT 0,
            result TEXT,
            error TEXT,
            created_at REAL,
            updated_at REAL
        )
    """)
    conn.commit()
    conn.close()

def enqueue(prompt):
    conn = sqlite3.connect(DB_PATH)
    task_id = str(uuid.uuid4())
    conn.execute(
        "INSERT INTO tasks (id, prompt, status, created_at, updated_at) VALUES (?, ?, 'pending', ?, ?)",
        (task_id, prompt, time.time(), time.time())
    )
    conn.commit()
    conn.close()
    return task_id

def claim_next():
    conn = sqlite3.connect(DB_PATH)
    row = conn.execute(
        "SELECT id, prompt FROM tasks WHERE status = 'pending' ORDER BY created_at LIMIT 1"
    ).fetchone()
    if row:
        conn.execute(
            "UPDATE tasks SET status = 'running', updated_at = ? WHERE id = ?",
            (time.time(), row[0])
        )
        conn.commit()
    conn.close()
    return row

The worker loop:

import os, time, sqlite3, httpx
from queue import init_db, claim_next

def complete(task_id, result):
    conn = sqlite3.connect("tasks.db")
    conn.execute(
        "UPDATE tasks SET status = 'done', result = ?, updated_at = ? WHERE id = ?",
        (result, time.time(), task_id)
    )
    conn.commit()
    conn.close()

def fail(task_id, error):
    conn = sqlite3.connect("tasks.db")
    conn.execute(
        "UPDATE tasks SET status = 'failed', error = ?, updated_at = ? WHERE id = ?",
        (error, time.time(), task_id)
    )
    conn.commit()
    conn.close()

def process(prompt):
    url = os.environ["LLM_URL"]
    key = os.environ["LLM_KEY"]
    model = os.environ["LLM_MODEL"]
    with httpx.Client(timeout=60) as client:
        r = client.post(
            url,
            headers={"Authorization": f"Bearer {key}"},
            json={"model": model,
                  "messages": [{"role": "user", "content": prompt}],
                  "max_tokens": 200},
        )
        r.raise_for_status()
        return r.json()["choices"][0]["message"]["content"]

def main():
    init_db()
    while True:
        row = claim_next()
        if not row:
            time.sleep(1)
            continue
        task_id, prompt = row
        try:
            result = process(prompt)
            complete(task_id, result)
        except Exception as exc:
            fail(task_id, type(exc).__name__)

if __name__ == "__main__":
    main()

Run three workers against one database:

python worker.py &
python worker.py &
python worker.py &

Three workers, one database. Each worker handles tasks independently. If one crashes, the tasks stay in the database.

Failed tasks need a second chance. Unlimited retries become an infinite loop.

Add a retry counter. After three attempts, move the task to a dead-letter state.

def main():
    init_db()
    while True:
        row = claim_next()
        if not row:
            time.sleep(1)
            continue
        task_id, prompt = row
        try:
            result = process(prompt)
            complete(task_id, result)
        except Exception as exc:
            conn = sqlite3.connect("tasks.db")
            conn.execute(
                "UPDATE tasks SET attempts = attempts + 1, status = 'pending' WHERE id = ?",
                (task_id,)
            )
            conn.commit()
            conn.close()

Simple but effective. A task fails, the counter increments, the task returns to the queue. After three failures, inspect it manually.

The queue's value is throughput. Synchronous code wastes idle time waiting for responses. Queue code is always working.

Load the queue with a batch:

from queue import init_db, enqueue

init_db()
for i in range(100):
    enqueue(f"Write a one-line Python function that returns {i}")
print("queued 100 tasks")
time python enqueue_batch.py
time python worker.py  # run until the queue is empty

Compare that wall-clock time against a synchronous loop that sends 100 requests and waits for each. The gap is your free quota's real capacity.

MonkeyCode's free server option is a natural home for these workers. It stays online, it is free, and it can reach the same network. Deploy the worker as a background process and let it drain the queue.

The free model access supplies the tokens for processing. The advertised 10 million token figure is a starting point, not a contract.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

SQLite is not a high-throughput queue. It performs well at thousands of tasks and degrades at millions. For production-scale workloads, use a real message broker.

The free server is shared infrastructure. Quotas shift. Latency fluctuates. Your results will differ from mine.

Who should not use this pattern? Teams that need real-time responses. Teams that need transactional guarantees. Teams that need throughput beyond what a free tier can deliver.

Who should use it? Teams processing batch jobs. People building prototypes. Anyone who wants actual work out of a free quota.

Synchronous calls are the enemy of free. Queues are the friend of free. Move work out of waiting, into a database, and let workers chew through it.

The pattern is not new. It is just rarely applied to LLM calls. Now you have the code. Run it, measure your throughput, and share the numbers.

── more in #large-language-models 4 stories · sorted by recency
── more on @sqlite 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/batch-llm-jobs-witho…] indexed:0 read:5min 2026-08-22 ·