cd /news/large-language-models/integrating-large-language-models-in… · home › topics › large-language-models › article
[ARTICLE · art-145683] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Integrating Large Language Models into SaaS Products: Practical Patterns and Pitfalls

A developer outlined three production patterns for integrating large language models into SaaS products: a retry-and-backoff HTTP wrapper for provider APIs, retrieval-augmented generation pairing a FAISS vector store with OpenAI embeddings, and batched multi-prompt requests for bulk generation tasks. The writeup includes Ruby, Python and Go code samples along with pitfalls such as never exposing API keys client-side, keeping embeddings fresh, and staying within token budgets.

by read3 min views1 publishedOct 5, 2026

Large language models (LLMs) have moved from research labs to production services. SaaS companies can add powerful text generation, summarization, or code assistance features, but the path from prototype to reliable service is full of decisions. This article walks through three patterns that work well in production, shows code snippets, and highlights common pitfalls to avoid.

The simplest way to add LLM capability is to wrap the provider’s HTTP API in a small library. The wrapper should handle retries, exponential back-off, and a timeout that matches your SLA. Below is a minimal Ruby example using net/http.

require 'net/http'
require 'json'

module LlmClient
  END_URL = URI('https://api.example.com/v1/completions')
  API_KEY = ENV['LLM_API_KEY']

  def self.complete(prompt, max_tokens: 200)
    payload = {
      model: 'gpt-4',
      prompt: prompt,
      max_tokens: max_tokens,
      temperature: 0.2
    }
    request = Net::HTTP::Post.new(API_URL)
    request['Authorization'] = "Bearer #{API_KEY}"
    request['Content-Type'] = 'application/json'
    request.body = payload.to_json
    response = Net::HTTP.start(API_URL.host, API_URL.port, use_ssl: true) do |http|
      http.request(request)
    end
    raise "LLM error #{response.code}" unless response.is_a?(Net::HTTPSuccess)
    JSON.parse(response.body)['choices'][0]['text'].strip
  end
end

Pitfalls: Do not expose the raw API key to the client side. Cache the wrapper instance per request to avoid repeated socket creation. Log the request and response IDs for observability.

For domain-specific knowledge, raw LLM prompts are insufficient. RAG combines a vector store of embeddings with the model to retrieve relevant passages before generation. The flow is:

Below is a Python snippet using faiss for similarity search and openai for completion.

import openai, faiss, numpy as np

 embed(text):
    resp = openai.Embedding.create(input=text, model='text-embedding-ada-002')
    return np.array(resp['data'][0]['embedding']).astype('float32')

 search(query, index, docs, k=3):
    q_vec = embed(query)
    distances, ids = index.search(np.expand_dims(q_vec, 0), k)
    return "\n".join([docs[i] for i in ids[0]])

def rag_completion(query, index, docs):
    context = search(query, index, docs)
    prompt = f"Context:\n{context}\n\nQuestion: {query}\nAnswer:"
    resp = openai.Completion.create(model='gpt-4', prompt=prompt, max_tokens=250)
    return resp['choices'][0]['text'].strip()

Pitfalls: Keep the vector store up-to-date; stale embeddings lead to irrelevant answers. Limit the amount of retrieved text to stay within the LLM token budget. Monitor latency; a FAISS search adds milliseconds but can be mitigated with caching.

When the use case involves bulk generation - such as summarizing thousands of support tickets - sending one request per item is wasteful. Batch the inputs and call the LLM with a single request that contains an array of prompts. The provider often returns an array of completions in the same order.

package llm

import (
    "bytes"
    "encoding/json"
    "net/http"
)

type BatchRequest struct {
    Model string `json:"model"`
    Prompts []string `json:"prompt"`
    MaxTokens int `json:"max_tokens"`
}

type BatchResponse struct {
    Choices []struct{ Text string `json:"text"` } `json:"choices"`
}

func CompleteBatch(prompts []string) ([]string, error) {
    reqBody:= BatchRequest{Model: "gpt-4", Prompts: prompts, MaxTokens: 150}
    data, _:= json.Marshal(reqBody)
    resp, err:= http.Post("https://api.example.com/v1/completions", "application/json", bytes.NewReader(data))
    if err!= nil { return nil, err }
    defer resp.Body.Close()
    var batchResp BatchResponse
    json.NewDecoder(resp.Body).Decode(&batchResp)
    results:= make([]string, len(batchResp.Choices))
    for i, c:= range batchResp.Choices {
        results[i] = c.Text
    }
    return results, nil
}

Pitfalls: The provider may impose a maximum number of prompts per batch; split large jobs accordingly. Preserve the order of inputs to match outputs. Use a job queue (e.g., Sidekiq or RabbitMQ) to handle retries and back-pressure.

Production LLM services generate costs that can grow quickly. Instrument every call with tags for model, token count, and latency. Set alerts on cost spikes and on error rates. A simple Prometheus metric can capture token usage:

llm_tokens_total{model="gpt-4",status="success"} 12345

Pitfalls: Do not rely on provider dashboards alone; they lag behind real traffic. Include request IDs in logs to correlate LLM calls with downstream processing.

LLM providers may retain input data for model improvement. If your SaaS handles PII, encrypt the payload before sending it, or use a provider that offers a “no-log” contract. Store only the hash of the input for audit purposes.

Pitfalls: Forgetting to redact sensitive fields before logging can leak data. Verify the provider’s data-handling policy in the contract.

Integrating LLMs into a SaaS product is more than a single API call. A robust wrapper, retrieval-augmented generation, batch processing, observability, and security together form a production-ready stack. By following the patterns and watching out for the listed pitfalls, engineering teams can deliver reliable AI features without surprise cost or downtime.

Author: senior engineer at developerz.ai, building AI-enhanced SaaS platforms.

── more in #large-language-models 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/integrating-large-la…] indexed:0 read:3min 2026-10-05 · —