Large language models (LLMs) have moved from research labs to production services. SaaS companies can add powerful text generation, summarization, or code assistance features, but the path from prototype to reliable service is full of decisions. This article walks through three patterns that work well in production, shows code snippets, and highlights common pitfalls to avoid.
The simplest way to add LLM capability is to wrap the provider’s HTTP API in a small library. The wrapper should handle retries, exponential back-off, and a timeout that matches your SLA. Below is a minimal Ruby example using net/http.
require 'net/http'
require 'json'
module LlmClient
END_URL = URI('https://api.example.com/v1/completions')
API_KEY = ENV['LLM_API_KEY']
def self.complete(prompt, max_tokens: 200)
payload = {
model: 'gpt-4',
prompt: prompt,
max_tokens: max_tokens,
temperature: 0.2
}
request = Net::HTTP::Post.new(API_URL)
request['Authorization'] = "Bearer #{API_KEY}"
request['Content-Type'] = 'application/json'
request.body = payload.to_json
response = Net::HTTP.start(API_URL.host, API_URL.port, use_ssl: true) do |http|
http.request(request)
end
raise "LLM error #{response.code}" unless response.is_a?(Net::HTTPSuccess)
JSON.parse(response.body)['choices'][0]['text'].strip
end
end
Pitfalls: Do not expose the raw API key to the client side. Cache the wrapper instance per request to avoid repeated socket creation. Log the request and response IDs for observability.
For domain-specific knowledge, raw LLM prompts are insufficient. RAG combines a vector store of embeddings with the model to retrieve relevant passages before generation. The flow is:
Below is a Python snippet using faiss for similarity search and openai for completion.
import openai, faiss, numpy as np
embed(text):
resp = openai.Embedding.create(input=text, model='text-embedding-ada-002')
return np.array(resp['data'][0]['embedding']).astype('float32')
search(query, index, docs, k=3):
q_vec = embed(query)
distances, ids = index.search(np.expand_dims(q_vec, 0), k)
return "\n".join([docs[i] for i in ids[0]])
def rag_completion(query, index, docs):
context = search(query, index, docs)
prompt = f"Context:\n{context}\n\nQuestion: {query}\nAnswer:"
resp = openai.Completion.create(model='gpt-4', prompt=prompt, max_tokens=250)
return resp['choices'][0]['text'].strip()
Pitfalls: Keep the vector store up-to-date; stale embeddings lead to irrelevant answers. Limit the amount of retrieved text to stay within the LLM token budget. Monitor latency; a FAISS search adds milliseconds but can be mitigated with caching.
When the use case involves bulk generation - such as summarizing thousands of support tickets - sending one request per item is wasteful. Batch the inputs and call the LLM with a single request that contains an array of prompts. The provider often returns an array of completions in the same order.
package llm
import (
"bytes"
"encoding/json"
"net/http"
)
type BatchRequest struct {
Model string `json:"model"`
Prompts []string `json:"prompt"`
MaxTokens int `json:"max_tokens"`
}
type BatchResponse struct {
Choices []struct{ Text string `json:"text"` } `json:"choices"`
}
func CompleteBatch(prompts []string) ([]string, error) {
reqBody:= BatchRequest{Model: "gpt-4", Prompts: prompts, MaxTokens: 150}
data, _:= json.Marshal(reqBody)
resp, err:= http.Post("https://api.example.com/v1/completions", "application/json", bytes.NewReader(data))
if err!= nil { return nil, err }
defer resp.Body.Close()
var batchResp BatchResponse
json.NewDecoder(resp.Body).Decode(&batchResp)
results:= make([]string, len(batchResp.Choices))
for i, c:= range batchResp.Choices {
results[i] = c.Text
}
return results, nil
}
Pitfalls: The provider may impose a maximum number of prompts per batch; split large jobs accordingly. Preserve the order of inputs to match outputs. Use a job queue (e.g., Sidekiq or RabbitMQ) to handle retries and back-pressure.
Production LLM services generate costs that can grow quickly. Instrument every call with tags for model, token count, and latency. Set alerts on cost spikes and on error rates. A simple Prometheus metric can capture token usage:
llm_tokens_total{model="gpt-4",status="success"} 12345
Pitfalls: Do not rely on provider dashboards alone; they lag behind real traffic. Include request IDs in logs to correlate LLM calls with downstream processing.
LLM providers may retain input data for model improvement. If your SaaS handles PII, encrypt the payload before sending it, or use a provider that offers a “no-log” contract. Store only the hash of the input for audit purposes.
Pitfalls: Forgetting to redact sensitive fields before logging can leak data. Verify the provider’s data-handling policy in the contract.
Integrating LLMs into a SaaS product is more than a single API call. A robust wrapper, retrieval-augmented generation, batch processing, observability, and security together form a production-ready stack. By following the patterns and watching out for the listed pitfalls, engineering teams can deliver reliable AI features without surprise cost or downtime.
Author: senior engineer at developerz.ai, building AI-enhanced SaaS platforms.