Integrating Large Language Models into SaaS Products: Practical Patterns and Pitfalls A developer outlined three production patterns for integrating large language models into SaaS products: a retry-and-backoff HTTP wrapper for provider APIs, retrieval-augmented generation pairing a FAISS vector store with OpenAI embeddings, and batched multi-prompt requests for bulk generation tasks. The writeup includes Ruby, Python and Go code samples along with pitfalls such as never exposing API keys client-side, keeping embeddings fresh, and staying within token budgets. Large language models LLMs have moved from research labs to production services. SaaS companies can add powerful text generation, summarization, or code assistance features, but the path from prototype to reliable service is full of decisions. This article walks through three patterns that work well in production, shows code snippets, and highlights common pitfalls to avoid. The simplest way to add LLM capability is to wrap the provider’s HTTP API in a small library. The wrapper should handle retries, exponential back-off, and a timeout that matches your SLA. Below is a minimal Ruby example using net/http . require 'net/http' require 'json' module LlmClient END URL = URI 'https://api.example.com/v1/completions' API KEY = ENV 'LLM API KEY' def self.complete prompt, max tokens: 200 payload = { model: 'gpt-4', prompt: prompt, max tokens: max tokens, temperature: 0.2 } request = Net::HTTP::Post.new API URL request 'Authorization' = "Bearer {API KEY}" request 'Content-Type' = 'application/json' request.body = payload.to json response = Net::HTTP.start API URL.host, API URL.port, use ssl: true do |http| http.request request end raise "LLM error {response.code}" unless response.is a? Net::HTTPSuccess JSON.parse response.body 'choices' 0 'text' .strip end end Pitfalls : Do not expose the raw API key to the client side. Cache the wrapper instance per request to avoid repeated socket creation. Log the request and response IDs for observability. For domain-specific knowledge, raw LLM prompts are insufficient. RAG combines a vector store of embeddings with the model to retrieve relevant passages before generation. The flow is: Below is a Python snippet using faiss for similarity search and openai for completion. python import openai, faiss, numpy as np embed text : resp = openai.Embedding.create input=text, model='text-embedding-ada-002' return np.array resp 'data' 0 'embedding' .astype 'float32' search query, index, docs, k=3 : q vec = embed query distances, ids = index.search np.expand dims q vec, 0 , k return "\n".join docs i for i in ids 0 def rag completion query, index, docs : context = search query, index, docs prompt = f"Context:\n{context}\n\nQuestion: {query}\nAnswer:" resp = openai.Completion.create model='gpt-4', prompt=prompt, max tokens=250 return resp 'choices' 0 'text' .strip Pitfalls : Keep the vector store up-to-date; stale embeddings lead to irrelevant answers. Limit the amount of retrieved text to stay within the LLM token budget. Monitor latency; a FAISS search adds milliseconds but can be mitigated with caching. When the use case involves bulk generation - such as summarizing thousands of support tickets - sending one request per item is wasteful. Batch the inputs and call the LLM with a single request that contains an array of prompts. The provider often returns an array of completions in the same order. package llm import "bytes" "encoding/json" "net/http" type BatchRequest struct { Model string json:"model" Prompts string json:"prompt" MaxTokens int json:"max tokens" } type BatchResponse struct { Choices struct{ Text string json:"text" } json:"choices" } func CompleteBatch prompts string string, error { reqBody:= BatchRequest{Model: "gpt-4", Prompts: prompts, MaxTokens: 150} data, := json.Marshal reqBody resp, err:= http.Post "https://api.example.com/v1/completions", "application/json", bytes.NewReader data if err = nil { return nil, err } defer resp.Body.Close var batchResp BatchResponse json.NewDecoder resp.Body .Decode &batchResp results:= make string, len batchResp.Choices for i, c:= range batchResp.Choices { results i = c.Text } return results, nil } Pitfalls : The provider may impose a maximum number of prompts per batch; split large jobs accordingly. Preserve the order of inputs to match outputs. Use a job queue e.g., Sidekiq or RabbitMQ to handle retries and back-pressure. Production LLM services generate costs that can grow quickly. Instrument every call with tags for model, token count, and latency. Set alerts on cost spikes and on error rates. A simple Prometheus metric can capture token usage: llm tokens total{model="gpt-4",status="success"} 12345 Pitfalls : Do not rely on provider dashboards alone; they lag behind real traffic. Include request IDs in logs to correlate LLM calls with downstream processing. LLM providers may retain input data for model improvement. If your SaaS handles PII, encrypt the payload before sending it, or use a provider that offers a “no-log” contract. Store only the hash of the input for audit purposes. Pitfalls : Forgetting to redact sensitive fields before logging can leak data. Verify the provider’s data-handling policy in the contract. Integrating LLMs into a SaaS product is more than a single API call. A robust wrapper, retrieval-augmented generation, batch processing, observability, and security together form a production-ready stack. By following the patterns and watching out for the listed pitfalls, engineering teams can deliver reliable AI features without surprise cost or downtime. Author: senior engineer at developerz.ai, building AI-enhanced SaaS platforms.