{"slug": "integrating-large-language-models-into-saas-products-practical-patterns-and", "title": "Integrating Large Language Models into SaaS Products: Practical Patterns and Pitfalls", "summary": "A developer outlined three production patterns for integrating large language models into SaaS products: a retry-and-backoff HTTP wrapper for provider APIs, retrieval-augmented generation pairing a FAISS vector store with OpenAI embeddings, and batched multi-prompt requests for bulk generation tasks. The writeup includes Ruby, Python and Go code samples along with pitfalls such as never exposing API keys client-side, keeping embeddings fresh, and staying within token budgets.", "body_md": "Large language models (LLMs) have moved from research labs to production services. SaaS companies can add powerful text generation, summarization, or code assistance features, but the path from prototype to reliable service is full of decisions. This article walks through three patterns that work well in production, shows code snippets, and highlights common pitfalls to avoid.\n\nThe simplest way to add LLM capability is to wrap the provider’s HTTP API in a small library. The wrapper should handle retries, exponential back-off, and a timeout that matches your SLA. Below is a minimal Ruby example using `net/http`.\n\n```\nrequire 'net/http'\nrequire 'json'\n\nmodule LlmClient\n  END_URL = URI('https://api.example.com/v1/completions')\n  API_KEY = ENV['LLM_API_KEY']\n\n  def self.complete(prompt, max_tokens: 200)\n    payload = {\n      model: 'gpt-4',\n      prompt: prompt,\n      max_tokens: max_tokens,\n      temperature: 0.2\n    }\n    request = Net::HTTP::Post.new(API_URL)\n    request['Authorization'] = \"Bearer #{API_KEY}\"\n    request['Content-Type'] = 'application/json'\n    request.body = payload.to_json\n    response = Net::HTTP.start(API_URL.host, API_URL.port, use_ssl: true) do |http|\n      http.request(request)\n    end\n    raise \"LLM error #{response.code}\" unless response.is_a?(Net::HTTPSuccess)\n    JSON.parse(response.body)['choices'][0]['text'].strip\n  end\nend\n```\n\n**Pitfalls**: Do not expose the raw API key to the client side. Cache the wrapper instance per request to avoid repeated socket creation. Log the request and response IDs for observability.\n\nFor domain-specific knowledge, raw LLM prompts are insufficient. RAG combines a vector store of embeddings with the model to retrieve relevant passages before generation. The flow is:\n\nBelow is a Python snippet using `faiss` for similarity search and `openai` for completion.\n\n``` python\nimport openai, faiss, numpy as np\n\n embed(text):\n    resp = openai.Embedding.create(input=text, model='text-embedding-ada-002')\n    return np.array(resp['data'][0]['embedding']).astype('float32')\n\n search(query, index, docs, k=3):\n    q_vec = embed(query)\n    distances, ids = index.search(np.expand_dims(q_vec, 0), k)\n    return \"\\n\".join([docs[i] for i in ids[0]])\n\ndef rag_completion(query, index, docs):\n    context = search(query, index, docs)\n    prompt = f\"Context:\\n{context}\\n\\nQuestion: {query}\\nAnswer:\"\n    resp = openai.Completion.create(model='gpt-4', prompt=prompt, max_tokens=250)\n    return resp['choices'][0]['text'].strip()\n```\n\n**Pitfalls**: Keep the vector store up-to-date; stale embeddings lead to irrelevant answers. Limit the amount of retrieved text to stay within the LLM token budget. Monitor latency; a FAISS search adds milliseconds but can be mitigated with caching.\n\nWhen the use case involves bulk generation - such as summarizing thousands of support tickets - sending one request per item is wasteful. Batch the inputs and call the LLM with a single request that contains an array of prompts. The provider often returns an array of completions in the same order.\n\n```\npackage llm\n\nimport (\n    \"bytes\"\n    \"encoding/json\"\n    \"net/http\"\n)\n\ntype BatchRequest struct {\n    Model string `json:\"model\"`\n    Prompts []string `json:\"prompt\"`\n    MaxTokens int `json:\"max_tokens\"`\n}\n\ntype BatchResponse struct {\n    Choices []struct{ Text string `json:\"text\"` } `json:\"choices\"`\n}\n\nfunc CompleteBatch(prompts []string) ([]string, error) {\n    reqBody:= BatchRequest{Model: \"gpt-4\", Prompts: prompts, MaxTokens: 150}\n    data, _:= json.Marshal(reqBody)\n    resp, err:= http.Post(\"https://api.example.com/v1/completions\", \"application/json\", bytes.NewReader(data))\n    if err!= nil { return nil, err }\n    defer resp.Body.Close()\n    var batchResp BatchResponse\n    json.NewDecoder(resp.Body).Decode(&batchResp)\n    results:= make([]string, len(batchResp.Choices))\n    for i, c:= range batchResp.Choices {\n        results[i] = c.Text\n    }\n    return results, nil\n}\n```\n\n**Pitfalls**: The provider may impose a maximum number of prompts per batch; split large jobs accordingly. Preserve the order of inputs to match outputs. Use a job queue (e.g., Sidekiq or RabbitMQ) to handle retries and back-pressure.\n\nProduction LLM services generate costs that can grow quickly. Instrument every call with tags for model, token count, and latency. Set alerts on cost spikes and on error rates. A simple Prometheus metric can capture token usage:\n\n```\nllm_tokens_total{model=\"gpt-4\",status=\"success\"} 12345\n```\n\n**Pitfalls**: Do not rely on provider dashboards alone; they lag behind real traffic. Include request IDs in logs to correlate LLM calls with downstream processing.\n\nLLM providers may retain input data for model improvement. If your SaaS handles PII, encrypt the payload before sending it, or use a provider that offers a “no-log” contract. Store only the hash of the input for audit purposes.\n\n**Pitfalls**: Forgetting to redact sensitive fields before logging can leak data. Verify the provider’s data-handling policy in the contract.\n\nIntegrating LLMs into a SaaS product is more than a single API call. A robust wrapper, retrieval-augmented generation, batch processing, observability, and security together form a production-ready stack. By following the patterns and watching out for the listed pitfalls, engineering teams can deliver reliable AI features without surprise cost or downtime.\n\n*Author: senior engineer at developerz.ai, building AI-enhanced SaaS platforms.*", "url": "https://wpnews.pro/news/integrating-large-language-models-into-saas-products-practical-patterns-and", "canonical_source": "https://dev.to/developerzai/integrating-large-language-models-into-saas-products-practical-patterns-and-pitfalls-3iah", "published_at": "2026-10-05 22:05:31+00:00", "updated_at": "2026-10-05 22:17:41.024895+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "developer-tools", "mlops", "ai-infrastructure"], "entities": ["OpenAI", "FAISS", "GPT-4", "text-embedding-ada-002", "Ruby", "Python", "Go"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/integrating-large-language-models-into-saas-products-practical-patterns-and", "markdown": "https://wpnews.pro/news/integrating-large-language-models-into-saas-products-practical-patterns-and.md", "text": "https://wpnews.pro/news/integrating-large-language-models-into-saas-products-practical-patterns-and.txt", "jsonld": "https://wpnews.pro/news/integrating-large-language-models-into-saas-products-practical-patterns-and.jsonld"}}