{"slug": "how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply", "title": "How to Build a RAG Pipeline with DynamoDB’s Native Vector Search, Explained Simply", "summary": "A developer demonstrated how to build a Retrieval-Augmented Generation (RAG) pipeline using DynamoDB's newly announced native vector indexes and plain Node.js, eliminating the need for a dedicated vector database. The approach stores 1536-dimensional Float32 embeddings alongside business data and uses a Global Secondary Index with cosine similarity for k-nearest-neighbor search, with Claude handling generation. The writeup argues that consolidating retrieval into DynamoDB cuts latency, operational overhead, and the risk of data mismatches between separate stores.", "body_md": "Most engineers reach for a dedicated vector database the moment they need similarity search, adding latency and operational overhead. DynamoDB just announced native vector indexes, letting you store embeddings alongside your business data. In this post you’ll see how to turn that new feature into a fully‑functional Retrieval‑Augmented Generation (RAG) flow with Claude, using only DynamoDB and plain Node.js.\n\n**In plain English** – DynamoDB now does the math to find “similar looking” vectors, so you no longer need a separate service just to ask “which pieces of text look like this query?”  \n\nA Retrieval‑Augmented Generation (RAG) system works in two steps:\n\nHistorically the *retrieve* step required a **vector store**—a specialized database that can compare high‑dimensional vectors (numeric representations of text) quickly. Adding a second store meant:\n\nDynamoDB’s new **vector index** removes the separate store. Your embeddings live right next to the rest of the item, and DynamoDB’s built‑in KNN (k‑nearest‑neighbors) engine does the similarity math for you.  \n\n**Key takeaway** – Cutting the stack in half also cuts the chances of data mismatches and reduces the time it takes to answer a user query.  \n\nBefore we can store or search embeddings, DynamoDB needs a table that knows **where** the vector lives and **how** to compare it. This is done by creating a **GSI** with a *vector* attribute type and declaring that the index should use **cosine similarity**.  \n\n``` js\n// src/createTable.ts\nimport {\n  DynamoDBClient,\n  CreateTableCommand,\n} from \"@aws-sdk/client-dynamodb\";\n\n// The low‑level client is used for the CreateTable call because\n// the lib‑dynamodb wrapper does not expose table‑creation APIs.\nconst client = new DynamoDBClient({ region: \"us-east-1\" });\n\nasync function createVectorTable() {\n  // Table name – keep it short and unique.\n  const TableName = \"RagDocs\";\n\n  // Primary key (partition key) – we use a simple UUID string.\n  const KeySchema = [{ AttributeName: \"docId\", KeyType: \"HASH\" }];\n\n  // The attribute that will hold the embedding as a binary blob.\n  const AttributeDefinitions = [\n    { AttributeName: \"docId\", AttributeType: \"S\" }, // S = String\n    { AttributeName: \"embedding\", AttributeType: \"B\" }, // B = Binary\n  ];\n\n  // Global Secondary Index (GSI) that enables vector search.\n  const GlobalSecondaryIndexes = [\n    {\n      IndexName: \"EmbeddingKNN\",\n      // No projection needed for this demo; we pull the whole item later.\n      Projection: { ProjectionType: \"ALL\" },\n      // The GSI's key schema uses a placeholder attribute; the\n      // vector is supplied at query time, not stored in the index key.\n      KeySchema: [{ AttributeName: \"docId\", KeyType: \"HASH\" }],\n      // Vector configuration tells DynamoDB to treat `embedding` as a\n      // 1536‑dimensional Float32 vector and to compare using cosine similarity.\n      // The `VectorConfig` field is only available on the GSI definition.\n      VectorConfig: {\n        VectorDimensions: 1536,\n        VectorDataType: \"FLOAT32\", // 32‑bit floating point numbers\n        MetricType: \"COSINE\", // similarity metric\n        VectorField: \"embedding\", // attribute that holds the binary vector\n      },\n      // Provisioned read capacity – adjust for your traffic.\n      ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 },\n    },\n  ];\n\n  const command = new CreateTableCommand({\n    TableName,\n    KeySchema,\n    AttributeDefinitions,\n    BillingMode: \"PROVISIONED\", // explicit to control cost\n    ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 },\n    GlobalSecondaryIndexes,\n    // Optional TTL (time‑to‑live) for automatic cleanup of old chunks.\n    TimeToLiveSpecification: {\n      AttributeName: \"expiresAt\", // epoch seconds\n      Enabled: true,\n    },\n  });\n\n  try {\n    const response = await client.send(command);\n    console.log(\"Table created:\", response.TableDescription?.TableName);\n  } catch (err) {\n    console.error(\"Failed to create table:\", err);\n  }\n}\n\ncreateVectorTable();\n```\n\n**Explanation of tricky bits**\n\n`VectorConfig` lives `B`) because the service expects a base64‑encoded `Float32Array`.\n**Tip** – Even though the GSI has a dummy hash key (`docId`), DynamoDB still requires it. Think of it as a “placeholder seat” that lets the engine focus on the vector field.  \n\nNow that the table exists, we need to **persist** text chunks together with their embeddings. The embeddings are produced elsewhere (e.g., OpenAI’s embedding endpoint) and are 1536‑dimensional `Float32` vectors. DynamoDB expects those vectors as binary data, so we must **encode** them as base64 before sending.  \n\nWhen we later need the most relevant chunks for a user question, we will:\n\n`EmbeddingKNN` GSI, asking for the top 3 nearest neighbors.\n\n``` js\n// src/putDocument.ts\nimport {\n  DynamoDBDocumentClient,\n  PutCommand,\n} from \"@aws-sdk/lib-dynamodb\";\nimport { DynamoDBClient } from \"@aws-sdk/client-dynamodb\";\n\n// Low‑level client needed only for the Document wrapper.\nconst ddbClient = new DynamoDBClient({ region: \"us-east-1\" });\nconst ddbDoc = DynamoDBDocumentClient.from(ddbClient);\n\n// Helper: turn a Float32Array into a base64 string.\nfunction encodeEmbedding(vec: Float32Array): string {\n  // Convert the raw bytes to a Uint8Array, then to a base64 string.\n  const bytes = new Uint8Array(vec.buffer);\n  return Buffer.from(bytes).toString(\"base64\");\n}\n\n// Example: a single text chunk and its embedding.\nasync function putChunk(\n  docId: string,\n  chunkText: string,\n  embedding: Float32Array,\n  ttlSeconds: number // optional expiration\n) {\n  const command = new PutCommand({\n    TableName: \"RagDocs\",\n    Item: {\n      docId, // partition key\n      chunk: chunkText,\n      embedding: encodeEmbedding(embedding), // stored as Binary (base64)\n      // DynamoDB treats a base64 string as Binary automatically.\n      // Adding a TTL helps keep the table size in check.\n      expiresAt: Math.floor(Date.now() / 1000) + ttlSeconds,\n    },\n  });\n\n  try {\n    await ddbDoc.send(command);\n    console.log(`Chunk ${docId} stored`);\n  } catch (err) {\n    console.error(\"Write error:\", err);\n  }\n}\n\n/* -------------------------------------------------------------\n   Imagine we already fetched an embedding from OpenAI:\n   const embedding = await getOpenAIEmbedding(chunkText);\n   ------------------------------------------------------------- */\njs\n// src/queryKnn.ts\nimport {\n  DynamoDBDocumentClient,\n  QueryCommand,\n} from \"@aws-sdk/lib-dynamodb\";\nimport { DynamoDBClient } from \"@aws-sdk/client-dynamodb\";\n\nconst ddbClient = new DynamoDBClient({ region: \"us-east-1\" });\nconst ddbDoc = DynamoDBDocumentClient.from(ddbClient);\n\n// Same encoder used for writes; DynamoDB expects the query vector\n// in the same binary format.\nfunction encodeEmbedding(vec: Float32Array): string {\n  const bytes = new Uint8Array(vec.buffer);\n  return Buffer.from(bytes).toString(\"base64\");\n}\n\n/**\n * Returns the `k` most similar chunks for a given query embedding.\n */\nasync function knnSearch(\n  queryEmbedding: Float32Array,\n  k: number = 3\n) {\n  const command = new QueryCommand({\n    TableName: \"RagDocs\",\n    IndexName: \"EmbeddingKNN\", // the vector‑enabled GSI\n    // The special `VectorSearch` field tells DynamoDB to perform a\n    // KNN lookup using the supplied binary vector.\n    // `KNN` is the number of neighbors to return.\n    // `Metric` defaults to the one defined in the GSI (COSINE).\n    VectorSearch: {\n      KNN: k,\n      VectorField: \"embedding\",\n      QueryVector: encodeEmbedding(queryEmbedding),\n    },\n    // ConsistentRead is not allowed on GSI; we accept eventual consistency.\n    // This matches the reality of DynamoDB’s GSI behavior.\n  });\n\n  try {\n    const result = await ddbDoc.send(command);\n    // Each item contains the original chunk text plus the embedding.\n    return result.Items?.map((item) => ({\n      docId: item.docId,\n      chunk: item.chunk,\n    })) ?? [];\n  } catch (err) {\n    console.error(\"KNN query error:\", err);\n    return [];\n  }\n}\n```\n\n**Gotchas highlighted**\n\n`ValidationException: Attribute type mismatch`.\n**In plain English** – Think of the embedding as a secret code. DynamoDB only understands the code when it’s wrapped in a base64 envelope; missing the envelope makes the service shout “I don’t know this type!”.  \n\nThe RAG pattern finishes by feeding the retrieved chunks to an LLM so it can craft a response. Claude’s HTTP API is straightforward: send a JSON payload with a `messages` array and receive a `completion` field. We’ll use Node 22’s built‑in **fetch** (no extra packages needed).  \n\n```\n// src/callClaude.ts\n/**\n * Sends a user question together with retrieved context chunks to Claude.\n * Returns Claude’s answer as plain text.\n */\nasync function askClaude(\n  question: string,\n  contextChunks: { chunk: string }[],\n  apiKey: string\n): Promise<string> {\n  // Build a single system prompt that concatenates the chunks.\n  const systemPrompt = `You are an assistant that answers questions using only the following information:\\n\\n${contextChunks\n    .map((c) => `- ${c.chunk}`)\n    .join(\"\\n\")}\\n\\nIf the answer cannot be derived from this data, say \"I don't know.\"`;\n\n  const payload = {\n    model: \"claude-3-5-sonnet-20241007\", // example model name\n    messages: [\n      { role: \"system\", content: systemPrompt },\n      { role: \"user\", content: question },\n    ],\n    max_tokens: 500,\n  };\n\n  const response = await fetch(\n    \"https://api.anthropic.com/v1/messages\",\n    {\n      method: \"POST\",\n      headers: {\n        \"Content-Type\": \"application/json\",\n        // Anthropic expects the key in the `x-api-key` header.\n        \"x-api-key\": apiKey,\n      },\n      body: JSON.stringify(payload),\n    }\n  );\n\n  if (!response.ok) {\n    const errorBody = await response.text();\n    throw new Error(`Claude API error ${response.status}: ${errorBody}`);\n  }\n\n  const data = await response.json();\n  // The completion text lives in `content[0].text` for most responses.\n  return data.content?.[0]?.text ?? \"No response\";\n}\n```\n\n**Tip** – Keep the system prompt short; Claude’s token limit includes both the prompt and the answer, so overly long context can truncate the model’s output.  \n\nHaving separate snippets is useful for learning, but production code needs a **single orchestrator** that:\n\nBelow is a compact Node.js script that ties everything together. It uses the same SDKs and fetch we already introduced, plus the OpenAI embedding endpoint (you can swap it for any model that returns a 1536‑dimensional vector).\n\n``` js\n// src/ragService.ts\nimport { DynamoDBDocumentClient } from \"@aws-sdk/lib-dynamodb\";\nimport { DynamoDBClient } from \"@aws-sdk/client-dynamodb\";\nimport { knnSearch } from \"./queryKnn\";\nimport { askClaude } from \"./callClaude\";\nimport fetch from \"node-fetch\"; // Node 22 includes global fetch, but keep for type safety\n\n// Re‑use the document client we created earlier.\nconst ddbClient = new DynamoDBClient({ region: \"us-east-1\" });\nconst ddbDoc = DynamoDBDocumentClient.from(ddbClient);\n\n// Replace with your own keys.\nconst OPENAI_API_KEY = process.env.OPENAI_API_KEY!;\nconst CLAUDE_API_KEY = process.env.CLAUDE_API_KEY!;\n\n/**\n * Calls OpenAI's embedding endpoint to turn text into a Float32Array.\n */\nasync function embed(text: string): Promise<Float32Array> {\n  const response = await fetch(\n    \"https://api.openai.com/v1/embeddings\",\n    {\n      method: \"POST\",\n      headers: {\n        \"Content-Type\": \"application/json\",\n        Authorization: `Bearer ${OPENAI_API_KEY}`,\n      },\n      body: JSON.stringify({\n        model: \"text-embedding-3-large\", // returns 1536‑dim vectors\n        input: text,\n      }),\n    }\n  );\n\n  if (!response.ok) {\n    const err = await response.text();\n    throw new Error(`OpenAI embed error ${response.status}: ${err}`);\n  }\n\n  const data = await response.json();\n  // OpenAI returns an array of numbers; convert to Float32Array.\n  const numbers: number[] = data.data[0].embedding;\n  return new Float32Array(numbers);\n}\n\n/**\n * Main entry point – receives a user question and returns Claude's answer.\n */\nexport async function answerQuestion(question: string): Promise<string> {\n  // 1️⃣ Turn the question into a vector.\n  const queryVec = await embed(question);\n\n  // 2️⃣ Find the 3 most similar stored chunks.\n  const matches = await knnSearch(queryVec, 3);\n\n  if (matches.length === 0) {\n    return \"I couldn't find any relevant information.\";\n  }\n\n  // 3️⃣ Pass the question + context to Claude.\n  const answer = await askClaude(question, matches, CLAUDE_API_KEY);\n  return answer;\n}\n\n/* -------------------------------------------------------------\n   Example usage as a simple HTTP server (Node's built‑in http).\n   ------------------------------------------------------------- */\nimport { createServer } from \"http\";\n\nconst server = createServer(async (req, res) => {\n  if (req.method !== \"POST\" || req.url !== \"/ask\") {\n    res.writeHead(404);\n    res.end(\"Not found\");\n    return;\n  }\n\n  try {\n    const body = await new Promise<string>((resolve, reject) => {\n      let data = \"\";\n      req.on(\"data\", (chunk) => (data += chunk));\n      req.on(\"end\", () => resolve(data));\n      req.on(\"error\", reject);\n    });\n    const { question } = JSON.parse(body);\n    const answer = await answerQuestion(question);\n    res.writeHead(200, { \"Content-Type\": \"application/json\" });\n    res.end(JSON.stringify({ answer }));\n  } catch (e) {\n    console.error(e);\n    res.writeHead(500);\n    res.end(\"Server error\");\n  }\n});\n\nserver.listen(3000, () => console.log(\"RAG service listening on :3000\"));\n```\n\n**What this script does**\n\n**Operational gotchas**\n\n`docId` range, DynamoDB can throttle. Mitigate by adding a random prefix to the partition key when you write chunks (sharding).\n`TransactionCanceledException`.\n**In plain English** – The whole pipeline is now a single DynamoDB table, one HTTP call to OpenAI, one KNN query, and one call to Claude. No extra services, no extra latency.  \n\n- DynamoDB’s native vector index lets you store embeddings next to your business data, removing the need for a separate vector database.\n- Create the table with a\n**vector‑enabled GSI**; remember to declare the embedding attribute as **Binary** and to base64‑encode the `Float32Array`.\n- Write and read embeddings using the\n`@aws-sdk/lib-dynamodb` package; the same client handles regular attributes and vectors.\n- A KNN query returns the most similar chunks; keep in mind the eventual‑consistency behavior of GSIs.\n- Use Node 22’s built‑in\n**fetch** to call Claude (or any LLM) after you have the context, building a short system prompt that forces the model to stay on‑topic.\n- Operational realities—hot partitions, TTL lag, and capacity pricing—still apply, so plan sharding and monitoring from day 1.\n\nWith these pieces in place, you have a lean, production‑ready Retrieval‑Augmented Generation service that lives entirely inside DynamoDB and plain Node.js. Happy building!\n\n**Transparency notice**\n\nThis article was written with the help of an AI system — [Groq](https://groq.com) (GPT OSS 120B).\n\n**Published:** 2026-10-07 · **Primary focus:** DynamoDB\n\nAll code blocks are intended to be correct and runnable, but please verify them\n\nagainst the official docs for the tools mentioned before using in production.\n\n*Find an error? Drop a comment — corrections are always welcome.*", "url": "https://wpnews.pro/news/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply", "canonical_source": "https://dev.to/dineshgowtham/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply-196k", "published_at": "2026-10-07 09:05:02+00:00", "updated_at": "2026-10-07 09:17:06.910219+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["DynamoDB", "Amazon Web Services", "Claude", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply", "markdown": "https://wpnews.pro/news/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply.md", "text": "https://wpnews.pro/news/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply.txt", "jsonld": "https://wpnews.pro/news/how-to-build-a-rag-pipeline-with-dynamodbs-native-vector-search-explained-simply.jsonld"}}