How to Build a RAG Pipeline with DynamoDB’s Native Vector Search, Explained Simply A developer demonstrated how to build a Retrieval-Augmented Generation (RAG) pipeline using DynamoDB's newly announced native vector indexes and plain Node.js, eliminating the need for a dedicated vector database. The approach stores 1536-dimensional Float32 embeddings alongside business data and uses a Global Secondary Index with cosine similarity for k-nearest-neighbor search, with Claude handling generation. The writeup argues that consolidating retrieval into DynamoDB cuts latency, operational overhead, and the risk of data mismatches between separate stores. Most engineers reach for a dedicated vector database the moment they need similarity search, adding latency and operational overhead. DynamoDB just announced native vector indexes, letting you store embeddings alongside your business data. In this post you’ll see how to turn that new feature into a fully‑functional Retrieval‑Augmented Generation RAG flow with Claude, using only DynamoDB and plain Node.js. In plain English – DynamoDB now does the math to find “similar looking” vectors, so you no longer need a separate service just to ask “which pieces of text look like this query?” A Retrieval‑Augmented Generation RAG system works in two steps: Historically the retrieve step required a vector store —a specialized database that can compare high‑dimensional vectors numeric representations of text quickly. Adding a second store meant: DynamoDB’s new vector index removes the separate store. Your embeddings live right next to the rest of the item, and DynamoDB’s built‑in KNN k‑nearest‑neighbors engine does the similarity math for you. Key takeaway – Cutting the stack in half also cuts the chances of data mismatches and reduces the time it takes to answer a user query. Before we can store or search embeddings, DynamoDB needs a table that knows where the vector lives and how to compare it. This is done by creating a GSI with a vector attribute type and declaring that the index should use cosine similarity . js // src/createTable.ts import { DynamoDBClient, CreateTableCommand, } from "@aws-sdk/client-dynamodb"; // The low‑level client is used for the CreateTable call because // the lib‑dynamodb wrapper does not expose table‑creation APIs. const client = new DynamoDBClient { region: "us-east-1" } ; async function createVectorTable { // Table name – keep it short and unique. const TableName = "RagDocs"; // Primary key partition key – we use a simple UUID string. const KeySchema = { AttributeName: "docId", KeyType: "HASH" } ; // The attribute that will hold the embedding as a binary blob. const AttributeDefinitions = { AttributeName: "docId", AttributeType: "S" }, // S = String { AttributeName: "embedding", AttributeType: "B" }, // B = Binary ; // Global Secondary Index GSI that enables vector search. const GlobalSecondaryIndexes = { IndexName: "EmbeddingKNN", // No projection needed for this demo; we pull the whole item later. Projection: { ProjectionType: "ALL" }, // The GSI's key schema uses a placeholder attribute; the // vector is supplied at query time, not stored in the index key. KeySchema: { AttributeName: "docId", KeyType: "HASH" } , // Vector configuration tells DynamoDB to treat embedding as a // 1536‑dimensional Float32 vector and to compare using cosine similarity. // The VectorConfig field is only available on the GSI definition. VectorConfig: { VectorDimensions: 1536, VectorDataType: "FLOAT32", // 32‑bit floating point numbers MetricType: "COSINE", // similarity metric VectorField: "embedding", // attribute that holds the binary vector }, // Provisioned read capacity – adjust for your traffic. ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 }, }, ; const command = new CreateTableCommand { TableName, KeySchema, AttributeDefinitions, BillingMode: "PROVISIONED", // explicit to control cost ProvisionedThroughput: { ReadCapacityUnits: 5, WriteCapacityUnits: 5 }, GlobalSecondaryIndexes, // Optional TTL time‑to‑live for automatic cleanup of old chunks. TimeToLiveSpecification: { AttributeName: "expiresAt", // epoch seconds Enabled: true, }, } ; try { const response = await client.send command ; console.log "Table created:", response.TableDescription?.TableName ; } catch err { console.error "Failed to create table:", err ; } } createVectorTable ; Explanation of tricky bits VectorConfig lives B because the service expects a base64‑encoded Float32Array . Tip – Even though the GSI has a dummy hash key docId , DynamoDB still requires it. Think of it as a “placeholder seat” that lets the engine focus on the vector field. Now that the table exists, we need to persist text chunks together with their embeddings. The embeddings are produced elsewhere e.g., OpenAI’s embedding endpoint and are 1536‑dimensional Float32 vectors. DynamoDB expects those vectors as binary data, so we must encode them as base64 before sending. When we later need the most relevant chunks for a user question, we will: EmbeddingKNN GSI, asking for the top 3 nearest neighbors. js // src/putDocument.ts import { DynamoDBDocumentClient, PutCommand, } from "@aws-sdk/lib-dynamodb"; import { DynamoDBClient } from "@aws-sdk/client-dynamodb"; // Low‑level client needed only for the Document wrapper. const ddbClient = new DynamoDBClient { region: "us-east-1" } ; const ddbDoc = DynamoDBDocumentClient.from ddbClient ; // Helper: turn a Float32Array into a base64 string. function encodeEmbedding vec: Float32Array : string { // Convert the raw bytes to a Uint8Array, then to a base64 string. const bytes = new Uint8Array vec.buffer ; return Buffer.from bytes .toString "base64" ; } // Example: a single text chunk and its embedding. async function putChunk docId: string, chunkText: string, embedding: Float32Array, ttlSeconds: number // optional expiration { const command = new PutCommand { TableName: "RagDocs", Item: { docId, // partition key chunk: chunkText, embedding: encodeEmbedding embedding , // stored as Binary base64 // DynamoDB treats a base64 string as Binary automatically. // Adding a TTL helps keep the table size in check. expiresAt: Math.floor Date.now / 1000 + ttlSeconds, }, } ; try { await ddbDoc.send command ; console.log Chunk ${docId} stored ; } catch err { console.error "Write error:", err ; } } / ------------------------------------------------------------- Imagine we already fetched an embedding from OpenAI: const embedding = await getOpenAIEmbedding chunkText ; ------------------------------------------------------------- / js // src/queryKnn.ts import { DynamoDBDocumentClient, QueryCommand, } from "@aws-sdk/lib-dynamodb"; import { DynamoDBClient } from "@aws-sdk/client-dynamodb"; const ddbClient = new DynamoDBClient { region: "us-east-1" } ; const ddbDoc = DynamoDBDocumentClient.from ddbClient ; // Same encoder used for writes; DynamoDB expects the query vector // in the same binary format. function encodeEmbedding vec: Float32Array : string { const bytes = new Uint8Array vec.buffer ; return Buffer.from bytes .toString "base64" ; } / Returns the k most similar chunks for a given query embedding. / async function knnSearch queryEmbedding: Float32Array, k: number = 3 { const command = new QueryCommand { TableName: "RagDocs", IndexName: "EmbeddingKNN", // the vector‑enabled GSI // The special VectorSearch field tells DynamoDB to perform a // KNN lookup using the supplied binary vector. // KNN is the number of neighbors to return. // Metric defaults to the one defined in the GSI COSINE . VectorSearch: { KNN: k, VectorField: "embedding", QueryVector: encodeEmbedding queryEmbedding , }, // ConsistentRead is not allowed on GSI; we accept eventual consistency. // This matches the reality of DynamoDB’s GSI behavior. } ; try { const result = await ddbDoc.send command ; // Each item contains the original chunk text plus the embedding. return result.Items?.map item = { docId: item.docId, chunk: item.chunk, } ?? ; } catch err { console.error "KNN query error:", err ; return ; } } Gotchas highlighted ValidationException: Attribute type mismatch . In plain English – Think of the embedding as a secret code. DynamoDB only understands the code when it’s wrapped in a base64 envelope; missing the envelope makes the service shout “I don’t know this type ”. The RAG pattern finishes by feeding the retrieved chunks to an LLM so it can craft a response. Claude’s HTTP API is straightforward: send a JSON payload with a messages array and receive a completion field. We’ll use Node 22’s built‑in fetch no extra packages needed . // src/callClaude.ts / Sends a user question together with retrieved context chunks to Claude. Returns Claude’s answer as plain text. / async function askClaude question: string, contextChunks: { chunk: string } , apiKey: string : Promise