{"slug": "vector-search-embeddings-in-java-building-semantic-search-engines", "title": "Vector Search & Embeddings in Java: Building Semantic Search Engines", "summary": "A developer has published a guide on building semantic search engines in Java using vector embeddings, walking through the architecture of an embedding model, a vector search engine, and similarity ranking. The writeup covers cosine similarity, Euclidean and Manhattan distance metrics, and demonstrates wiring OpenAI's text-embedding-ada-002 model to Pinecone via the openai-gpt3-java and pinecone-client libraries.", "body_md": "Traditional keyword-based search is dead. When users search for \"best restaurants near me,\" they don't expect results matching those exact words—they expect restaurants that *mean* the right thing.\n\nThis is where **vector search and embeddings** enter the picture.\n\nVector search is the technology powering modern AI applications: ChatGPT's retrieval-augmented generation (RAG), Netflix's recommendation engine, Spotify's \"Discover Weekly,\" and enterprise semantic search platforms. Yet many Java developers still think of search as Elasticsearch queries with exact terms.\n\nThis gap is costing you:\n\nIn this guide, you'll learn:\n\nBy the end, you'll understand why vector search is essential for modern applications and how to build it with Java.\n\nAn **embedding** is a numerical representation of text, images, or other data. Instead of storing \"restaurant recommendations,\" you store a vector of numbers: `[0.25, -0.15, 0.82, ..., 0.41]`.\n\nThese aren't random numbers. They're learned through **neural networks** trained on massive datasets. Similar concepts produce similar vectors. This property is the entire foundation of vector search.\n\n**Example:**\n\n`[0.12, 0.88, -0.31, 0.45, ...]`\n`[0.14, 0.87, -0.29, 0.46, ...]`\nThese vectors are *close* in vector space. Measuring that distance (using cosine similarity, Euclidean distance, etc.) gives you a relevance score.\n\nIn production, you don't hand-craft embeddings. You use pre-trained embedding models:\n\nThese models map text → vector in a way that preserves semantic meaning.\n\nAll vectors live in an **N-dimensional space**. When you have 1,536-dimensional vectors (from OpenAI), you're working in 1,536-dimensional space.\n\n**Similarity metrics:**\n\n**Cosine Similarity** (most common)\n\n`(A · B) / (||A|| × ||B||)`\n**Euclidean Distance** (L2)\n\n`√(Σ(ai - bi)²)`\n**Manhattan Distance** (L1)\n\nFor text search, **cosine similarity** is almost always the right choice.\n\nA semantic search system has these components:\n\n```\n┌─────────────────────────────────────────────┐\n│  User Query                                  │\n└─────────────┬───────────────────────────────┘\n              │\n┌─────────────▼───────────────────────────────┐\n│  1. Embedding Model (Convert text → vector) │\n│     (OpenAI API / Local Sentence Transformer)│\n└─────────────┬───────────────────────────────┘\n              │\n┌─────────────▼───────────────────────────────┐\n│  2. Vector Search Engine                     │\n│     (Pinecone / PostgreSQL pgvector)        │\n└─────────────┬───────────────────────────────┘\n              │\n┌─────────────▼───────────────────────────────┐\n│  3. Similarity Ranking                       │\n│     Return top-K most relevant results      │\n└─────────────┬───────────────────────────────┘\n              │\n┌─────────────▼───────────────────────────────┐\n│  Results with scores (0.0 - 1.0)            │\n└─────────────────────────────────────────────┘\n```\n\n**Step 1: Add Dependencies**\n\n``` php\n<!-- pom.xml -->\n<dependency>\n    <groupId>com.theokanning.openai-gpt3-java</groupId>\n    <artifactId>api</artifactId>\n    <version>0.18.1</version>\n</dependency>\n\n<dependency>\n    <groupId>io.pinecone</groupId>\n    <artifactId>pinecone-client</artifactId>\n    <version>0.1.0</version>\n</dependency>\n```\n\n**Step 2: Create Embedding Service**\n\n``` python\nimport com.theokanning.openai.embedding.Embedding;\nimport com.theokanning.openai.embedding.EmbeddingRequest;\nimport com.theokanning.openai.embedding.EmbeddingResult;\nimport com.theokanning.openai.service.OpenAiService;\n\npublic class EmbeddingService {\n    private final OpenAiService openAiService;\n    private final String modelId = \"text-embedding-ada-002\";\n\n    public EmbeddingService(String apiKey) {\n        this.openAiService = new OpenAiService(apiKey);\n    }\n\n    public List<Double> embedText(String text) {\n        EmbeddingRequest request = EmbeddingRequest.builder()\n            .model(modelId)\n            .input(Collections.singletonList(text))\n            .build();\n\n        EmbeddingResult result = openAiService.createEmbeddings(request);\n        return result.getData().get(0).getEmbedding();\n    }\n\n    public List<List<Double>> embedTexts(List<String> texts) {\n        EmbeddingRequest request = EmbeddingRequest.builder()\n            .model(modelId)\n            .input(texts)\n            .build();\n\n        EmbeddingResult result = openAiService.createEmbeddings(request);\n        return result.getData().stream()\n            .sorted(Comparator.comparingInt(Embedding::getIndex))\n            .map(Embedding::getEmbedding)\n            .collect(Collectors.toList());\n    }\n}\n```\n\n**Step 3: Pinecone Vector Search**\n\n``` python\nimport io.pinecone.clients.Pinecone;\nimport io.pinecone.clients.Index;\n\npublic class PineconeVectorStore {\n    private final Index index;\n    private final EmbeddingService embeddingService;\n\n    public PineconeVectorStore(String apiKey, String projectName, \n                               String indexName, String embeddingApiKey) {\n        Pinecone client = new Pinecone.Builder()\n            .withApiKey(apiKey)\n            .build();\n\n        this.index = client.getIndex(projectName, indexName);\n        this.embeddingService = new EmbeddingService(embeddingApiKey);\n    }\n\n    // Index documents with embeddings\n    public void indexDocument(String docId, String content, \n                             Map<String, String> metadata) {\n        List<Double> embedding = embeddingService.embedText(content);\n\n        index.upsert(\n            docId,\n            embedding,\n            metadata\n        );\n    }\n\n    // Search: returns top K similar documents\n    public List<SearchResult> search(String query, int topK) {\n        List<Double> queryEmbedding = embeddingService.embedText(query);\n\n        var results = index.query(queryEmbedding)\n            .withTopK(topK)\n            .withIncludeMetadata(true)\n            .execute();\n\n        return results.getMatches().stream()\n            .map(match -> new SearchResult(\n                match.getId(),\n                match.getScore(),\n                match.getMetadata()\n            ))\n            .collect(Collectors.toList());\n    }\n}\n\nrecord SearchResult(String id, Double score, Map<String, String> metadata) {}\n```\n\n**Step 4: End-to-End Usage**\n\n```\npublic class SemanticSearchApp {\n    public static void main(String[] args) {\n        PineconeVectorStore store = new PineconeVectorStore(\n            System.getenv(\"PINECONE_API_KEY\"),\n            \"my-project\",\n            \"restaurant-index\",\n            System.getenv(\"OPENAI_API_KEY\")\n        );\n\n        // Index documents\n        store.indexDocument(\"rest-1\", \"Best pizza in town, authentic Italian, family-owned\", \n            Map.of(\"name\", \"Pasta Paradise\", \"type\", \"pizza\"));\n\n        store.indexDocument(\"rest-2\", \"Fast casual ramen, Tokyo-style noodles, great broth\",\n            Map.of(\"name\", \"Noodle House\", \"type\", \"ramen\"));\n\n        // Search with semantic understanding\n        var results = store.search(\"where to find good Italian food\", 3);\n\n        results.forEach(r -> \n            System.out.printf(\"ID: %s, Score: %.4f, Name: %s%n\",\n                r.id(), r.score(), r.metadata().get(\"name\"))\n        );\n    }\n}\n```\n\nFor privacy-sensitive applications, you might want embeddings to stay within your infrastructure.\n\n**Step 1: Setup PostgreSQL pgvector**\n\n```\n-- Enable pgvector extension\nCREATE EXTENSION IF NOT EXISTS vector;\n\n-- Create table\nCREATE TABLE documents (\n    id SERIAL PRIMARY KEY,\n    content TEXT NOT NULL,\n    embedding vector(1536),\n    metadata JSONB,\n    created_at TIMESTAMP DEFAULT NOW()\n);\n\n-- Create HNSW index for fast similarity search\nCREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);\n```\n\n**Step 2: Java Implementation with Sentence Transformers**\n\n``` python\nimport org.springframework.ai.document.Document;\nimport org.springframework.ai.embedding.Embedding;\nimport org.springframework.ai.embedding.EmbeddingModel;\nimport org.springframework.jdbc.core.JdbcTemplate;\n\n@Component\npublic class PostgresVectorStore {\n    private final JdbcTemplate jdbcTemplate;\n    private final EmbeddingModel embeddingModel;\n\n    @Autowired\n    public PostgresVectorStore(JdbcTemplate jdbcTemplate, \n                               EmbeddingModel embeddingModel) {\n        this.jdbcTemplate = jdbcTemplate;\n        this.embeddingModel = embeddingModel;\n    }\n\n    public void indexDocument(String content, String metadata) {\n        Embedding embedding = embeddingModel.embed(content);\n\n        String sql = \"INSERT INTO documents (content, embedding, metadata) \" +\n                     \"VALUES (?, ?::vector, ?::jsonb)\";\n\n        jdbcTemplate.update(sql, content, vectorToString(embedding), metadata);\n    }\n\n    public List<DocumentResult> similaritySearch(String query, int limit) {\n        Embedding queryEmbedding = embeddingModel.embed(query);\n\n        String sql = \"SELECT id, content, metadata, \" +\n                     \"1 - (embedding <=> ?::vector) as similarity \" +\n                     \"FROM documents \" +\n                     \"ORDER BY embedding <=> ?::vector \" +\n                     \"LIMIT ?\";\n\n        return jdbcTemplate.query(sql, new Object[]{\n            vectorToString(queryEmbedding),\n            vectorToString(queryEmbedding),\n            limit\n        }, (rs, rowNum) -> new DocumentResult(\n            rs.getInt(\"id\"),\n            rs.getString(\"content\"),\n            rs.getDouble(\"similarity\"),\n            rs.getString(\"metadata\")\n        ));\n    }\n\n    private String vectorToString(Embedding embedding) {\n        return \"[\" + embedding.getOutput().stream()\n            .map(String::valueOf)\n            .collect(Collectors.joining(\",\")) + \"]\";\n    }\n}\n\nrecord DocumentResult(int id, String content, double score, String metadata) {}\n```\n\nTraditional: Search for \"comfortable shoes\" → matches only products with those exact words\n\nSemantic: Matches \"ergonomic footwear,\" \"supportive sneakers,\" \"cushioned athletic shoes\"\n\n```\npublic class ProductSearchService {\n    private final VectorStore vectorStore;\n\n    public List<Product> findSimilarProducts(String query) {\n        return vectorStore.search(query, 10).stream()\n            .map(this::toProduct)\n            .collect(Collectors.toList());\n    }\n}\n```\n\nRetrieval-Augmented Generation: Embed your knowledge base, find relevant docs, feed them to LLM\n\n```\n@RestController\npublic class SupportChatbot {\n    private final VectorStore knowledgeBase;\n    private final OpenAiService openAiService;\n\n    @PostMapping(\"/ask\")\n    public ResponseEntity<String> ask(@RequestBody String question) {\n        // Step 1: Find relevant docs using vector search\n        var relevantDocs = knowledgeBase.search(question, 3);\n\n        // Step 2: Build context from retrieved docs\n        String context = relevantDocs.stream()\n            .map(SearchResult::content)\n            .collect(Collectors.joining(\"\\n\\n\"));\n\n        // Step 3: Ask LLM with context\n        String prompt = String.format(\n            \"Based on this knowledge base:\\n%s\\n\\nAnswer: %s\",\n            context, question\n        );\n\n        ChatCompletionRequest request = ChatCompletionRequest.builder()\n            .model(\"gpt-4\")\n            .messages(List.of(new ChatMessage(ChatMessageRole.USER.value(), prompt)))\n            .build();\n\n        String answer = openAiService.createChatCompletion(request)\n            .getChoices().get(0).getMessage().getContent();\n\n        return ResponseEntity.ok(answer);\n    }\n}\n```\n\nFind near-duplicate documents or suspicious fraud patterns\n\n```\npublic class DuplicateDetector {\n    private final VectorStore vectorStore;\n\n    public boolean isProbablyDuplicate(String document, double threshold) {\n        var similarDocs = vectorStore.search(document, 1);\n\n        return !similarDocs.isEmpty() && \n               similarDocs.get(0).score() > threshold;\n    }\n}\n```\n\nDon't embed one document at a time. Batch them.\n\n```\npublic void indexManyDocuments(List<Document> documents) {\n    // ❌ Slow: N API calls\n    // documents.forEach(doc -> index(doc));\n\n    // ✅ Fast: 1 API call per batch\n    Iterables.partition(documents, 100).forEach(batch -> {\n        List<String> texts = batch.stream()\n            .map(Document::getContent)\n            .collect(Collectors.toList());\n\n        List<List<Double>> embeddings = embeddingService.embedTexts(texts);\n\n        for (int i = 0; i < batch.size(); i++) {\n            vectorStore.index(batch.get(i).getId(), embeddings.get(i));\n        }\n    });\n}\n```\n\nStore computed embeddings to avoid redundant API calls\n\n```\n@Component\npublic class CachedEmbeddingService {\n    private final EmbeddingService service;\n    private final Map<String, List<Double>> cache = new ConcurrentHashMap<>();\n\n    public List<Double> embed(String text) {\n        return cache.computeIfAbsent(text, key -> service.embedText(key));\n    }\n}\n```\n\nNot all 1,536 dimensions are necessary for your use case. PCA can reduce them:\n\n```\npublic List<Double> reduceDimensions(List<Double> embedding, int targetDim) {\n    // Use Apache Commons Math or similar\n    // This trades accuracy for speed/storage\n    return PCA.reduce(embedding, targetDim);\n}\n```\n\nIf this is too slow, cache frequent queries.\n\nFor high-volume applications, self-hosting saves money.\n\nWhen document content changes, re-embed and update:\n\n```\npublic void updateDocument(String docId, String newContent) {\n    List<Double> newEmbedding = embeddingService.embedText(newContent);\n    vectorStore.update(docId, newEmbedding, newContent);\n}\n```\n\nTrack embedding quality:\n\n```\n@Component\npublic class EmbeddingQualityMonitor {\n    private final MeterRegistry meterRegistry;\n\n    public void recordSimilarityScore(double score) {\n        Timer.builder(\"search.similarity.score\")\n            .publishPercentiles(0.5, 0.95, 0.99)\n            .register(meterRegistry)\n            .record(Duration.ofMillis((long) (score * 1000)));\n    }\n}\n```\n\nVector search and embeddings are no longer bleeding-edge. They're essential for:\n\n**Key Takeaways:**\n\nYour next semantic search implementation is just these components away. Build it today.", "url": "https://wpnews.pro/news/vector-search-embeddings-in-java-building-semantic-search-engines", "canonical_source": "https://dev.to/said_olano/vector-search-embeddings-in-java-building-semantic-search-engines-279c", "published_at": "2026-10-02 01:41:21+00:00", "updated_at": "2026-10-02 01:44:22.735356+00:00", "lang": "en", "topics": ["natural-language-processing", "ai-tools", "developer-tools", "large-language-models"], "entities": ["Java", "OpenAI", "Pinecone", "text-embedding-ada-002", "openai-gpt3-java", "pinecone-client", "PostgreSQL pgvector", "Elasticsearch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vector-search-embeddings-in-java-building-semantic-search-engines", "markdown": "https://wpnews.pro/news/vector-search-embeddings-in-java-building-semantic-search-engines.md", "text": "https://wpnews.pro/news/vector-search-embeddings-in-java-building-semantic-search-engines.txt", "jsonld": "https://wpnews.pro/news/vector-search-embeddings-in-java-building-semantic-search-engines.jsonld"}}