# Vector Search & Embeddings in Java: Building Semantic Search Engines

> Source: <https://dev.to/said_olano/vector-search-embeddings-in-java-building-semantic-search-engines-279c>
> Published: 2026-10-02 01:41:21+00:00

Traditional keyword-based search is dead. When users search for "best restaurants near me," they don't expect results matching those exact words—they expect restaurants that *mean* the right thing.

This is where **vector search and embeddings** enter the picture.

Vector search is the technology powering modern AI applications: ChatGPT's retrieval-augmented generation (RAG), Netflix's recommendation engine, Spotify's "Discover Weekly," and enterprise semantic search platforms. Yet many Java developers still think of search as Elasticsearch queries with exact terms.

This gap is costing you:

In this guide, you'll learn:

By the end, you'll understand why vector search is essential for modern applications and how to build it with Java.

An **embedding** is a numerical representation of text, images, or other data. Instead of storing "restaurant recommendations," you store a vector of numbers: `[0.25, -0.15, 0.82, ..., 0.41]`.

These aren't random numbers. They're learned through **neural networks** trained on massive datasets. Similar concepts produce similar vectors. This property is the entire foundation of vector search.

**Example:**

`[0.12, 0.88, -0.31, 0.45, ...]`
`[0.14, 0.87, -0.29, 0.46, ...]`
These vectors are *close* in vector space. Measuring that distance (using cosine similarity, Euclidean distance, etc.) gives you a relevance score.

In production, you don't hand-craft embeddings. You use pre-trained embedding models:

These models map text → vector in a way that preserves semantic meaning.

All vectors live in an **N-dimensional space**. When you have 1,536-dimensional vectors (from OpenAI), you're working in 1,536-dimensional space.

**Similarity metrics:**

**Cosine Similarity** (most common)

`(A · B) / (||A|| × ||B||)`
**Euclidean Distance** (L2)

`√(Σ(ai - bi)²)`
**Manhattan Distance** (L1)

For text search, **cosine similarity** is almost always the right choice.

A semantic search system has these components:

```
┌─────────────────────────────────────────────┐
│  User Query                                  │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  1. Embedding Model (Convert text → vector) │
│     (OpenAI API / Local Sentence Transformer)│
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  2. Vector Search Engine                     │
│     (Pinecone / PostgreSQL pgvector)        │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  3. Similarity Ranking                       │
│     Return top-K most relevant results      │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  Results with scores (0.0 - 1.0)            │
└─────────────────────────────────────────────┘
```

**Step 1: Add Dependencies**

``` php
<!-- pom.xml -->
<dependency>
    <groupId>com.theokanning.openai-gpt3-java</groupId>
    <artifactId>api</artifactId>
    <version>0.18.1</version>
</dependency>

<dependency>
    <groupId>io.pinecone</groupId>
    <artifactId>pinecone-client</artifactId>
    <version>0.1.0</version>
</dependency>
```

**Step 2: Create Embedding Service**

``` python
import com.theokanning.openai.embedding.Embedding;
import com.theokanning.openai.embedding.EmbeddingRequest;
import com.theokanning.openai.embedding.EmbeddingResult;
import com.theokanning.openai.service.OpenAiService;

public class EmbeddingService {
    private final OpenAiService openAiService;
    private final String modelId = "text-embedding-ada-002";

    public EmbeddingService(String apiKey) {
        this.openAiService = new OpenAiService(apiKey);
    }

    public List<Double> embedText(String text) {
        EmbeddingRequest request = EmbeddingRequest.builder()
            .model(modelId)
            .input(Collections.singletonList(text))
            .build();

        EmbeddingResult result = openAiService.createEmbeddings(request);
        return result.getData().get(0).getEmbedding();
    }

    public List<List<Double>> embedTexts(List<String> texts) {
        EmbeddingRequest request = EmbeddingRequest.builder()
            .model(modelId)
            .input(texts)
            .build();

        EmbeddingResult result = openAiService.createEmbeddings(request);
        return result.getData().stream()
            .sorted(Comparator.comparingInt(Embedding::getIndex))
            .map(Embedding::getEmbedding)
            .collect(Collectors.toList());
    }
}
```

**Step 3: Pinecone Vector Search**

``` python
import io.pinecone.clients.Pinecone;
import io.pinecone.clients.Index;

public class PineconeVectorStore {
    private final Index index;
    private final EmbeddingService embeddingService;

    public PineconeVectorStore(String apiKey, String projectName, 
                               String indexName, String embeddingApiKey) {
        Pinecone client = new Pinecone.Builder()
            .withApiKey(apiKey)
            .build();

        this.index = client.getIndex(projectName, indexName);
        this.embeddingService = new EmbeddingService(embeddingApiKey);
    }

    // Index documents with embeddings
    public void indexDocument(String docId, String content, 
                             Map<String, String> metadata) {
        List<Double> embedding = embeddingService.embedText(content);

        index.upsert(
            docId,
            embedding,
            metadata
        );
    }

    // Search: returns top K similar documents
    public List<SearchResult> search(String query, int topK) {
        List<Double> queryEmbedding = embeddingService.embedText(query);

        var results = index.query(queryEmbedding)
            .withTopK(topK)
            .withIncludeMetadata(true)
            .execute();

        return results.getMatches().stream()
            .map(match -> new SearchResult(
                match.getId(),
                match.getScore(),
                match.getMetadata()
            ))
            .collect(Collectors.toList());
    }
}

record SearchResult(String id, Double score, Map<String, String> metadata) {}
```

**Step 4: End-to-End Usage**

```
public class SemanticSearchApp {
    public static void main(String[] args) {
        PineconeVectorStore store = new PineconeVectorStore(
            System.getenv("PINECONE_API_KEY"),
            "my-project",
            "restaurant-index",
            System.getenv("OPENAI_API_KEY")
        );

        // Index documents
        store.indexDocument("rest-1", "Best pizza in town, authentic Italian, family-owned", 
            Map.of("name", "Pasta Paradise", "type", "pizza"));

        store.indexDocument("rest-2", "Fast casual ramen, Tokyo-style noodles, great broth",
            Map.of("name", "Noodle House", "type", "ramen"));

        // Search with semantic understanding
        var results = store.search("where to find good Italian food", 3);

        results.forEach(r -> 
            System.out.printf("ID: %s, Score: %.4f, Name: %s%n",
                r.id(), r.score(), r.metadata().get("name"))
        );
    }
}
```

For privacy-sensitive applications, you might want embeddings to stay within your infrastructure.

**Step 1: Setup PostgreSQL pgvector**

```
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Create table
CREATE TABLE documents (
    id SERIAL PRIMARY KEY,
    content TEXT NOT NULL,
    embedding vector(1536),
    metadata JSONB,
    created_at TIMESTAMP DEFAULT NOW()
);

-- Create HNSW index for fast similarity search
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
```

**Step 2: Java Implementation with Sentence Transformers**

``` python
import org.springframework.ai.document.Document;
import org.springframework.ai.embedding.Embedding;
import org.springframework.ai.embedding.EmbeddingModel;
import org.springframework.jdbc.core.JdbcTemplate;

@Component
public class PostgresVectorStore {
    private final JdbcTemplate jdbcTemplate;
    private final EmbeddingModel embeddingModel;

    @Autowired
    public PostgresVectorStore(JdbcTemplate jdbcTemplate, 
                               EmbeddingModel embeddingModel) {
        this.jdbcTemplate = jdbcTemplate;
        this.embeddingModel = embeddingModel;
    }

    public void indexDocument(String content, String metadata) {
        Embedding embedding = embeddingModel.embed(content);

        String sql = "INSERT INTO documents (content, embedding, metadata) " +
                     "VALUES (?, ?::vector, ?::jsonb)";

        jdbcTemplate.update(sql, content, vectorToString(embedding), metadata);
    }

    public List<DocumentResult> similaritySearch(String query, int limit) {
        Embedding queryEmbedding = embeddingModel.embed(query);

        String sql = "SELECT id, content, metadata, " +
                     "1 - (embedding <=> ?::vector) as similarity " +
                     "FROM documents " +
                     "ORDER BY embedding <=> ?::vector " +
                     "LIMIT ?";

        return jdbcTemplate.query(sql, new Object[]{
            vectorToString(queryEmbedding),
            vectorToString(queryEmbedding),
            limit
        }, (rs, rowNum) -> new DocumentResult(
            rs.getInt("id"),
            rs.getString("content"),
            rs.getDouble("similarity"),
            rs.getString("metadata")
        ));
    }

    private String vectorToString(Embedding embedding) {
        return "[" + embedding.getOutput().stream()
            .map(String::valueOf)
            .collect(Collectors.joining(",")) + "]";
    }
}

record DocumentResult(int id, String content, double score, String metadata) {}
```

Traditional: Search for "comfortable shoes" → matches only products with those exact words

Semantic: Matches "ergonomic footwear," "supportive sneakers," "cushioned athletic shoes"

```
public class ProductSearchService {
    private final VectorStore vectorStore;

    public List<Product> findSimilarProducts(String query) {
        return vectorStore.search(query, 10).stream()
            .map(this::toProduct)
            .collect(Collectors.toList());
    }
}
```

Retrieval-Augmented Generation: Embed your knowledge base, find relevant docs, feed them to LLM

```
@RestController
public class SupportChatbot {
    private final VectorStore knowledgeBase;
    private final OpenAiService openAiService;

    @PostMapping("/ask")
    public ResponseEntity<String> ask(@RequestBody String question) {
        // Step 1: Find relevant docs using vector search
        var relevantDocs = knowledgeBase.search(question, 3);

        // Step 2: Build context from retrieved docs
        String context = relevantDocs.stream()
            .map(SearchResult::content)
            .collect(Collectors.joining("\n\n"));

        // Step 3: Ask LLM with context
        String prompt = String.format(
            "Based on this knowledge base:\n%s\n\nAnswer: %s",
            context, question
        );

        ChatCompletionRequest request = ChatCompletionRequest.builder()
            .model("gpt-4")
            .messages(List.of(new ChatMessage(ChatMessageRole.USER.value(), prompt)))
            .build();

        String answer = openAiService.createChatCompletion(request)
            .getChoices().get(0).getMessage().getContent();

        return ResponseEntity.ok(answer);
    }
}
```

Find near-duplicate documents or suspicious fraud patterns

```
public class DuplicateDetector {
    private final VectorStore vectorStore;

    public boolean isProbablyDuplicate(String document, double threshold) {
        var similarDocs = vectorStore.search(document, 1);

        return !similarDocs.isEmpty() && 
               similarDocs.get(0).score() > threshold;
    }
}
```

Don't embed one document at a time. Batch them.

```
public void indexManyDocuments(List<Document> documents) {
    // ❌ Slow: N API calls
    // documents.forEach(doc -> index(doc));

    // ✅ Fast: 1 API call per batch
    Iterables.partition(documents, 100).forEach(batch -> {
        List<String> texts = batch.stream()
            .map(Document::getContent)
            .collect(Collectors.toList());

        List<List<Double>> embeddings = embeddingService.embedTexts(texts);

        for (int i = 0; i < batch.size(); i++) {
            vectorStore.index(batch.get(i).getId(), embeddings.get(i));
        }
    });
}
```

Store computed embeddings to avoid redundant API calls

```
@Component
public class CachedEmbeddingService {
    private final EmbeddingService service;
    private final Map<String, List<Double>> cache = new ConcurrentHashMap<>();

    public List<Double> embed(String text) {
        return cache.computeIfAbsent(text, key -> service.embedText(key));
    }
}
```

Not all 1,536 dimensions are necessary for your use case. PCA can reduce them:

```
public List<Double> reduceDimensions(List<Double> embedding, int targetDim) {
    // Use Apache Commons Math or similar
    // This trades accuracy for speed/storage
    return PCA.reduce(embedding, targetDim);
}
```

If this is too slow, cache frequent queries.

For high-volume applications, self-hosting saves money.

When document content changes, re-embed and update:

```
public void updateDocument(String docId, String newContent) {
    List<Double> newEmbedding = embeddingService.embedText(newContent);
    vectorStore.update(docId, newEmbedding, newContent);
}
```

Track embedding quality:

```
@Component
public class EmbeddingQualityMonitor {
    private final MeterRegistry meterRegistry;

    public void recordSimilarityScore(double score) {
        Timer.builder("search.similarity.score")
            .publishPercentiles(0.5, 0.95, 0.99)
            .register(meterRegistry)
            .record(Duration.ofMillis((long) (score * 1000)));
    }
}
```

Vector search and embeddings are no longer bleeding-edge. They're essential for:

**Key Takeaways:**

Your next semantic search implementation is just these components away. Build it today.
