Traditional keyword-based search is dead. When users search for "best restaurants near me," they don't expect results matching those exact wordsβthey expect restaurants that mean the right thing.
This is where vector search and embeddings enter the picture.
Vector search is the technology powering modern AI applications: ChatGPT's retrieval-augmented generation (RAG), Netflix's recommendation engine, Spotify's "Discover Weekly," and enterprise semantic search platforms. Yet many Java developers still think of search as Elasticsearch queries with exact terms.
This gap is costing you:
In this guide, you'll learn:
By the end, you'll understand why vector search is essential for modern applications and how to build it with Java.
An embedding is a numerical representation of text, images, or other data. Instead of storing "restaurant recommendations," you store a vector of numbers: [0.25, -0.15, 0.82, ..., 0.41].
These aren't random numbers. They're learned through neural networks trained on massive datasets. Similar concepts produce similar vectors. This property is the entire foundation of vector search.
Example:
[0.12, 0.88, -0.31, 0.45, ...]
[0.14, 0.87, -0.29, 0.46, ...]
These vectors are close in vector space. Measuring that distance (using cosine similarity, Euclidean distance, etc.) gives you a relevance score.
In production, you don't hand-craft embeddings. You use pre-trained embedding models:
These models map text β vector in a way that preserves semantic meaning.
All vectors live in an N-dimensional space. When you have 1,536-dimensional vectors (from OpenAI), you're working in 1,536-dimensional space.
Similarity metrics:
Cosine Similarity (most common)
(A Β· B) / (||A|| Γ ||B||)
Euclidean Distance (L2)
β(Ξ£(ai - bi)Β²)
Manhattan Distance (L1)
For text search, cosine similarity is almost always the right choice.
A semantic search system has these components:
βββββββββββββββββββββββββββββββββββββββββββββββ
β User Query β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 1. Embedding Model (Convert text β vector) β
β (OpenAI API / Local Sentence Transformer)β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 2. Vector Search Engine β
β (Pinecone / PostgreSQL pgvector) β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 3. Similarity Ranking β
β Return top-K most relevant results β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β Results with scores (0.0 - 1.0) β
βββββββββββββββββββββββββββββββββββββββββββββββ
Step 1: Add Dependencies
<!-- pom.xml -->
<dependency>
<groupId>com.theokanning.openai-gpt3-java</groupId>
<artifactId>api</artifactId>
<version>0.18.1</version>
</dependency>
<dependency>
<groupId>io.pinecone</groupId>
<artifactId>pinecone-client</artifactId>
<version>0.1.0</version>
</dependency>
Step 2: Create Embedding Service
import com.theokanning.openai.embedding.Embedding;
import com.theokanning.openai.embedding.EmbeddingRequest;
import com.theokanning.openai.embedding.EmbeddingResult;
import com.theokanning.openai.service.OpenAiService;
public class EmbeddingService {
private final OpenAiService openAiService;
private final String modelId = "text-embedding-ada-002";
public EmbeddingService(String apiKey) {
this.openAiService = new OpenAiService(apiKey);
}
public List<Double> embedText(String text) {
EmbeddingRequest request = EmbeddingRequest.builder()
.model(modelId)
.input(Collections.singletonList(text))
.build();
EmbeddingResult result = openAiService.createEmbeddings(request);
return result.getData().get(0).getEmbedding();
}
public List<List<Double>> embedTexts(List<String> texts) {
EmbeddingRequest request = EmbeddingRequest.builder()
.model(modelId)
.input(texts)
.build();
EmbeddingResult result = openAiService.createEmbeddings(request);
return result.getData().stream()
.sorted(Comparator.comparingInt(Embedding::getIndex))
.map(Embedding::getEmbedding)
.collect(Collectors.toList());
}
}
Step 3: Pinecone Vector Search
import io.pinecone.clients.Pinecone;
import io.pinecone.clients.Index;
public class PineconeVectorStore {
private final Index index;
private final EmbeddingService embeddingService;
public PineconeVectorStore(String apiKey, String projectName,
String indexName, String embeddingApiKey) {
Pinecone client = new Pinecone.Builder()
.withApiKey(apiKey)
.build();
this.index = client.getIndex(projectName, indexName);
this.embeddingService = new EmbeddingService(embeddingApiKey);
}
// Index documents with embeddings
public void indexDocument(String docId, String content,
Map<String, String> metadata) {
List<Double> embedding = embeddingService.embedText(content);
index.upsert(
docId,
embedding,
metadata
);
}
// Search: returns top K similar documents
public List<SearchResult> search(String query, int topK) {
List<Double> queryEmbedding = embeddingService.embedText(query);
var results = index.query(queryEmbedding)
.withTopK(topK)
.withIncludeMetadata(true)
.execute();
return results.getMatches().stream()
.map(match -> new SearchResult(
match.getId(),
match.getScore(),
match.getMetadata()
))
.collect(Collectors.toList());
}
}
record SearchResult(String id, Double score, Map<String, String> metadata) {}
Step 4: End-to-End Usage
public class SemanticSearchApp {
public static void main(String[] args) {
PineconeVectorStore store = new PineconeVectorStore(
System.getenv("PINECONE_API_KEY"),
"my-project",
"restaurant-index",
System.getenv("OPENAI_API_KEY")
);
// Index documents
store.indexDocument("rest-1", "Best pizza in town, authentic Italian, family-owned",
Map.of("name", "Pasta Paradise", "type", "pizza"));
store.indexDocument("rest-2", "Fast casual ramen, Tokyo-style noodles, great broth",
Map.of("name", "Noodle House", "type", "ramen"));
// Search with semantic understanding
var results = store.search("where to find good Italian food", 3);
results.forEach(r ->
System.out.printf("ID: %s, Score: %.4f, Name: %s%n",
r.id(), r.score(), r.metadata().get("name"))
);
}
}
For privacy-sensitive applications, you might want embeddings to stay within your infrastructure.
Step 1: Setup PostgreSQL pgvector
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Create table
CREATE TABLE documents (
id SERIAL PRIMARY KEY,
content TEXT NOT NULL,
embedding vector(1536),
metadata JSONB,
created_at TIMESTAMP DEFAULT NOW()
);
-- Create HNSW index for fast similarity search
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Step 2: Java Implementation with Sentence Transformers
import org.springframework.ai.document.Document;
import org.springframework.ai.embedding.Embedding;
import org.springframework.ai.embedding.EmbeddingModel;
import org.springframework.jdbc.core.JdbcTemplate;
@Component
public class PostgresVectorStore {
private final JdbcTemplate jdbcTemplate;
private final EmbeddingModel embeddingModel;
@Autowired
public PostgresVectorStore(JdbcTemplate jdbcTemplate,
EmbeddingModel embeddingModel) {
this.jdbcTemplate = jdbcTemplate;
this.embeddingModel = embeddingModel;
}
public void indexDocument(String content, String metadata) {
Embedding embedding = embeddingModel.embed(content);
String sql = "INSERT INTO documents (content, embedding, metadata) " +
"VALUES (?, ?::vector, ?::jsonb)";
jdbcTemplate.update(sql, content, vectorToString(embedding), metadata);
}
public List<DocumentResult> similaritySearch(String query, int limit) {
Embedding queryEmbedding = embeddingModel.embed(query);
String sql = "SELECT id, content, metadata, " +
"1 - (embedding <=> ?::vector) as similarity " +
"FROM documents " +
"ORDER BY embedding <=> ?::vector " +
"LIMIT ?";
return jdbcTemplate.query(sql, new Object[]{
vectorToString(queryEmbedding),
vectorToString(queryEmbedding),
limit
}, (rs, rowNum) -> new DocumentResult(
rs.getInt("id"),
rs.getString("content"),
rs.getDouble("similarity"),
rs.getString("metadata")
));
}
private String vectorToString(Embedding embedding) {
return "[" + embedding.getOutput().stream()
.map(String::valueOf)
.collect(Collectors.joining(",")) + "]";
}
}
record DocumentResult(int id, String content, double score, String metadata) {}
Traditional: Search for "comfortable shoes" β matches only products with those exact words
Semantic: Matches "ergonomic footwear," "supportive sneakers," "cushioned athletic shoes"
public class ProductSearchService {
private final VectorStore vectorStore;
public List<Product> findSimilarProducts(String query) {
return vectorStore.search(query, 10).stream()
.map(this::toProduct)
.collect(Collectors.toList());
}
}
Retrieval-Augmented Generation: Embed your knowledge base, find relevant docs, feed them to LLM
@RestController
public class SupportChatbot {
private final VectorStore knowledgeBase;
private final OpenAiService openAiService;
@PostMapping("/ask")
public ResponseEntity<String> ask(@RequestBody String question) {
// Step 1: Find relevant docs using vector search
var relevantDocs = knowledgeBase.search(question, 3);
// Step 2: Build context from retrieved docs
String context = relevantDocs.stream()
.map(SearchResult::content)
.collect(Collectors.joining("\n\n"));
// Step 3: Ask LLM with context
String prompt = String.format(
"Based on this knowledge base:\n%s\n\nAnswer: %s",
context, question
);
ChatCompletionRequest request = ChatCompletionRequest.builder()
.model("gpt-4")
.messages(List.of(new ChatMessage(ChatMessageRole.USER.value(), prompt)))
.build();
String answer = openAiService.createChatCompletion(request)
.getChoices().get(0).getMessage().getContent();
return ResponseEntity.ok(answer);
}
}
Find near-duplicate documents or suspicious fraud patterns
public class DuplicateDetector {
private final VectorStore vectorStore;
public boolean isProbablyDuplicate(String document, double threshold) {
var similarDocs = vectorStore.search(document, 1);
return !similarDocs.isEmpty() &&
similarDocs.get(0).score() > threshold;
}
}
Don't embed one document at a time. Batch them.
public void indexManyDocuments(List<Document> documents) {
// β Slow: N API calls
// documents.forEach(doc -> index(doc));
// β
Fast: 1 API call per batch
Iterables.partition(documents, 100).forEach(batch -> {
List<String> texts = batch.stream()
.map(Document::getContent)
.collect(Collectors.toList());
List<List<Double>> embeddings = embeddingService.embedTexts(texts);
for (int i = 0; i < batch.size(); i++) {
vectorStore.index(batch.get(i).getId(), embeddings.get(i));
}
});
}
Store computed embeddings to avoid redundant API calls
@Component
public class CachedEmbeddingService {
private final EmbeddingService service;
private final Map<String, List<Double>> cache = new ConcurrentHashMap<>();
public List<Double> embed(String text) {
return cache.computeIfAbsent(text, key -> service.embedText(key));
}
}
Not all 1,536 dimensions are necessary for your use case. PCA can reduce them:
public List<Double> reduceDimensions(List<Double> embedding, int targetDim) {
// Use Apache Commons Math or similar
// This trades accuracy for speed/storage
return PCA.reduce(embedding, targetDim);
}
If this is too slow, cache frequent queries.
For high-volume applications, self-hosting saves money.
When document content changes, re-embed and update:
public void updateDocument(String docId, String newContent) {
List<Double> newEmbedding = embeddingService.embedText(newContent);
vectorStore.update(docId, newEmbedding, newContent);
}
Track embedding quality:
@Component
public class EmbeddingQualityMonitor {
private final MeterRegistry meterRegistry;
public void recordSimilarityScore(double score) {
Timer.builder("search.similarity.score")
.publishPercentiles(0.5, 0.95, 0.99)
.register(meterRegistry)
.record(Duration.ofMillis((long) (score * 1000)));
}
}
Vector search and embeddings are no longer bleeding-edge. They're essential for:
Key Takeaways:
Your next semantic search implementation is just these components away. Build it today.