Voyage AI vs Cohere Embed v4 vs Nemotron 3 Embed: Production RAG Trade-offs A developer benchmarked Voyage 4, Cohere Embed v4, and NVIDIA Nemotron 3 Embed for production RAG, concluding that the models' published scores are not directly comparable because Nemotron's figures come from retrieval-specific RTEB runs (78.46 for the 8B checkpoint, 72.38 for the 1B) while Voyage and Cohere numbers circulating publicly are general MTEB aggregates. The writeup argues embedding selection should be treated as an index architecture decision, with model choice validated against a customer's own labeled corpus and the metric harness verified before any vectors are scored. Embedding selection looks reversible until a production index contains tens of millions of vectors. At that point, changing the model can mean re-reading every source document, reproducing the original chunking pipeline, paying for another embedding pass, building a second index, and switching traffic without mixing incompatible vector spaces. That operational commitment—not a leaderboard position—was why we evaluated Voyage 4, Cohere Embed v4, and NVIDIA Nemotron 3 Embed. The three options represent materially different deployment choices: Under the hood, each model converts documents and queries into vectors that an index compares by cosine similarity, dot product, or an equivalent distance function. The implementation details differ, but the production contract is the same: document and query preprocessing must remain stable, dimensions must match, and we need continued access to a query encoder compatible with the stored vectors, rather than necessarily the exact model version originally used. We found an immediate problem with the public comparison narrative. The available benchmark figures were not generated on one common evaluation suite. The Nemotron 3 figures we examined included 78.46 on RTEB and 75.45 on MMTEB Retrieval for the 8B checkpoint. The 1B checkpoint was listed at 72.38 RTEB and 71.04 MMTEB Retrieval. Figures circulating for Voyage and Cohere were general MTEB aggregates, not directly comparable RTEB runs. We therefore refused to place those values into a single “winner” ranking. A 78.46 retrieval score and a roughly 67 general MTEB aggregate do not measure the same thing. MTEB can blend retrieval, clustering, classification, and semantic-similarity tasks. RTEB is retrieval-specific. Different task sets, relevance judgments, languages, pooling policies, and gain conventions can reverse an apparent lead. This matters because a model can perform well on public web retrieval and still miss exact clauses in contracts, versioned API behavior, product identifiers, or cross-language support content. Our deployment recommendation consequently depends on a customer’s labeled corpus rather than the headline score. We treat the model choice as an index architecture decision. Teams evaluating the surrounding stack can compare it with the vector and retrieval tools in our AI tools collection https://dev.to/tools , but they should freeze the embedding contract before optimizing the database. We started by validating the metric harness rather than sending three sets of vectors into an untrusted scorer. That step caught a class of benchmark errors more damaging than a slow API: document-ID misalignment can turn perfect retrieval into a score of zero while every individual array still has the expected shape. We ran the tests in Python 3.12 with these pinned dependencies: python -m venv .venv source .venv/bin/activate python -m pip install \ numpy==2.1.3 \ scipy==1.14.1 \ scikit-learn==1.5.2 python benchmark metric contract.py The following compact fixture reproduces the key cutoff and all-zero-relevance checks; the document-ID alignment results are reported separately below. It is runnable without an API key or GPU. python import json import numpy as np from sklearn.metrics import ndcg score cutoff = 10 document count = 12 def score relevance, ranking scores : y true = np.asarray relevance , dtype=float y score = np.asarray ranking scores , dtype=float return float ndcg score y true, y score, k=cutoff Higher retrieval scores produce earlier ranks. descending scores = list range document count, 0, -1 def one relevant at rank rank : relevance = 0.0 document count relevance rank - 1 = 1.0 return score relevance, descending scores results = { "single relevant at rank 1": one relevant at rank 1 , "single relevant at rank 2": one relevant at rank 2 , "single relevant at rank 10": one relevant at rank 10 , "single relevant at rank 11": one relevant at rank 11 , "all zero relevance query": score 0.0 document count, descending scores, , } assert results "single relevant at rank 1" == 1.0 assert results "single relevant at rank 10" 0.0 assert results "single relevant at rank 11" == 0.0 assert results "all zero relevance query" == 0.0 print json.dumps results, indent=2 Our executed fixture produced: { "single relevant at rank 1": 1.0, "single relevant at rank 2": 0.6309297535714574, "single relevant at rank 10": 0.2890648263178878, "single relevant at rank 11": 0.0, "all zero relevance query": 0.0 } We also tested graded relevance. Swapping two graded results produced an nDCG@10 of 0.9224945116765988 with linear gains and 0.842828264880938 when we supplied exponential gains. That absolute difference of 0.07966624679566081 is large enough to manufacture a benchmark “win” if two runs use different gain conventions. The alignment fixture exposed an even larger scoring error: aligned score=1.0 joint column permutation score=1.0 judgments only permutation score=0.0 all assertions passed=true Permuting document IDs in both predictions and judgments preserved the score. Permuting only the relevance judgments reduced it from 1.0 to 0.0. We now require one canonical document-ID map across chunk storage, relevance judgments, cached vectors, and result files. This was a measurement-contract test, not a model-quality run. We did not produce vendor embeddings in this executed sandbox, so we do not present invented nDCG, recall, latency, throughput, or multilingual results for Voyage, Cohere, or Nemotron. For a real head-to-head, our required run manifest is: { "corpus": { "chunks": "2000-5000", "chunker version": "frozen", "content hashes": "required" }, "queries": { "relevance judgments": "graded", "zero relevance policy": "declared", "languages": "stratified" }, "retrieval": { "metric": "nDCG@10", "Recall@5", "Recall@10", "Recall@50" , "similarity": "cosine", "normalization": "explicit", "reranker": "disabled" }, "execution": { "model version": "pinned", "dimension": "recorded", "input type": "recorded", "truncation": "disabled-or-counted", "batch and realtime": "separate" } } Voyage’s embedding reference https://docs.voyageai.com/docs/embeddings and official Python client https://github.com/voyage-ai/voyageai-python expose the necessary model, input type, output dimension, truncation, and output data type controls. Teams should pin Cohere’s equivalent configuration using its Embed reference https://docs.cohere.com/docs/embed , rather than infer it from a framework wrapper. The first failure was not an API error. It was evidence comparability. We could not reproduce a valid three-model leaderboard from the supplied material because no common fixed-corpus run existed. Any article naming an unconditional winner from these inputs would be overstating the evidence. We hit six additional production hazards. 1. Dimensions are a schema, not a tuning flag. Voyage 4 supports 256, 512, 1,024, and 2,048 dimensions, with 1,024 as the default. Cohere Embed v4 configurations include 256, 512, 1,024, and 1,536 dimensions. The Nemotron 3 material we reviewed described 2,048 dimensions for the 1B model and 4,096 for 8B. A collection created at 1,024 dimensions cannot accept a 1,536- or 4,096-element vector. We plan a new collection, validation, and traffic cutover for a dimension change. Switching to an incompatible embedding space requires re-embedding; reducing compatible Matryoshka vectors may allow us to reuse existing embeddings after validating the transformation and retrieval quality. We store model, revision, dimension, normalization, data type, and preprocessing hash beside every index. 2. Voyage silently truncates by default. Voyage’s truncation option defaults to True . An oversized input can therefore return a valid vector for incomplete content. That is worse than an explicit failure when the missing text contains the relevant clause. We set truncation=False in ingestion jobs and route over-length chunks to a dead-letter queue. We would rather fail visibly than index partial documents silently. 3. Query and document modes are operationally significant. Voyage prepends different retrieval prompts when input type is query or document . Cohere also distinguishes retrieval input types. We treat this field as part of the model version. A missing query mode is not harmless configuration drift. 4. Request size is not the same as rate limit. Voyage accepts at most 1,000 texts per request, with additional token ceilings. The documented total is up to one million tokens for voyage-4-lite , 320,000 for voyage-4 , and 120,000 for voyage-4-large . Those are payload limits, not throughput guarantees. We could not verify per-key requests-per-minute, token-per-minute limits, cooldown behavior, or reserved-capacity terms across all three providers. Teams should load-test their own account tier rather than extrapolate from maximum batch size. 5. Quantization support is uneven. Voyage’s hosted 4-series can return float, int8, uint8, binary, and unsigned-binary output. The local voyage-4-nano path does not support int8 or uint8 because its required calibration ranges are not shipped; those modes raise NotImplementedError . Float and binary formats remain available locally. The local extra also requires Python 3.10 or newer and installs PyTorch plus Sentence Transformers. We kept API-only and local-inference environments separate to avoid turning a lightweight client image into a large ML runtime. 6. Nemotron naming can produce the wrong comparison. We found material mixing Nemotron 3 Embed with the older llama-nemotron-embed-1b-v2 . The older model has an 8,192-token input limit and Matryoshka dimensions down from 2,048. Those specifications must not be copied onto the newer Nemotron 3 family described with a 32,768-token context. We pin exact model identifiers and artifact revisions. “Nemotron Embed” is not precise enough for an index manifest. We normalized the product configurations available during the review, but we did not measure service latency or GPU throughput. The latency column is therefore deliberately absent rather than filled with synthetic numbers. | Option | Deployment | Context limit | Dimensions reviewed | Multimodal | Public price used for planning | Main operational trade-off | |---|---|---|---|---|---|---| | Voyage 4 Lite | Managed API | 32,000 | 256-2,048 | No | $0.02 per 1M tokens | Lowest listed API rate; quality must be tested by domain | | Voyage 4 | Managed API | 32,000 | 256-2,048 | No | $0.06 per 1M tokens | Balanced managed text retrieval | | Voyage 4 Large | Managed API | 32,000 | 256-2,048 | No | $0.12 per 1M tokens | Higher-cost quality tier | | Cohere Embed v4 | Managed API | 128,000 | 256-1,536 | Yes | $0.12 per 1M text tokens | Best fit here for mixed text and image retrieval | | Nemotron 3 Embed 1B | Self-hosted | 32,768 | 2,048 | No | No per-token license charge | GPU serving, observability, and capacity become our responsibility | | Nemotron 3 Embed 8B | Self-hosted | 32,768 | 4,096 | No | No per-token license charge | Largest raw vector payload at the reviewed dimensions; serving capacity must be measured | At one billion text tokens, the simple API bill is approximately: These numbers expose an important reality: initial embedding API charges can be smaller than vector storage, indexing, replicas, backups, and migration engineering. At float32, one million raw vectors require approximately: That excludes IDs, graph edges, metadata, allocator overhead, replicas, and snapshots. Moving from 1,024 to 4,096 dimensions quadruples the raw vector payload before the vector database adds its own structures. For Nemotron, we use a break-even equation rather than pretending open weights are free: self hosted cost per million = gpu hourly cost / millions of tokens per hour At a hypothetical $2 per GPU-hour, self-hosting must sustain more than 16.67 million tokens per hour to beat a $0.12-per-million API on raw compute. It must exceed 33.33 million tokens per hour to beat a $0.06-per-million API. Engineering labor, idle capacity, failover, deployment, and monitoring move the real threshold higher. We did not measure Nemotron throughput, memory use, or NVFP4 quality loss, so those break-even examples are formulas, not performance claims. For teams planning a migration or capacity model, our AI infrastructure services https://dev.to/services focus on the complete cost boundary: embedding generation, index memory, replicas, refresh cadence, and rollback—not just the token invoice. We cannot honestly name a universal retrieval-quality winner from the available evidence. Nemotron 3 has the strongest retrieval-specific public numbers in this comparison, but those numbers were not generated in the same run as Voyage 4 and Cohere Embed v4. Our executed work validated the metric harness, not the three models. That distinction defines our verdict. Deploy Voyage 4 if: input type , reject truncation, and keep a model manifest. Deploy Cohere Embed v4 if: Deploy Nemotron 3 Embed if: Hold off or avoid all three as a final choice if: Our practical default is Voyage 4 at 1,024 dimensions for managed, text-only evaluation; Cohere Embed v4 when multimodal retrieval is a real requirement; and Nemotron 3 only when self-hosting is an architectural objective rather than a reaction to per-token pricing. We would not approve production deployment until all candidates had run over the same 2,000-to-5,000-chunk corpus, the same graded queries, the same distance metric, and the same nDCG and recall implementation. The winning model is the one that clears the application’s recall threshold at the lowest total operating cost—not the one with the largest vector or the most flattering public benchmark. If the index migration risk is already material, contact our engineering team https://dev.to/contact before changing dimensions or model families. Rebuilding a benchmark is cheap. Rebuilding a live retrieval index under customer traffic is not.