{"slug": "cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search", "title": "Cohere launches Embed 5 with a cheaper model for live enterprise search", "summary": "Cohere released Embed 5 on September 30th, a pair of embedding models for enterprise search and retrieval that share an embedding space so customers can index with the higher-quality Pro model and query with the cheaper Fast model without rebuilding their index. Cohere priced Pro at $0.12 per million text tokens and Fast at $0.08, with image inputs at $0.40 per million tokens for either model, and reported company-run scores of 85.8 on ViDoRe V3 for Pro and 84.5 for Fast, ahead of Voyage 4 Large at 83.7 and Google's Gemini Embedding 2 at 83.2. Cohere also introduced RCP-nDCG@10, an AI-judge retrieval-scoring method used in its own evaluations, which the company says it validated against human judgment.", "body_md": "# Cohere launches Embed 5 with a cheaper model for live enterprise search\n\n**The Pro and Fast models share an embedding space, letting customers index with one and run queries with the other without rebuilding their index.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [X](https://x.com/cohere/status/2105285142394896435)\n\n## Why it matters\n\nCohere is making enterprise retrieval a cost-and-latency choice: customers can index with Pro and query with the cheaper Fast model in the same embedding space. Its performance claims rely on company-run evaluations, including a new Cohere scoring method.\n\n[Cohere](https://cohere.com/blog/embed-5) released Embed 5 on September 30th, a pair of [embedding](https://runtimewire.com/models/huggingface/pyannote-embedding-1dac8935eac8dbcc) models aimed at enterprise search and retrieval systems: Pro for retrieval quality and Fast for lower-latency, higher-volume queries. The models share an embedding space, so customers can build an index with Pro and query it with Fast. Cohere co-founder and CEO [Aidan Gomez (@aidangomez)](https://x.com/aidangomez), one of the co-authors of the 2017 paper \"Attention Is All You Need,\" is leading a company still competing on the practical infrastructure around enterprise AI, not only on generative models. Cohere announced Embed 5 in [a thread on X](https://x.com/cohere/status/2105285142394896435).\n\nThe shared space is the product's central operating proposition. Embedding models turn text, images, or other inputs into vectors that software can search by meaning. Companies use them to find relevant records before a generative model drafts an answer or an AI agent takes an action. Embed 5 Pro costs $0.12 per million text tokens; Fast costs $0.08. Image inputs cost $0.40 per million tokens for either model, according to [Cohere's product announcement](https://cohere.com/blog/embed-5).\n\nThe pairing gives customers a potential way to reserve the more expensive model for indexing large collections and use the cheaper one on the repeated query path. Cohere says the two versions can be mixed without rebuilding the index, and recommends indexing with Pro and querying with Fast. That is a potentially useful cost and deployment option; the published price alone does not establish total savings, which depend on workload, data volume, latency needs and downstream search infrastructure.\n\nCohere reports that Pro scored 85.8 on ViDoRe V3, compared with 83.7 for Voyage 4 Large and 83.2 for Google's Gemini Embedding 2. Fast scored 84.5. These are company-reported results, not an independent test. Cohere also introduced RCP-nDCG@10, a retrieval-scoring method used in its evaluations. The method grades documents against query-specific relevance criteria with an AI judge, rather than relying only on pre-existing relevance labels. Cohere says it validated the method against human judgment, but its own explanation describes the evaluation as a new approach the company developed and used to optimize its models. Comparisons using it deserve that context.\n\nThe benchmark details narrow the claim further. Cohere says the ViDoRe V3 documents were parsed as text and that ViDoRe's authors curated the evaluation material. For a separate parsed-document suite, Cohere reports an average of 84.8 for Pro, ahead of Voyage 4 Large at 83.6 and Gemini Embedding 2 at 80.8; the documents in that suite were parsed using Gemini 1.5 Flash. Those results support a specific proposition about retrieval on the tested collections. They do not establish that Embed 5 will outperform competitors on every company's internal records, parsing pipeline or search task.\n\nCohere's financial-document comparisons are especially relevant to its enterprise positioning. The company says Pro ranked first on its tests of FinanceBench, FinQA and ViDoRe V3 Finance, while Fast placed second on each. Financial filings, reports and spreadsheets often combine dense text with tables and layout that can be lost when documents are converted to plain text. Embed 5 accepts text, images and mixed text-image inputs, and both versions support a 128K-token context window, more than 100 languages and several vector sizes and output formats.\n\nFast's value proposition also rests on throughput. Cohere reports that it processed documents at an average 2.4 times Pro's throughput in its tests. The company positions Pro for offline indexing and Fast for interactive search, high-volume retrieval and agent workflows, where repeated searches can make latency consequential. Those throughput figures are also company-reported; the announcement does not provide a customer workload or independent deployment result demonstrating the same performance in production.\n\nEmbed 5 is available through the Cohere API, Model Vault, [Microsoft Foundry](https://ai.azure.com/) and [Amazon SageMaker](https://aws.amazon.com/). Cohere's [model documentation](https://docs.cohere.com/changelog/embed-v5) says it supports single-tenant deployment through Model Vault, giving customers a deployment option beyond shared API access. The company's established focus on enterprise deployments sits behind the model's emphasis on private infrastructure, multilingual documents and retrieval quality.\n\nGomez's technical history gives the launch a founder-specific throughline: before co-founding Cohere, he worked at Google Brain and co-authored \"Attention Is All You Need,\" the paper that introduced the Transformer architecture. His own [biography](https://aidangomez.ca/) says he also led researchers at For.ai before Cohere. Embed 5 puts that research legacy to work on a narrower commercial problem: making enterprise data easier to find, at a cost and speed that companies can use in search and agent systems.\n\nCohere raised $500 million at a $6.8 billion valuation in August 2025, in a round led by Radical Ventures and Inovia Capital, according to the company's [funding announcement](https://cohere.com/blog/august-2025-funding-round). Embed 5 does not disclose customer adoption, revenue or independent benchmark validation in its launch material. The immediate bet is measurable in the product design: Cohere is selling retrieval as an operational choice between a higher-scoring index model and a cheaper query model, tied together by a shared vector space.", "url": "https://wpnews.pro/news/cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search", "canonical_source": "https://runtimewire.com/article/cohere-embed-5-retrieval-models", "published_at": "2026-10-01 01:48:09+00:00", "updated_at": "2026-10-01 02:17:48.315479+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-infrastructure", "large-language-models", "ai-search"], "entities": ["Cohere", "Embed 5", "Aidan Gomez", "Voyage 4 Large", "Gemini Embedding 2", "ViDoRe V3", "FinanceBench", "RCP-nDCG@10"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search", "markdown": "https://wpnews.pro/news/cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search.md", "text": "https://wpnews.pro/news/cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search.txt", "jsonld": "https://wpnews.pro/news/cohere-launches-embed-5-with-a-cheaper-model-for-live-enterprise-search.jsonld"}}