

Cohere · v5 · 2× · last seen Oct 01, 2026
Embed 5 is Cohere's new embedding model family for enterprise search, RAG, and agent workflows, consisting of Pro (highest retrieval quality) and Fast (low latency, high throughput) variants. Both models share a common vector space, enabling a corpus indexed with Pro to be queried with Fast. The models process text, images, and mixed text-image inputs (e.g., PDF pages) in over 100 languages with a context window of 128,000 Tokens and support Matryoshka embeddings in multiple dimensions as well as float, int8, and binary output formats. Embed 5 is available via the Cohere API, Model Vault (single-tenant), Microsoft Foundry, and Amazon SageMaker, as well as for private VPC and on-premises deployments via vLLM.
Features
| Deployment (Self-Hosted/Cloud) | Cloud via API/Model Vault/Foundry/SageMaker; private VPC or on-premises via vLLM |
| Throughput/Latency | Fast processes ~377 documents/sec vs. ~160 documents/sec for Pro (~2.4x higher throughput) |
| License | Proprietary model; weights not publicly released, licensed for private deployment via vLLM |
| Platform | Cohere API, Model Vault, Microsoft Foundry, Amazon SageMaker, North |
| Price | Pro: $0.12/1M tokens (text), $0.40/1M tokens (image); Fast: $0.08/1M tokens (text), $0.40/1M tokens (image) |
| Protocol Compatibility | 128k token context window, Matryoshka embeddings in 256/512/768/1024/1536/2048 dims, float/int8/binary |
| Release Date | September 30, 2026 |
| Supported Models/Providers | Two variants: embed-v5.0-pro and embed-v5.0-fast, sharing a common embedding space |