cd /news/artificial-intelligence/jina-embeddings-v4-as-an-openai-comp… · home topics artificial-intelligence article
[ARTICLE · art-108560] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

jina-embeddings-v4 as an OpenAI-Compatible Embeddings Server

A developer has released jina-embeddings-v4, a self-hosted server for the jina-embeddings-v4 embedding model with an OpenAI-compatible /v1/embeddings endpoint. The server runs on a single NVIDIA GPU and allows applications that call OpenAI for embeddings to switch by only changing the base URL. It supports tasks such as text-matching, retrieval, and code, and returns 2048-dimensional float32 vectors.

read3 min views1 publishedAug 24, 2026

jina-embeddings-v4 is a self-hosted server for the jina-embeddings-v4 embedding model with an OpenAI-compatible /v1/embeddings

endpoint. It runs on a single NVIDIA GPU. An application that calls OpenAI for embeddings can call this server instead. The request and response bodies are the same, so setting the client's base URL is the only change needed.

task

is code

.float32

vectorsCreate a docker-compose.yml

:

services:
  jina:
    image: edgaras0x4e/jina-embeddings-v4:latest
    ports:
      - "8081:80"
    volumes:
      - jina-cache:/root/.cache/huggingface
    environment:
      HF_HOME: /root/.cache/huggingface
      API_KEY: your-api-key-here
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    restart: unless-stopped

volumes:
  jina-cache:
docker compose up -d

The image download is about 7.8 GB, and 14.7 GB unpacked on disk. The prebuilt image needs no build step. Building from source (git clone

the repo, then docker compose up -d --build

) compiles flash-attn

against PyTorch 2.5.1 and CUDA 12.4, which takes 10 to 20 minutes.

The weights are not in the image: on first start the server downloads about 7 GB from Hugging Face and loads them into GPU memory. On later starts the server reads the weights from the volume instead of down them again.

When /health

returns ok

, the server is ready to embed:

curl http://localhost:8081/health
{"status":"ok"}
curl http://localhost:8081/v1/embeddings \
  -H "Authorization: Bearer your-api-key-here" \
  -H "Content-Type: application/json" \
  -d '{"input": ["The train leaves at eight in the morning"]}'
{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "embedding": [-0.008472833782434464, -0.024404730647802353, 0.007389210630208254, ...],
      "index": 0
    }
  ],
  "model": "jinaai/jina-embeddings-v4",
  "usage": {"prompt_tokens": 8, "total_tokens": 8}
}

The array holds 2048 values. Three are shown.

The response uses the OpenAI list format. data

holds one object per input text, each with an embedding

array and its index

. usage.prompt_tokens

counts the input tokens. total_tokens

equals it, since embedding produces no output tokens.

input

also accepts a list of strings, one request for the whole batch.

The full request body:

Field Required Description
input
yes One string or a list of strings
model
no Echoed back in the response. The server always serves the model set by MODEL_ID .
task
no
text-matching (default), retrieval , or code
prompt_name
no
query or passage . Used only when task is retrieval , defaults to passage .
encoding_format
no Accepted for OpenAI compatibility and ignored. Vectors are always float32 arrays.

The model produces a different embedding for the same text depending on the value of task

. The default, text-matching

, is for comparing two texts of the same kind, such as two support tickets or two product descriptions.

retrieval

is for search, where a short query is matched against longer documents. Embed the documents with prompt_name

set to passage

and the query with prompt_name

set to query

:

curl http://localhost:8081/v1/embeddings \
  -H "Authorization: Bearer your-api-key-here" \
  -H "Content-Type: application/json" \
  -d '{
        "input": ["when does the train leave"],
        "task": "retrieval",
        "prompt_name": "query"
      }'

code

is for source code and code search.

from openai import OpenAI

client = OpenAI(api_key="your-api-key-here", base_url="http://localhost:8081/v1")

resp = client.embeddings.create(
    model="jinaai/jina-embeddings-v4",
    input=["how long does the journey take"],
)
print(len(resp.data[0].embedding))
2048

The task

and prompt_name

fields are not OpenAI parameters, so the SDK passes them through extra_body

:

resp = client.embeddings.create(
    model="jinaai/jina-embeddings-v4",
    input=["SELECT id, name FROM users WHERE active = 1"],
    extra_body={"task": "code"},
)

If the server runs without API_KEY

, the OpenAI SDK still rejects an empty api_key

string. Pass any non-empty placeholder.

Environment variables on the jina

service:

Variable Default Purpose
MODEL_ID
jinaai/jina-embeddings-v4
Hugging Face model id. Override only for a fork or finetune with the same architecture.
HF_HOME
/root/.cache/huggingface
Cache path inside the container. The compose file mounts the jina-cache volume there, so a new container reuses the downloaded weights.
API_KEY
unset (optional) Bearer token for /v1/embeddings . If unset, the endpoint accepts requests without a token.

The compose file maps host port 8081 to port 80 in the container. If another service already listens on 8081, change the first number in 8081:80

.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @jina-embeddings-v4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jina-embeddings-v4-a…] indexed:0 read:3min 2026-08-24 ·