# Generate embeddings with Neon AI Gateway

> Source: <https://neon.com/blog/generate-embeddings-with-neon-ai-gateway>
> Published: 2026-10-02 12:00:00+00:00

#### We're building backends

When a coding agent ships an app today, it deploys the [Neon backend](https://neon.com/blog/neon-backend-is-ga) - [Lakebase Postgres](https://neon.com/docs/postgres/overview) (our database) plus [Object Storage](https://neon.com/docs/storage/overview), [Functions](https://neon.com/docs/compute/functions/overview), [Managed Better Auth](https://neon.com/docs/auth/overview), and [AI Gateway](https://neon.com/docs/ai-gateway/overview). Every primitive branches with your data.

Neon now includes [AI Gateway](https://neon.com/blog/llms-belong-in-your-backend), our primitive for calling LLMs. You get frontier and open-weight models hosted by Databricks, billed through Neon at the labs' own prices (no markup).

**The latest addition to AI Gateway: it now [serves embedding models](https://neon.com/docs/ai-gateway/embeddings) on the same OpenAI-compatible endpoint and with the same credential you already use for chat.** Point your SDK at `/v1/embeddings`, write the vectors into [Lakebase Postgres](https://neon.com/lakebase), and query them with [Lakebase Search](https://neon.com/docs/ai/lakebase-search). The whole retrieval pipeline, from the uploaded file to the generated answer, runs inside one Neon branch.

Ask your agent to set it up:

Or set it up yourself with the OpenAI SDK:

## Quick intro on the primitives

The pipeline above involves three pieces of the Neon backend: [Lakebase Postgres](https://neon.com/docs/postgres/overview) (our database), [AI Gateway](https://neon.com/docs/ai-gateway/overview), and [Lakebase Search](https://neon.com/docs/ai/lakebase-search). In case you're new here, here's a quick primer on each.

### Lakebase Postgres

[Lakebase Postgres](https://neon.com/docs/postgres/overview) is the Neon database, the center of the Neon backend:

- It's 100% Postgres, but serverless: instant to provision, with autoscaling, and scale to zero, with [usage-based pricing](https://neon.com/pricing) and a[generous free plan](https://neon.com/pricing)
- It's built on the [lakebase architecture](https://neon.com/docs/introduction/architecture-overview) : compute is separated from versioned, copy-on-write storage
- Lakebase Postgres also branches: a [branch](https://neon.com/docs/introduction/branching) is an isolated copy of your database that's ready in seconds and doesn't duplicate storage - teams use it to automate all kinds of workflows that would otherwise involve a "dev instance". The rest of the Neon backend primitives also branch, following the database

### AI Gateway

[AI Gateway](https://neon.com/docs/ai-gateway/overview) lets you call models from Neon instead of collecting lab accounts:

- A Neon credential with the `ai_gateway:invoke` scope reaches the whole[model catalog](https://neon.com/docs/ai-gateway/models) . Switching models means changing a string
- Models are served on Databricks [Foundation Model APIs](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/) , and we pass through each lab's published per-token price
- Chat completions and embeddings sit on an OpenAI-compatible `/v1` path, so the OpenAI SDK works once you change the base URL and key
- Every branch gets its own gateway host, and `neon env pull` writes`NEON_AI_GATEWAY_TOKEN` and`NEON_AI_GATEWAY_BASE_URL` for the branch you're on. Inside a[Neon Function](https://neon.com/functions) , both are injected for you

### Lakebase Search

[Lakebase Search](https://neon.com/docs/ai/lakebase-search) is vector, keyword, and hybrid search inside Lakebase Postgres, delivered as two extensions:

- **`lakebase_vector`** adds the` lakebase_ann` index for vector similarity search. It uses the same`vector` types, distance operators, and query syntax as`pgvector` , so there's nothing to migrate, and a single index scales past 1 billion vectors.
- **`lakebase_text`** adds the` lakebase_bm25` index for keyword search, with real BM25 ranking and top-K pushdown on standard`tsvector` columns.

Both indexes live in the same Postgres as your app data, so one query can combine them, join your tables, and filter by tenant. **Lakebase Search is also built for scale to zero:** indexes live in storage, so they're ready after a cold start and available on every branch with no rebuild.

#### Leading on price-performance

We recently ran VectorDBBench on LAION-100M, and Lakebase Search led the tested systems on price-performance. See the results in [Lakebase Search: the retrieval primitive for agents on Neon](https://neon.com/blog/lakebase-search-retrieval-agents).

## What's new: embeddings on AI Gateway

AI Gateway now exposes `POST /v1/embeddings`. It accepts a single string or a batch of up to 150 strings in one request, and returns vectors in the standard OpenAI response shape. Two models are available at launch:

| Model | Dimensions | Normalized | Price (as of Oct 2026) | 
|---|---|---|---|
| `qwen3-embedding-0-6b` | 1024 (configurable) | Yes | $0.02 per 1M input tokens | 
| `gte-large-en` | 1024 | No, use cosine distance | $0.13 per 1M input tokens | 

#### Availability

AI Gateway is available on the Launch and Scale plans, paid with [prepaid credits](https://neon.com/docs/ai-gateway/prepaid-credits). If you'd like to try it for free, [tell us on Discord](https://discord.gg/92vNTzKDGp): we have credits to give.

#### Choosing a model

If you're unsure what to choose, start with `qwen3-embedding-0-6b`, since it's the cheapest option. Use `gte-large-en` if you already have vectors produced by it and need new embeddings to match.

The [Embeddings guide](https://neon.com/docs/ai-gateway/embeddings) covers the request and response fields, Python and curl examples, the `dimensions` parameter, and how to use the embedding models with the Vercel AI SDK through [`@neon/ai-sdk-provider`](https://www.npmjs.com/package/@neon/ai-sdk-provider).

## Build a complete retrieval pipeline on a Neon branch

Having embeddings in AI Gateway is useful on its own, but it gets more interesting when you zoom out and look at the whole pipeline they're part of.

If you think about the features teams are constantly shipping these days (a support bot that answers from your docs, search across the files your users upload, an agent with memory), under the hood, they all have a similar shape. Content comes in → gets turned into vectors → gets stored → gets searched when a question arrives → the best matches go to a model that writes the answer.

Now that AI Gateway serves embeddings, every step is backed by a Neon primitive:

1. **Ingest → Object Storage + Functions:** e.g. a user uploads a file to[Object Storage](https://neon.com/docs/storage/overview) , and a[Function Trigger](https://neon.com/docs/compute/functions/triggers/object-storage) invokes a[Neon Function](https://neon.com/docs/compute/functions/overview) to process it
2. **Embed → AI Gateway:** the Function splits the file into chunks and turns each one into a vector with AI Gateway's`/v1/embeddings` endpoint
3. **Store → Lakebase Postgres:** the chunks and their vectors go into Lakebase Postgres, next to the rest of your app data, so they can be joined and filtered like any other rows
4. **Retrieve → Lakebase Search:** when a question arrives, Lakebase Search finds the most relevant chunks with vector, keyword, or hybrid search
5. **Generate → AI Gateway:** the matches go to a chat model through the same gateway and credential, and the model writes the answer

To build something like this, hand a prompt like this to your agent:

### Branching is the connective tissue of this experience

Create a Neon branch and the whole pipeline follows: it reflects production exactly, it's ready immediately, and it stays lightweight, because none of your rows or files are duplicated.

For a retrieval pipeline, this makes it safe and easy to switch embedding models or vector sizes. On a branch, you can re-embed, run your evals against real production data, and keep the change only if it wins. The same loop works for all kinds of experiments: a new chunking strategy, a different hybrid-search weighting, or a different chat model.

## Try it

Generate, store, and search embeddings in one Neon backend. If you're building with a coding agent, start by [setting it up with Neon](https://neon.com/docs/get-started/with-an-agent): one command installs the Neon CLI, agent skills, and MCP server.

It's also good to point your agent to these docs:

- [Embeddings on AI Gateway](https://neon.com/docs/ai-gateway/embeddings) - for the endpoint, models, and request options
- [Lakebase Search quickstart](https://neon.com/docs/ai/lakebase-search-get-started) - for vector, keyword, and hybrid search
- [Trigger a function on object upload](https://neon.com/docs/compute/functions/triggers/object-storage) - to automate the ingest step
- [Tour the Neon backend](https://neon.com/docs/get-started/backend-overview) - for how the primitives fit together in one`neon.ts`

Every docs page is also available as markdown (add `.md` to the URL), and [llms.txt](https://neon.com/docs/llms.txt) indexes all of them, so your agent can read the exact reference.
