AI Gateway models
Neon AI Gateway serves models hosted by Databricks. Use short model IDs in the model field, for example gpt 5 mini or gemini 3 flash. The databricks prefixed form is also accepted. The Neon Console an...
Neon AI Gateway serves models hosted by Databricks. Use short model IDs in the model field, for example gpt-5-mini or gemini-3-flash. The databricks- prefixed form is also accepted. The Neon Console and most examples use the short form.
Model availability may vary by region, and the catalog expands over time, so check back for new additions.
The full catalog is served as JSON at neon.com/models.json, the machine-readable source of truth, and mirrored as the neon provider on models.dev.
Model access
Section titled “Model access”Neon AI Gateway gives you one credential for both open-weight and foundation models. The catalog grows continuously as new models roll out, so the table below is always the source of truth for what you can call today.
Using the AI Gateway requires a paid plan with prepaid credits, which gives you the open-weight models. Foundation models are rolled out gradually. See Model access for what's included and how to request access to foundation models.
Available models
Section titled “Available models”Browse the full catalog below. Switch between the Text, Image, and Embeddings tabs, filter by provider or open weights, sort any column, and click a model for a copy-paste quickstart. Chat and image models get AI SDK, Mastra, Python, TypeScript, and cURL; embedding models get Python, TypeScript, and cURL. The endpoint each snippet targets is baked into its base URL: /v1 for chat completions and embeddings, /openai/v1 for the Responses API (image generation).
Open weights only
| Inputs | ||||||||
|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | text, image, pdf | 1.1M | Sep 2026 | $10 | $50 | — | |
| GLM-5.3 Flash | Zhipu AI | text, image | 1M | Aug 2026 | $0.15 | $0.50 | Open weights | |
| Grok 4.6 | xAI | text, image | 500K | Aug 2026 | $2 | $6 | — | |
| Gemini 3.5 Flash Lite | text, image, video, audio, pdf | 1M | Jul 2026 | $0.30 | $2.50 | — | ||
| Gemini 3.6 Flash | text, image, video, audio, pdf | 1M | Jul 2026 | $1.50 | $7.50 | — | ||
| Kimi K3 | Moonshot AI | text, image, video | 1M | Jul 2026 | $3 | $15 | Open weights | |
| Inkling | Thinking Machines | text, image, audio | 1M | Jul 2026 | $1 | $4.05 | Open weights | |
| GPT-5.6 Luna | OpenAI | text, image, pdf | 1.1M | Jul 2026 | $0.20 | $1.20 | — | |
| GPT-5.6 Sol | OpenAI | text, image, pdf | 1.1M | Jul 2026 | $5 | $30 | — | |
| GPT-5.6 Terra | OpenAI | text, image, pdf | 1.1M | Jul 2026 | $2 | $12 | — | |
| GLM-5.2 | Zhipu AI | text | 1M | Jun 2026 | $1.40 | $4.40 | Open weights | |
| Gemini 3.5 Flash | text, image, video, audio, pdf | 1M | May 2026 | $1.50 | $9 | — | ||
| GPT-5.5 | OpenAI | text, image, pdf | 1.1M | Apr 2026 | $5 | $30 | — | |
| GPT-5.5 Pro | OpenAI | text, image, pdf | 1.1M | Apr 2026 | $30 | $180 | — | |
| GPT-5.4 mini | OpenAI | text, image | 400K | Mar 2026 | $0.75 | $4.50 | — | |
| GPT-5.4 nano | OpenAI | text, image | 400K | Mar 2026 | $0.20 | $1.25 | — | |
| GPT-5.4 | OpenAI | text, image, pdf | 1.1M | Mar 2026 | $2.50 | $15 | — | |
| Gemini 3.1 Flash Lite Preview | text, image, video, audio, pdf | 1M | Mar 2026 | $0.25 | $1.50 | — | ||
| Qwen3.5 122B-A10B | Alibaba | text | 262K | Feb 2026 | $0.22 | $2.20 | Open weights | |
| Gemini 3.1 Pro Preview Custom Tools | text, image, video, audio, pdf | 1M | Feb 2026 | $2 | $12 | — | ||
| GPT-5.3 Codex | OpenAI | text, image, pdf | 400K | Feb 2026 | $1.75 | $14 | — | |
| Gemini 3 Flash Preview | text, image, video, audio, pdf | 1M | Dec 2025 | $0.50 | $3 | — | ||
| GPT-5.2 | OpenAI | text, image | 400K | Dec 2025 | $1.75 | $14 | — | |
| GPT-5.1 | OpenAI | text, image | 400K | Nov 2025 | $1.25 | $10 | — | |
| Qwen3-Next 80B-A3B Instruct | Alibaba | text | 131K | Sep 2025 | $0.15 | $1.20 | Open weights | |
| GPT-5 | OpenAI | text, image | 400K | Aug 2025 | $1.25 | $10 | — | |
| GPT-5 Mini | OpenAI | text, image | 400K | Aug 2025 | $0.25 | $2 | — | |
| GPT-5 Nano | OpenAI | text, image | 400K | Aug 2025 | $0.05 | $0.40 | — | |
| GPT OSS 120B | OpenAI | text | 131K | Aug 2025 | $0.15 | $0.60 | Open weights | |
| GPT OSS 20B | OpenAI | text | 131K | Aug 2025 | $0.07 | $0.30 | Open weights | |
| Qwen3 Embedding 0.6B | Alibaba | — | — | Jun 2025 | $0.02 | — | Open weights | |
| Llama 4 Maverick 17B Instruct | Meta | text, image | 1M | Apr 2025 | $0.50 | $1.50 | Open weights | |
| Gemma 3 12B | text, image | 131K | Mar 2025 | $0.15 | $0.50 | Open weights | ||
| Llama-3.3-70B-Instruct | Meta | text | 128K | Dec 2024 | $0.50 | $1.50 | Open weights | |
| Llama 3.1 8B Instruct | Meta | text | 131K | Jul 2024 | $0.15 | $0.45 | Open weights | |
| GTE Large EN | Alibaba | — | — | Jul 2023 | $0.13 | — | Open weights |
Prices are provider list prices per million tokens. Inference is free during the private preview. Click a model for a copy-paste quickstart.
For full request paths and when to prefer each endpoint, see Which endpoint to use. For embedding models (dimensions, normalization, and choosing a distance operator), see Embeddings.
Rate limits
Section titled “Rate limits”The following limit applies per account:
| Limit | Value |
|---|---|
| Tokens per minute (TPM) | 200,000 |
If you hit the limit, you'll receive a 429 Too Many Requests response with a message like ai gateway per-minute token limit exceeded for model "<model-id>". Requests resume when the rate limit window resets.
The TPM limit is counted against total tokens (input and output combined), not input alone. Upstream output token limits (20,000 OTPM for most models) apply independently, so you can hit a 429 on output tokens without reaching the gateway's TPM limit. See Databricks Foundation Model API limits for details.
The 200,000 TPM ceiling is a soft limit. If you need a higher limit, contact Support.
A separate account-level daily spend cap also applies and can block AI Gateway requests with a 429 / REQUEST_LIMIT_EXCEEDED. It isn't a fixed published number and can vary by account. See Pricing for details, or Troubleshooting if you hit it.
Pricing
Section titled “Pricing”See AI Gateway pricing for details.
Independent of billing, Neon enforces an account-level daily spend cap on AI Gateway usage, separate from the per-minute rate limits above. If your account exceeds it, every AI Gateway endpoint returns 429 Too Many Requests with error code REQUEST_LIMIT_EXCEEDED until the cap resets or the block is lifted. Neon hasn't published a fixed cap value; it isn't a flat number and can vary by account. See Troubleshooting if you hit this.
Which endpoint to use
Section titled “Which endpoint to use”Most models work with the Chat completions endpoint. It is the recommended starting point and works with all providers. Use a provider-specific endpoint when required:
All paths below are appended to your branch's bare AI Gateway host (NEON_AI_GATEWAY_BASE_URL).
| Provider | Recommended endpoint | Notes |
|---|---|---|
| OpenAI (most models) | /v1/chat/completions |
Use /openai/v1/responses for Responses API features |
OpenAI (gpt-5-3-codex, gpt-5-5-pro) |
/openai/v1/responses |
These models require the Responses API and don't work with chat/completions |
| Google Gemini | /v1/chat/completions |
Use /gemini/v1beta/models/{model}:generateContent with the google-genai SDK |
| Google Gemma 3 12B | /v1/chat/completions |
Chat completions only. Doesn't support the Gemini SDK endpoint |
| Meta, Zhipu AI, Thinking Machines, Moonshot AI | /v1/chat/completions |
Chat completions only |
| Alibaba (chat models) | /v1/chat/completions |
Chat completions only |
| Alibaba (embedding models) | /v1/embeddings |
No chat completions — returns a vector, not text |
Shorter paths
Section titled “Shorter paths”Each inference dialect is reachable at two equivalent paths: a shorter top-level path (recommended, and what most examples and the @neon/ai-sdk-provider use) and a longer /ai-gateway/<dialect>/v1 path. Both forms behave identically, using the same branch host, bearer token, request body, response body, model routing, rate limits, and quota, and neither is deprecated. The longer /ai-gateway/... paths keep working indefinitely.
The shorter form isn't a uniform /v1/<dialect> rule. The unified chat completions endpoint is a bare /v1/chat/completions, matching the OpenAI and OpenRouter convention. The native dialects are prefixed by provider instead, and each keeps its own upstream version segment so the path matches what that provider's SDK expects: /openai/v1/... and /gemini/v1beta/....
Use the shorter paths when you want OpenAI/OpenRouter-style URLs. Use the /ai-gateway/... paths when a framework or existing Neon example expects the older dialect-specific route.
| Shorter path | Equivalent to |
|---|---|
POST /v1/chat/completions |
/ai-gateway/mlflow/v1/chat/completions |
POST /openai/v1/responses |
/ai-gateway/openai/v1/responses |
POST /gemini/v1beta/models/{model}:generateContent |
/ai-gateway/gemini/v1beta/models/{model}:generateContent |
List available models
Section titled “List available models”GET /v1/models lists the model catalog in an OpenRouter-shaped response, authenticated the same way as the endpoints above. Unlike the inference dialects, the model list has only this /v1/models path, with no /ai-gateway/... form.
curl "$NEON_AI_GATEWAY_BASE_URL/v1/models" \
-H "Authorization: Bearer $NEON_AI_GATEWAY_TOKEN"{
"object": "list",
"data": [
{
"id": "gpt-5-mini",
"canonical_slug": "gpt-5-mini",
"name": "GPT-5 Mini",
"object": "model",
"owned_by": "openai",
"created": 0,
"enabled": true,
"context_length": null,
"architecture": {
"modality": "text->text",
"input_modalities": ["text"],
"output_modalities": ["text"],
"tokenizer": "GPT",
"instruct_type": null
},
"top_provider": {
"is_moderated": false,
"context_length": null,
"max_completion_tokens": null
},
"pricing": null,
"per_request_limits": null
}
]
}The response returns one object per model. Key fields:
enabledis whether your account can call the model. Iffalse, a request returns a403(see Troubleshooting). Gated models are sometimes left out of the list entirely, so useenabled: trueas your check. See Model access for what determines access and how to request more models.id,name, andowned_byidentify the model. Useid(or itsdatabricks-prefixed form) in themodelfield of a request.canonical_slug,architecture, andtop_providerare OpenRouter-compatible descriptive fields.createdis always0, andpricing,per_request_limits, andcontext_lengthare currently alwaysnull. Use the tables earlier on this page for context windows and model details.
Check what your account can call
Section titled “Check what your account can call”To list only the models your account can call, filter on enabled:
curl "$NEON_AI_GATEWAY_BASE_URL/v1/models" \
-H "Authorization: Bearer $NEON_AI_GATEWAY_TOKEN" \
| jq -r '.data[] | select(.enabled) | .id'Provider terms
Section titled “Provider terms”Models are hosted by Databricks and served through Neon AI Gateway. You are responsible for complying with each provider's applicable terms of use.
| Provider | Terms |
|---|---|
| OpenAI | OpenAI Usage Policies |
| Google Gemini | Google Cloud Acceptable Use Policy · Google Generative AI Prohibited Use Policy |
| Google Gemma | Gemma Terms of Use · Gemma Prohibited Use Policy |
| Meta | Terms differ by Llama version. See the Notes column in the Meta models table. |
Need help?
Section titled “Need help?”Join our Discord Server to ask questions or see what others are doing with Neon. For paid plan support options, see Support.