Anthropic Messages API
The Anthropic Messages endpoint exposes the Anthropic Messages API through Neon AI Gateway. Use it when you need extended thinking or prompt caching, which require the native Anthropic SDK. For standa...
The Anthropic Messages endpoint exposes the Anthropic Messages API through Neon AI Gateway. Use it when you need extended thinking or prompt caching, which require the native Anthropic SDK. For standard completions, the chat completions endpoint works with all Anthropic models and doesn't require the Anthropic SDK.
Base URL: https://<branch-host>/anthropic
This endpoint is also reachable at the longer /ai-gateway/anthropic/v1/messages path. Both behave identically and neither is deprecated. See Shorter paths for the full list of aliases.
Set these environment variables. See Get started for how to obtain them.
NEON_AI_GATEWAY_TOKEN=nt_live_...
NEON_AI_GATEWAY_BASE_URL=https://br-winter-pond-aptw82ef-api.ai.c-2.us-east-2.aws.neon.techSupported models
Section titled “Supported models”This endpoint accepts Anthropic models only. See the AI Gateway catalog for the full list. Supported models:
claude-opus-5,claude-sonnet-5,claude-fable-5claude-opus-4-8,claude-opus-4-7,claude-opus-4-6,claude-opus-4-5,claude-opus-4-1claude-sonnet-4-6,claude-sonnet-4-5claude-haiku-4-5
Sending a non-Anthropic model ID returns 400 model "<model-id>" is not available on the anthropic_messages endpoint, naming whichever model you sent. Use the chat completions endpoint if you need to call multiple providers from the same code.
Basic request
Section titled “Basic request”import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic({
authToken: process.env.NEON_AI_GATEWAY_TOKEN,
baseURL: `${process.env.NEON_AI_GATEWAY_BASE_URL}/anthropic`,
});
const message = await client.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 1024,
messages: [{ role: 'user', content: 'What is Neon?' }],
});
console.log(message.content[0].text);import anthropic
import os
client = anthropic.Anthropic(
auth_token=os.environ['NEON_AI_GATEWAY_TOKEN'],
base_url=f"{os.environ['NEON_AI_GATEWAY_BASE_URL']}/anthropic",
)
message = client.messages.create(
model='claude-sonnet-4-6',
max_tokens=1024,
messages=[{'role': 'user', 'content': 'What is Neon?'}],
)
print(message.content[0].text)curl -X POST "$NEON_AI_GATEWAY_BASE_URL/anthropic/v1/messages" \
-H "Authorization: Bearer $NEON_AI_GATEWAY_TOKEN" \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "What is Neon?"}]
}'Streaming
Section titled “Streaming”Streaming works the same as with the Anthropic SDK directly. Use client.messages.stream() or pass "stream": true in a cURL request. The only change from standard usage is base_url.
Prompt caching
Section titled “Prompt caching”The gateway forwards the cache_control field to Anthropic unchanged. Prompt caching works exactly as described in the Anthropic prompt caching docs.
const message = await client.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 1024,
system: [
{ type: 'text', text: 'You are a helpful assistant.' },
{
type: 'text',
text: longDocumentContent,
cache_control: { type: 'ephemeral' },
},
],
messages: [{ role: 'user', content: 'Summarize the key points.' }],
});
console.log(message.usage);
// { input_tokens: 50, output_tokens: 200,
// cache_creation_input_tokens: 10000, cache_read_input_tokens: 0 }message = client.messages.create(
model='claude-sonnet-4-6',
max_tokens=1024,
system=[
{'type': 'text', 'text': 'You are a helpful assistant.'},
{
'type': 'text',
'text': long_document_content,
'cache_control': {'type': 'ephemeral'},
},
],
messages=[{'role': 'user', 'content': 'Summarize the key points.'}],
)
print(message.usage)
# input_tokens=50, output_tokens=200,
# cache_creation_input_tokens=10000, cache_read_input_tokens=0Extended thinking
Section titled “Extended thinking”The gateway forwards the thinking parameter to Anthropic unchanged. Set budget_tokens to control how many tokens Claude can use for thinking. max_tokens must be greater than budget_tokens.
const message = await client.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 16000,
thinking: {
type: 'enabled',
budget_tokens: 10000,
},
messages: [{ role: 'user', content: 'Design a database schema for a multi-tenant SaaS app.' }],
});
for (const block of message.content) {
if (block.type === 'thinking') {
console.log('Thinking:', block.thinking);
} else if (block.type === 'text') {
console.log(block.text);
}
}message = client.messages.create(
model='claude-sonnet-4-6',
max_tokens=16000,
thinking={
'type': 'enabled',
'budget_tokens': 10000,
},
messages=[{'role': 'user', 'content': 'Design a database schema for a multi-tenant SaaS app.'}],
)
for block in message.content:
if block.type == 'thinking':
print('Thinking:', block.thinking)
elif block.type == 'text':
print(block.text)Forwarded headers
Section titled “Forwarded headers”The gateway forwards these request headers to the upstream provider: Accept, Anthropic-Beta, Anthropic-Version, Content-Type, User-Agent.
All other headers are stripped. The Authorization header is replaced with the workspace credential before forwarding. Your NEON_AI_GATEWAY_TOKEN is never sent to Anthropic directly.
Error handling
Section titled “Error handling”| Status | Message | Cause |
|---|---|---|
400 Bad Request |
unknown model "<model-id>" |
Model ID not in the catalog |
400 Bad Request |
model "<model-id>" is not available on the anthropic_messages endpoint |
Non-Anthropic model sent to this endpoint |
For authentication, quota, and upstream errors, see Troubleshooting.
Next steps
Section titled “Next steps”- Models: full model catalog
- Chat completions: use any model including Anthropic via the unified endpoint
- Authentication: credential scopes and branch binding
Need help?
Section titled “Need help?”Join our Discord Server to ask questions or see what others are doing with Neon. For paid plan support options, see Support.