SDKs
Calling the gateway from code -- the first-party TypeScript provider, and how every other language points its existing vendor SDK at the same endpoint.
SDKs
The gateway speaks the OpenAI Chat Completions wire, so every language
already has a client for it -- point your existing vendor SDK at your
gateway's base URL and give it an issued sk-dz- key.
TypeScript additionally has a first-party provider,
@nautir/ai-sdk-provider,
which plugs into the Vercel AI SDK and makes the gateway's
own controls -- routing, sessions, cost attribution, prompt caching -- typed
options rather than headers you assemble by hand.
You need a base URL and a key before any of this. Both come from the API Keys screen; see API access.
TypeScript
Install
npm install @nautir/ai-sdk-provider aiMake a request
import { createNautir } from "@nautir/ai-sdk-provider";
import { generateText } from "ai";
// Reads NAUTIR_GATEWAY_URL and NAUTIR_API_KEY from the environment.
const nautir = createNautir();
const { text } = await generateText({
model: nautir("anthropic/claude-haiku-4.5"),
prompt: "Say hello from the Nautir gateway.",
});
console.log(text);Model ids use the gateway's author/slug vocabulary. A native vendor
spelling such as gpt-4o is not resolvable -- see
Product surfaces.
Configuration
| Option | Environment variable | Notes |
|---|---|---|
baseURL | NAUTIR_GATEWAY_URL | Your own gateway deployment. |
apiKey | NAUTIR_API_KEY | An issued key, beginning sk-dz-. |
const nautir = createNautir({
baseURL: "https://gateway.internal.example.com",
apiKey: process.env.NAUTIR_API_KEY,
});There is deliberately no default base URL. Nothing should silently point at an endpoint that is not yours, so the provider requires one -- from the option or the environment variable.
One wire
Every model rides Chat Completions (/v1/chat/completions), Claude
included. Which upstream actually serves is a routing decision the gateway
makes, and it can move per request -- fallback chains, auto mode, a
X-Dz-Router override -- so the client cannot know it in advance and does not
try to.
Gateway controls
Gateway controls ride providerOptions.nautir and leave as X-Dz-* request
headers, which are their canonical carrier -- never body fields.
await generateText({
model: nautir("anthropic/claude-sonnet-4-5"),
prompt: "Hello!",
providerOptions: {
nautir: {
// Routing override (X-Dz-Router), bounded by your team's Routing Profile.
router: { provider_sort: "cost", fallback_models: ["openai/gpt-4o"] },
// Conversation pinning (X-Dz-Session-Id).
sessionId: "conv-42",
// Cost attribution (X-Dz-User/Team/Project/Env/Labels).
user: "alice",
team: "ml-platform",
project: "search",
env: "prod",
labels: { feature: "summarize" },
// Service identity (X-Dz-Workload / X-Dz-Namespace).
workload: "checkout-api",
namespace: "checkout",
workloadType: "service",
},
},
});Each of these is the typed form of a header documented in Request headers. Two are worth calling out:
sessionIdis the cheapest single thing you can send. It is what makes the routing pin, cache-miss classification and holdout stickiness work across turns.workloadandnamespacemust both be set. Workload attribution keys on the(namespace, workload)pair and skips a row missing either half, soworkloadalone records a name and resolves no attribution.
Prompt caching
Mark a cacheable prefix with a breakpoint and the provider puts cache_control
on that content part:
await generateText({
model: nautir("anthropic/claude-haiku-4.5"),
messages: [
{
role: "user",
content: [
{
type: "text",
text: bigStableContext,
providerOptions: { nautir: { cacheControl: { type: "ephemeral" } } },
},
{ type: "text", text: question },
],
},
],
});Two things make caching silently do nothing -- no error is raised in either case.
The prefix must clear the model's minimum. That minimum is not monotonic across generations, so it cannot be guessed from how new a model is: a 3K-token prefix caches on Opus 4.8 and does not on Haiku 4.5. Ask instead of hardcoding:
import { minimumCacheableTokens } from "@nautir/ai-sdk-provider";
minimumCacheableTokens("anthropic/claude-haiku-4.5"); // 4096The marked bytes must be identical across calls. A prefix that varies per request is written to the cache and never read.
Reads and writes are reported separately on usage.inputTokenDetails, as
cacheReadTokens and cacheWriteTokens, because they are priced separately --
a write costs more than an ordinary input token and a read a fraction of one,
so a single "cached" number cannot be turned back into money.
For apps already written against @ai-sdk/anthropic, four spellings of the
breakpoint are accepted unchanged: nautir.cacheControl,
nautir.cache_control, anthropic.cacheControl, anthropic.cache_control.
Python and other languages
Use the official OpenAI or Anthropic SDK and set its base URL to your gateway:
from openai import OpenAI
client = OpenAI(
base_url=f"{os.environ['NAUTIR_GATEWAY_URL']}/v1",
api_key=os.environ["NAUTIR_API_KEY"],
)
response = client.chat.completions.create(
model="anthropic/claude-haiku-4.5",
messages=[{"role": "user", "content": "Hello"}],
)Gateway controls are plain X-Dz-* request headers on this path -- see
Request headers. The API Keys screen carries
ready-to-run snippets with your own base URL already filled in.
Errors
| Code | Cause |
|---|---|
invalid_model_id | The model id is not author/slug. Native vendor spellings are not resolvable through the gateway. |
dz_router_conflict | A routing override arrived on both the header and a dz_router body field. The provider only ever sends the header; remove any dz_router you inject yourself. |
dz_router_too_large | The override is past the 8KB header cap. Move that configuration into a Routing Profile. |
Upgrading from 0.1.x
0.2.0 renames the SDK's own names from DevZero to Nautir, with no aliases:
createDevzero() becomes createNautir(), providerOptions.devzero becomes
providerOptions.nautir, and DEVZERO_GATEWAY_URL / DEVZERO_API_KEY become
NAUTIR_GATEWAY_URL / NAUTIR_API_KEY.
A leftover providerOptions.devzero throws rather than being ignored. An
ignored namespace would drop your routing and attribution headers and your cache
breakpoints while every request still succeeded -- a silent downgrade is worse
than a loud failure.
The gateway wire is unchanged. Requests still carry X-Dz-* headers and
issued keys still begin sk-dz-, so a 0.2.0 client talks to the same gateways
as 0.1.x.