Skip to main content

Migrate from Vercel AI Gateway to NexLLM: update clients, authentication, model IDs, routing configuration, and observability.

Move your Vercel AI Gateway integration to NexLLM while keeping your application’s existing AI SDK, OpenAI-compatible, or Claude-native request format. This guide separates the connection changes from the gateway-specific behavior you need to review before production rollout.
You can continue using the Vercel AI SDK and hosting your application on Vercel. This migration changes the inference gateway, not necessarily your application framework or deployment platform.

Migration Overview

See API Reference and Authentication.

Prerequisites

1

Create a NexLLM API key

Create a dedicated key for the migration and store it in your application’s server-side secrets.Review its expiration, quota, model restrictions, IP allowlist, and group.See API Keys.
2

Confirm model access and endpoint compatibility

List the models available to your key. Check the selected model’s details for the endpoint and capabilities your application needs.See Models API and Model Details.
3

Inventory your existing integration

Identify client initialization, gateway credentials, model strings, provider options, custom headers, response metadata, and telemetry.Include background jobs, server routes, scheduled tasks, and local scripts.
4

Prepare acceptance tests and rollback

Preserve your current configuration until the migrated integration passes representative tests. Define rollback criteria before changing production traffic.

Quick Start for Claude Code Users

You can give Claude Code the following prompt from your repository. This is a suggested migration prompt, not an installable NexLLM migration skill.
Migrating application API calls does not automatically configure Claude Code itself to use NexLLM. Treat the coding agent’s connection settings as a separate integration.

Step 1: Update Your Environment Variables

Replace the credential used by the migrated inference calls:
For the examples below, choose a model available to your key:
NEXLLM_MODEL and NEXLLM_CLAUDE_MODEL are application configuration variables used in this guide, not special API parameters.

Replace the Gateway Authentication Fallback

Change application logic such as:
to:
Use the NexLLM key for NexLLM requests rather than forwarding a Vercel gateway key or OIDC token. For OpenAI-compatible requests, authentication uses:
See Authentication.
Do not remove VERCEL_OIDC_TOKEN, cloud credentials, or provider keys globally without checking their other uses. They may still be required by unrelated services or your rollback path.Keep NexLLM credentials server-side. Do not expose them through public environment variables or browser bundles.
Review the new key’s configured expiration rather than assuming it never expires. See API Keys.

Step 2: Update Your Client

Choose the example that matches your current request format. The examples use environment variables for model selection so that your application can use the exact ID verified in Step 5.

Vercel AI SDK

Keep the AI SDK, but replace implicit gateway model resolution with an explicit NexLLM-backed provider. Install a version of @ai-sdk/openai compatible with your project’s ai package. Avoid combining the gateway migration with an unrelated SDK major-version upgrade.
Before:
After:
Use nexllm.chat(modelId) when targeting Chat Completions.The OpenAI provider’s default model factory can select the Responses API. Do not replace .chat(modelId) with nexllm(modelId) without deliberately reviewing the endpoint change.
Check the AI SDK OpenAI provider documentation for client-library behavior and NexLLM’s Chat Completions reference for the gateway endpoint. Search all call sites, including streamText, shared provider factories, and provider registries. Updating one client does not migrate independently configured clients.

OpenAI SDK and HTTP Clients

For an existing OpenAI client, replace:
with:
Then update the key and model ID.
See Quickstart.

Anthropic SDK and Claude-Native Requests

You do not need to convert a Claude-native integration to Chat Completions solely to use NexLLM. NexLLM documents:
with these headers:
See Native Authentication.
An SDK base URL and a complete HTTP endpoint are not interchangeable.The Anthropic SDK example uses https://www.nexllm.ai because its Messages method supplies /v1/messages. The OpenAI SDK examples use https://www.nexllm.ai/v1.Confirm that the final request path contains exactly one /v1 segment.
Preserve your native content-block handling when keeping the Messages API. If you deliberately switch to Chat Completions, translate system instructions, tools, content blocks, streaming events, and response parsing separately.

Step 3: Review Vercel-Specific Headers and Metadata

Remove gateway-specific configuration from the NexLLM client after deciding how to preserve any behavior that depended on it. Do not indiscriminately delete application tracing headers or Vercel deployment configuration used elsewhere.
See Authentication.
See Chat Completions and Usage & Logs.

Step 4: Decompose Your Gateway Options

Review providerOptions.gateway and any related request fields before removing them. A connection change alone does not establish equivalent routing, caching, privacy, or timeout behavior.
The table below describes migration actions, not guaranteed one-to-one API replacements.“Verify separately” means this guide does not establish an equivalent NexLLM control from the referenced documentation. It does not necessarily mean the capability is unavailable.
For NexLLM-specific access and pricing configuration, see Channel Groups and Pricing.

Use Channel Groups Deliberately

Each NexLLM key belongs to one channel group. The group controls model access and the pricing ratio applied to usage. Treat this as a key-level configuration decision, not a direct replacement for every per-request routing rule. If separate workloads need different configurations, consider dedicated keys and select the appropriate key in trusted server-side code. See Channel Groups.

Preserve Required Constraints

Before migrating a request that has provider or privacy restrictions:
  1. Record the original requirement.
  2. Identify how it will be enforced after migration.
  3. Add a test or operational check.
  4. Keep the old route until the replacement is verified.
Do not silently relax provider allowlists, output schemas, or privacy constraints to make requests succeed.
Do not copy routing.model.fallbacks, routing.provider.fallbacks, routing.model.sort, or model: "auto" from another gateway’s migration guide unless NexLLM separately documents and supports that configuration for your use case.

Test Fallbacks Without Duplicating Work

If your application implements fallback:
  • Distinguish retryable failures from invalid requests or credentials.
  • Limit attempts and overall elapsed time.
  • Avoid restarting after partial streamed output without an explicit policy.
  • Prevent duplicate tool execution.
  • Record which model was attempted and which one produced the accepted result.

Step 5: Update Model Identifiers

Retrieve the model IDs available to the key that will serve production traffic:
Use the exact data[].id value in inference requests. See Models API.

Example Model Mapping

These are candidates to verify, not unconditional replacements.
Do not apply a universal “remove the provider prefix” rule.NexLLM documents both unprefixed and prefixed IDs. A mapping can also change the hosting provider, so matching the model family alone is insufficient.
Keep model-version changes separate from the gateway migration where practical. If the original model is unavailable, evaluate the substitute explicitly. The Models API establishes availability to your key. Check model details separately for endpoint compatibility and required capabilities. See Model Details.

Step 6: Reconnect Observability

NexLLM Usage Logs include request time, key name, model, timing, token consumption, and cost or quota deducted. See Usage & Logs. See API Keys and Wallet.

Keep Application Correlation IDs

For each inference operation, consider recording:
  • Your application request ID.
  • Environment and application version.
  • Requested model.
  • Returned completion ID, where applicable.
  • Duration and outcome.
  • Retry or fallback attempt.
  • Token usage when returned.
Do not log credentials. Apply your existing controls to prompt and response content.

Preserve Your AI Gateway History

Before reducing access to the previous gateway:
  1. Identify historical records required for operations or audits.
  2. Check the export or retention mechanisms available to your account.
  3. Preserve available logs, billing records, and routing configuration.
  4. Verify that archived records remain readable.
  5. Document the cutover time.
Do not assume historical Vercel records will automatically appear in NexLLM.

Step 7 (Optional): Evaluate the Responses API

NexLLM lists an OpenAI Responses-compatible endpoint:
See API Reference. If your application already uses Responses, evaluate that endpoint before converting the workflow to Chat Completions. Select a model whose details show Responses compatibility. The model browser supports filtering by endpoint type. See Model Details.

Minimal Compatibility Test

The following is a minimal Responses-format request pattern. Replace the placeholder with a compatible model available to your key:
The documented endpoint does not, by itself, establish feature parity with every gateway or provider.Verify response chaining, storage, reasoning, tools, structured output, multimodal inputs, and streaming for your selected model. Do not assume previous_response_id or provider-specific tools work identically.
For the AI SDK, select .responses(modelId) only as a deliberate endpoint choice. Retest response parsing and streaming behavior.

Why Migrate to NexLLM

Keep an OpenAI-Compatible Integration

Use the documented Chat Completions interface while accessing GPT, Claude, and Gemini models. See Quickstart.

Choose Model Access and Pricing Through Channel Groups

Select a group appropriate for your workload rather than assuming the previous gateway’s routing policy transfers unchanged. See Channel Groups.

Configure Purpose-Specific Credentials

Separate development, production, or workload credentials using NexLLM’s key controls. See API Keys.

Review Usage in One Dashboard

Use request logs and consumption statistics to evaluate the migration. See Usage & Logs.
Validate the benefits against your own acceptance criteria. This guide does not promise identical routing, automatic fallback, BYOK terms, retention policies, service tiers, or latency.

Troubleshooting

Call GET /v1/models using the application’s actual key.Check the exact identifier, the assigned channel group, and any model restrictions. Do not assume the previous gateway’s prefix is accepted.See Models API.
Confirm the application sends a NexLLM key rather than an AI Gateway key, OIDC token, or upstream provider credential.Check secret availability in the running process, configured expiration, IP restrictions, and accidental whitespace.Use bearer authentication for the OpenAI-compatible examples and native headers for the Claude Messages example.See Authentication.
Inspect every model construction site.Replace implicit gateway strings and gateway(...) instances with the configured NexLLM provider where migration is intended.Check shared registries, background jobs, streaming routes, and default provider configuration. Preserve intentional gateway usage elsewhere.
For Chat Completions, use:
Do not rely on the default OpenAI provider factory to select your intended endpoint.
Inspect the final URL.The complete Messages endpoint is:
For the Anthropic SDK example in this guide, use the host-only base URL:
Check for duplicated /v1 segments, missing path segments, or a client still pointing to the previous gateway.
Verify the actual destination, credential, account, and log time range.Check independently initialized clients and background workers. Avoid printing authorization headers while debugging.See Usage & Logs.
Review the removed gateway options and the requirements behind them.A channel-group selection is not proof that a per-request provider allowlist, ordering rule, or fallback chain has been reproduced.Restore affected traffic to the previous route until required constraints are verified.
Compare model IDs, request formats, prompt structure, cache behavior, token accounting, and the applicable group ratio.Do not assume the original gateway’s automatic cache configuration transfers to NexLLM.See Pricing.
Test the documented Models endpoint first:
Then run a small inference request.A successful model-list request does not validate every inference feature. Do not invent a /health endpoint based on another gateway’s documentation.

Verification Checklist

Before completing the cutover:
  • Production clients use the intended NexLLM endpoint.
  • Credentials remain server-side.
  • Exact model IDs are verified using the production key.
  • Channel-group and key restrictions are reviewed.
  • Plain-text requests pass.
  • Streaming passes, including cancellation and partial failures.
  • Tool calls pass without duplicate execution.
  • Structured outputs meet the original validation requirements.
  • Required multimodal inputs work.
  • Routing, retries, timeouts, and fallbacks are explicitly tested.
  • Privacy and provider restrictions are preserved.
  • Logging and consumption monitoring work.
  • Representative quality, latency, and cost checks pass.
  • Historical records remain accessible.
  • A rollback procedure is tested.
  • Old credentials are removed only after remaining dependencies are checked.

Next Steps

API Reference

Review supported endpoint formats.

Available Models

Discover model IDs available to your key.

Channel Groups

Review access and pricing configuration.

Usage & Logs

Inspect requests after migration.

Feedback

When reporting migration issues, include:
  • The client package and version.
  • The endpoint and model ID.
  • The behavior expected from the previous integration.
  • A sanitized request example.
  • The HTTP status and error body.
  • The approximate request time.
  • A completion or request identifier, if available.
Remove credentials, private prompts, and personal data before sharing diagnostics.