Migrate from Vercel AI Gateway to NexLLM: update clients, authentication, model IDs, routing configuration, and observability.
Move your Vercel AI Gateway integration to NexLLM while keeping your application’s existing AI SDK, OpenAI-compatible, or Claude-native request format. This guide separates the connection changes from the gateway-specific behavior you need to review before production rollout.You can continue using the Vercel AI SDK and hosting your application on Vercel. This migration changes the inference gateway, not necessarily your application framework or deployment platform.
Migration Overview
See API Reference and Authentication.
Prerequisites
1
Create a NexLLM API key
Create a dedicated key for the migration and store it in your application’s server-side secrets.Review its expiration, quota, model restrictions, IP allowlist, and group.See API Keys.
2
Confirm model access and endpoint compatibility
List the models available to your key. Check the selected model’s details for the endpoint and capabilities your application needs.See Models API and Model Details.
3
Inventory your existing integration
Identify client initialization, gateway credentials, model strings, provider options, custom headers, response metadata, and telemetry.Include background jobs, server routes, scheduled tasks, and local scripts.
4
Prepare acceptance tests and rollback
Preserve your current configuration until the migrated integration passes representative tests. Define rollback criteria before changing production traffic.
Quick Start for Claude Code Users
You can give Claude Code the following prompt from your repository. This is a suggested migration prompt, not an installable NexLLM migration skill.Step 1: Update Your Environment Variables
Replace the credential used by the migrated inference calls:NEXLLM_MODEL and NEXLLM_CLAUDE_MODEL are application configuration variables used in this guide, not special API parameters.
Replace the Gateway Authentication Fallback
Change application logic such as:Step 2: Update Your Client
Choose the example that matches your current request format. The examples use environment variables for model selection so that your application can use the exact ID verified in Step 5.Vercel AI SDK
Keep the AI SDK, but replace implicit gateway model resolution with an explicit NexLLM-backed provider. Install a version of@ai-sdk/openai compatible with your project’s ai package. Avoid combining the gateway migration with an unrelated SDK major-version upgrade.
streamText, shared provider factories, and provider registries. Updating one client does not migrate independently configured clients.
OpenAI SDK and HTTP Clients
For an existing OpenAI client, replace:- Python
- JavaScript
- cURL
Anthropic SDK and Claude-Native Requests
You do not need to convert a Claude-native integration to Chat Completions solely to use NexLLM. NexLLM documents:- Python — Anthropic SDK
- cURL
Step 3: Review Vercel-Specific Headers and Metadata
Remove gateway-specific configuration from the NexLLM client after deciding how to preserve any behavior that depended on it. Do not indiscriminately delete application tracing headers or Vercel deployment configuration used elsewhere.Request headers and authentication
Request headers and authentication
See Authentication.
Response metadata and headers
Response metadata and headers
See Chat Completions and Usage & Logs.
Step 4: Decompose Your Gateway Options
ReviewproviderOptions.gateway and any related request fields before removing them.
A connection change alone does not establish equivalent routing, caching, privacy, or timeout behavior.
The table below describes migration actions, not guaranteed one-to-one API replacements.“Verify separately” means this guide does not establish an equivalent NexLLM control from the referenced documentation. It does not necessarily mean the capability is unavailable.
For NexLLM-specific access and pricing configuration, see Channel Groups and Pricing.
Use Channel Groups Deliberately
Each NexLLM key belongs to one channel group. The group controls model access and the pricing ratio applied to usage. Treat this as a key-level configuration decision, not a direct replacement for every per-request routing rule. If separate workloads need different configurations, consider dedicated keys and select the appropriate key in trusted server-side code. See Channel Groups.Preserve Required Constraints
Before migrating a request that has provider or privacy restrictions:- Record the original requirement.
- Identify how it will be enforced after migration.
- Add a test or operational check.
- Keep the old route until the replacement is verified.
Test Fallbacks Without Duplicating Work
If your application implements fallback:- Distinguish retryable failures from invalid requests or credentials.
- Limit attempts and overall elapsed time.
- Avoid restarting after partial streamed output without an explicit policy.
- Prevent duplicate tool execution.
- Record which model was attempted and which one produced the accepted result.
Step 5: Update Model Identifiers
Retrieve the model IDs available to the key that will serve production traffic:data[].id value in inference requests.
See Models API.
Example Model Mapping
These are candidates to verify, not unconditional replacements.
Keep model-version changes separate from the gateway migration where practical. If the original model is unavailable, evaluate the substitute explicitly.
The Models API establishes availability to your key. Check model details separately for endpoint compatibility and required capabilities.
See Model Details.
Step 6: Reconnect Observability
NexLLM Usage Logs include request time, key name, model, timing, token consumption, and cost or quota deducted. See Usage & Logs.
See API Keys and Wallet.
Keep Application Correlation IDs
For each inference operation, consider recording:- Your application request ID.
- Environment and application version.
- Requested model.
- Returned completion ID, where applicable.
- Duration and outcome.
- Retry or fallback attempt.
- Token usage when returned.
Preserve Your AI Gateway History
Before reducing access to the previous gateway:- Identify historical records required for operations or audits.
- Check the export or retention mechanisms available to your account.
- Preserve available logs, billing records, and routing configuration.
- Verify that archived records remain readable.
- Document the cutover time.
Step 7 (Optional): Evaluate the Responses API
NexLLM lists an OpenAI Responses-compatible endpoint:Minimal Compatibility Test
The following is a minimal Responses-format request pattern. Replace the placeholder with a compatible model available to your key:.responses(modelId) only as a deliberate endpoint choice. Retest response parsing and streaming behavior.
Why Migrate to NexLLM
Keep an OpenAI-Compatible Integration
Use the documented Chat Completions interface while accessing GPT, Claude, and Gemini models. See Quickstart.Choose Model Access and Pricing Through Channel Groups
Select a group appropriate for your workload rather than assuming the previous gateway’s routing policy transfers unchanged. See Channel Groups.Configure Purpose-Specific Credentials
Separate development, production, or workload credentials using NexLLM’s key controls. See API Keys.Review Usage in One Dashboard
Use request logs and consumption statistics to evaluate the migration. See Usage & Logs.Validate the benefits against your own acceptance criteria. This guide does not promise identical routing, automatic fallback, BYOK terms, retention policies, service tiers, or latency.
Troubleshooting
The model is not found
The model is not found
Call
GET /v1/models using the application’s actual key.Check the exact identifier, the assigned channel group, and any model restrictions. Do not assume the previous gateway’s prefix is accepted.See Models API.Authentication fails
Authentication fails
Confirm the application sends a NexLLM key rather than an AI Gateway key, OIDC token, or upstream provider credential.Check secret availability in the running process, configured expiration, IP restrictions, and accidental whitespace.Use bearer authentication for the OpenAI-compatible examples and native headers for the Claude Messages example.See Authentication.
AI SDK requests still go through Vercel AI Gateway
AI SDK requests still go through Vercel AI Gateway
Inspect every model construction site.Replace implicit gateway strings and
gateway(...) instances with the configured NexLLM provider where migration is intended.Check shared registries, background jobs, streaming routes, and default provider configuration. Preserve intentional gateway usage elsewhere.The AI SDK calls /responses unexpectedly
The AI SDK calls /responses unexpectedly
For Chat Completions, use:Do not rely on the default OpenAI provider factory to select your intended endpoint.
Claude-native requests return 404
Claude-native requests return 404
Inspect the final URL.The complete Messages endpoint is:For the Anthropic SDK example in this guide, use the host-only base URL:Check for duplicated
/v1 segments, missing path segments, or a client still pointing to the previous gateway.Requests succeed but do not appear in NexLLM logs
Requests succeed but do not appear in NexLLM logs
Verify the actual destination, credential, account, and log time range.Check independently initialized clients and background workers. Avoid printing authorization headers while debugging.See Usage & Logs.
Provider routing or fallback no longer applies
Provider routing or fallback no longer applies
Review the removed gateway options and the requirements behind them.A channel-group selection is not proof that a per-request provider allowlist, ordering rule, or fallback chain has been reproduced.Restore affected traffic to the previous route until required constraints are verified.
Cache behavior or cost changed
Cache behavior or cost changed
Compare model IDs, request formats, prompt structure, cache behavior, token accounting, and the applicable group ratio.Do not assume the original gateway’s automatic cache configuration transfers to NexLLM.See Pricing.
Connection checks fail
Connection checks fail
Test the documented Models endpoint first:Then run a small inference request.A successful model-list request does not validate every inference feature. Do not invent a
/health endpoint based on another gateway’s documentation.Verification Checklist
Before completing the cutover:- Production clients use the intended NexLLM endpoint.
- Credentials remain server-side.
- Exact model IDs are verified using the production key.
- Channel-group and key restrictions are reviewed.
- Plain-text requests pass.
- Streaming passes, including cancellation and partial failures.
- Tool calls pass without duplicate execution.
- Structured outputs meet the original validation requirements.
- Required multimodal inputs work.
- Routing, retries, timeouts, and fallbacks are explicitly tested.
- Privacy and provider restrictions are preserved.
- Logging and consumption monitoring work.
- Representative quality, latency, and cost checks pass.
- Historical records remain accessible.
- A rollback procedure is tested.
- Old credentials are removed only after remaining dependencies are checked.
Next Steps
API Reference
Review supported endpoint formats.
Available Models
Discover model IDs available to your key.
Channel Groups
Review access and pricing configuration.
Usage & Logs
Inspect requests after migration.
Feedback
When reporting migration issues, include:- The client package and version.
- The endpoint and model ID.
- The behavior expected from the previous integration.
- A sanitized request example.
- The HTTP status and error body.
- The approximate request time.
- A completion or request identifier, if available.