Migrate from Merge Gateway to NexLLM: update your API client, review routing policies, and preserve project controls and observability.
Move your Merge Gateway integration to NexLLM while retaining the request format that best fits your application. This guide covers client configuration, Merge-specific extensions, routing policies, model identifiers, project attribution, budgets, and observability.Migration overview
Choose the migration path that matches your existing integration.| Existing integration | NexLLM migration path |
|---|---|
Native merge-gateway-sdk | Replace the Merge client with the OpenAI SDK and evaluate the Responses endpoint. |
| OpenAI SDK | Update the base URL, API key, and model identifier. |
| Anthropic Messages integration | Keep the native Messages format and use NexLLM’s documented authentication. |
| Vercel AI SDK | Replace the Merge provider with a configured OpenAI provider and explicitly select the API format. |
| LangChain or another wrapper | Update its underlying client and verify the final request URL and payload. |
| Direct HTTP requests | Update the endpoint, authentication, model, and Merge-specific request extensions. |
| Purpose | Full endpoint |
|---|---|
| Chat Completions | https://www.nexllm.ai/v1/chat/completions |
| Responses | https://www.nexllm.ai/v1/responses |
| Claude Messages | https://www.nexllm.ai/v1/messages |
| Model discovery | https://www.nexllm.ai/v1/models |
The OpenAI SDK base URL is
https://www.nexllm.ai/v1. Endpoint paths shown above already include /v1; do not append it twice.Prerequisites
- Create a NexLLM API key. Use a dedicated migration key so you can isolate testing from existing production traffic. See API Keys.
- Check account funding. Review your available balance and top up if necessary. See Wallet.
- Choose a channel group. Confirm that it includes the models you need. See Channel Groups.
- Inventory the existing integration. Include client code, deployment secrets, routing policies, project settings, budgets, compression, and telemetry.
- Prepare rollback. Retain the previous configuration until the migration passes your acceptance tests.
Quick start for Claude Code users
You can give Claude Code the following prompt to assist with application changes. This is a migration prompt, not a downloadable NexLLM skill or a change to Claude Code’s own provider configuration.Step 1: Update environment variables
Keep the new configuration separate during testing.Step 2: Update your client
OpenAI SDK and Chat Completions
Replace Merge’s OpenAI compatibility URL:- Python
- JavaScript
- cURL
Native Merge SDK
ReplaceMergeGateway with an OpenAI client configured for NexLLM.
If the existing application calls client.responses.create(...), start with the Responses example in Step 8.
Do not automatically rewrite a Responses application into Chat Completions. Such a change also requires reviewing input structure, output parsing, tool results, streaming events, and conversation state.
Anthropic Messages integrations
You can retain the native Messages request format instead of converting every request to Chat Completions. NexLLM’s documented native example uses:POST /v1/messagesx-api-keycontaining the NexLLM keyanthropic-version: 2023-06-01
/v1/messages. Do not retain Merge’s /v1/anthropic prefix.
See the Quickstart.
Vercel AI SDK
Replacemerge-gateway-ai-sdk-provider with @ai-sdk/openai.
.chat(...) factory explicitly selects Chat Completions. Do not rely on the provider factory’s default API format.
See the AI SDK OpenAI provider documentation.
Step 3: Remove Merge-specific fields and headers
Record the purpose of each extension before removing it. The table below describes migration actions, not guaranteed one-to-one feature replacements.Request-body fields
Headers and SDK options
Search application code, shared clients, and deployment configuration for:project_id or tags nested inside wrapper options. Preserve unrelated headers and supported request parameters.
Replace Merge authentication with the appropriate NexLLM authentication from Step 2.
Response dependencies
Audit consumers of:- Gateway request or trace IDs.
- Resolved-provider headers.
- Routing-decision metadata.
- Rate-limit headers.
id.
See Chat Completions response fields.
Step 4: Decompose your routing policy
Channel groups define model access and pricing. Each NexLLM API key belongs to one group. A group is not a general-purpose replacement for Merge’s named routing policies or tag conditions. See Channel Groups. Use the following migration design:Selecting between channel groups
When a workload needs different groups, prepare separate keys with the required assignments and select the appropriate client in application code. Do not invent a request-bodygroup parameter.
Failure handling
Before enabling retries or fallbacks:- Define eligible error conditions.
- Set bounded retry counts and total deadlines.
- Account for SDK retries as well as application retries.
- Avoid restarting a partially streamed answer without an explicit recovery strategy.
- Preserve tool and structured-output requirements.
- Never relax data-handling requirements merely to obtain a successful response.
Step 5: Update model identifiers
Discover model IDs using the same key that will serve the migrated workload.data[].id exactly. The catalog reflects the key’s channel group.
NexLLM’s documented examples include:
See the Models API.
Do not apply a universal transformation such as:
- Removing every provider prefix.
- Preserving every Merge alias.
- Replacing
aws/withbedrock/. - Adding
openai/oranthropic/to every model.
openai/gpt-4o to gpt-4o only after verifying that the latter is available for your key.
Model listing alone is not a substitute for checking endpoint and feature compatibility. Review the model’s details in Model Square before testing Responses, tools, or multimodal requests.
See Model details and pricing.
Step 6: Map projects, budgets, and compression
Treat access controls, accounting, and context management as separate migration tasks.
NexLLM documents these key controls:
- Expiration.
- Remaining quota and unlimited-quota configuration.
- Model restrictions.
- IP allowlists.
- Channel-group assignment.
Recalculate effective cost
Account for model pricing, applicable token categories, and the assigned group ratio.Preserve context behavior
If Merge previously compressed requests, compare the actual model inputs before and after migration. Test long conversations, system instructions, tool-call/result pairs, structured content, and output quality. Do not silently discard required history just to make a request fit.Step 7: Reconnect observability
NexLLM Usage Logs document these per-request fields:- Request time.
- API key name.
- Model name.
- Timing.
- Input and output token consumption.
- Cost or quota deducted.
Usage logs alone do not establish distributed tracing, full prompt logging, audit-log parity, geographic processing guarantees, zero data retention, prompt-injection protection, or DLP enforcement.
Preserve Merge history
Before deprovisioning Merge:- Identify historical logs and configuration versions you must retain.
- Check available export mechanisms.
- Export permitted records into access-controlled storage.
- Verify completeness and readability.
- Record retention and deletion responsibilities.
Step 8: Optionally use the Responses API
NexLLM lists an OpenAI Responses-compatible endpoint atPOST /v1/responses.
For a native Merge Responses application, evaluate this path before converting to another API format. See the API Reference.
Set a separate model variable after confirming Responses compatibility in the model’s details:
- Python
- JavaScript
- cURL
Why migrate to NexLLM
Reuse familiar interfaces
Use documented OpenAI-compatible endpoints or native Claude Messages where appropriate. See API Reference.Access multiple model families
Build against supported GPT, Claude, and Gemini models while selecting the appropriate model for each workload. See GPT models, Claude models, and Gemini models.Choose access and pricing through channel groups
Review model availability and group ratios together instead of assuming one configuration fits every workload. See Channel Groups.Review usage after cutover
Use request-level records and dashboard statistics to compare the migrated workload with your baseline. See Usage & Logs.Troubleshooting
Model not found or unavailable
Query/v1/models with the failing application’s key. Compare the exact identifier with your configuration, then check its group and model restrictions.
Do not assume an old Merge alias remains valid.
Authentication fails
Check for stale secrets, extra whitespace, expired credentials, and source-IP restrictions. Confirm that the request uses NexLLM credentials and the authentication format appropriate to its endpoint.Requests succeed but no NexLLM logs appear
Inspect the actual outbound host and the dashboard account you are viewing. Check separately initialized workers and background jobs; changing one client may not update every caller.Project attribution or tags disappeared
Inspect where Merge project IDs and tags were previously attached. Preserve those values in application telemetry and verify your project-to-key-name mapping.A routing policy no longer applies
Removingrouting_policy_id does not preserve the policy.
Keep affected traffic on the previous integration until every required routing condition has a tested replacement.
Long conversations fail or cost more
Compare the context actually sent before and after migration. Check whether compression or trimming was lost, then review context limits and applicable pricing.Response parsing fails
Confirm the endpoint before debugging the parser. Do not use Chat Completions-specific parsing for a Responses or native Messages result. Review streaming and tool-result handling as well.Connection errors or incorrect paths
For the OpenAI SDK, use:/openai, /anthropic, or /ai-sdk.
Test authenticated model discovery:
Production migration checklist
- Create and secure a dedicated NexLLM key.
- Check account funding and key controls.
- Confirm model IDs, group access, and endpoint compatibility.
- Update every client, worker, and deployment environment.
- Remove Merge-specific extensions without dropping required behavior.
- Validate routing, retries, fallbacks, and deadlines.
- Verify budget enforcement and project attribution.
- Test long-context and compression requirements.
- Test streaming, tools, structured output, and multimodal inputs where used.
- Preserve telemetry and historical records.
- Revalidate required security and data-handling controls.
- Compare output quality, latency, and effective cost.
- Roll out gradually with a tested rollback path.
- Remove unused dependencies and secrets after acceptance.