Skip to main content

Migrate from Merge Gateway to NexLLM: update your API client, review routing policies, and preserve project controls and observability.

Move your Merge Gateway integration to NexLLM while retaining the request format that best fits your application. This guide covers client configuration, Merge-specific extensions, routing policies, model identifiers, project attribution, budgets, and observability.
API compatibility does not transfer gateway configuration. Removing a Merge field does not recreate the routing, attribution, or budget behavior it previously controlled.

Migration overview

Choose the migration path that matches your existing integration.
Existing integrationNexLLM migration path
Native merge-gateway-sdkReplace the Merge client with the OpenAI SDK and evaluate the Responses endpoint.
OpenAI SDKUpdate the base URL, API key, and model identifier.
Anthropic Messages integrationKeep the native Messages format and use NexLLM’s documented authentication.
Vercel AI SDKReplace the Merge provider with a configured OpenAI provider and explicitly select the API format.
LangChain or another wrapperUpdate its underlying client and verify the final request URL and payload.
Direct HTTP requestsUpdate the endpoint, authentication, model, and Merge-specific request extensions.
NexLLM documents these endpoints:
PurposeFull endpoint
Chat Completionshttps://www.nexllm.ai/v1/chat/completions
Responseshttps://www.nexllm.ai/v1/responses
Claude Messageshttps://www.nexllm.ai/v1/messages
Model discoveryhttps://www.nexllm.ai/v1/models
See the API Reference.
The OpenAI SDK base URL is https://www.nexllm.ai/v1. Endpoint paths shown above already include /v1; do not append it twice.

Prerequisites

  1. Create a NexLLM API key. Use a dedicated migration key so you can isolate testing from existing production traffic. See API Keys.
  2. Check account funding. Review your available balance and top up if necessary. See Wallet.
  3. Choose a channel group. Confirm that it includes the models you need. See Channel Groups.
  4. Inventory the existing integration. Include client code, deployment secrets, routing policies, project settings, budgets, compression, and telemetry.
  5. Prepare rollback. Retain the previous configuration until the migration passes your acceptance tests.

Quick start for Claude Code users

You can give Claude Code the following prompt to assist with application changes. This is a migration prompt, not a downloadable NexLLM skill or a change to Claude Code’s own provider configuration.

Step 1: Update environment variables

Keep the new configuration separate during testing.
These environment variable names are conventions used by this guide; the examples pass them explicitly to each client. For OpenAI-compatible requests, use the NexLLM key as a bearer token. See Authentication. Do not forward Merge or upstream provider credentials to NexLLM. Keep credentials needed for rollback or other services until those dependencies are retired.

Step 2: Update your client

OpenAI SDK and Chat Completions

Replace Merge’s OpenAI compatibility URL:
with:
The following examples use the environment variables from Step 1.
See Chat Completions.

Native Merge SDK

Replace MergeGateway with an OpenAI client configured for NexLLM. If the existing application calls client.responses.create(...), start with the Responses example in Step 8. Do not automatically rewrite a Responses application into Chat Completions. Such a change also requires reviewing input structure, output parsing, tool results, streaming events, and conversation state.

Anthropic Messages integrations

You can retain the native Messages request format instead of converting every request to Chat Completions. NexLLM’s documented native example uses:
  • POST /v1/messages
  • x-api-key containing the NexLLM key
  • anthropic-version: 2023-06-01
This model is a documented example, not a recommendation to replace your production model with a different capability tier. For SDK wrappers, verify that the configured base URL produces exactly /v1/messages. Do not retain Merge’s /v1/anthropic prefix. See the Quickstart.

Vercel AI SDK

Replace merge-gateway-ai-sdk-provider with @ai-sdk/openai.
The .chat(...) factory explicitly selects Chat Completions. Do not rely on the provider factory’s default API format. See the AI SDK OpenAI provider documentation.

Step 3: Remove Merge-specific fields and headers

Record the purpose of each extension before removing it. The table below describes migration actions, not guaranteed one-to-one feature replacements.

Request-body fields

Headers and SDK options

Search application code, shared clients, and deployment configuration for:
Remove the Merge-specific entries, including project_id or tags nested inside wrapper options. Preserve unrelated headers and supported request parameters. Replace Merge authentication with the appropriate NexLLM authentication from Step 2.

Response dependencies

Audit consumers of:
  • Gateway request or trace IDs.
  • Resolved-provider headers.
  • Routing-decision metadata.
  • Rate-limit headers.
Do not assume header names or semantics carry over. Retain your own correlation IDs and, for Chat Completions, record the returned completion id. See Chat Completions response fields.

Step 4: Decompose your routing policy

Channel groups define model access and pricing. Each NexLLM API key belongs to one group. A group is not a general-purpose replacement for Merge’s named routing policies or tag conditions. See Channel Groups. Use the following migration design:
This guide does not assume support for routing.model.fallbacks, routing.provider.fallbacks, routing.model.sort, or model: "auto".Do not copy these Concentrate-specific examples into NexLLM requests without separate confirmation.

Selecting between channel groups

When a workload needs different groups, prepare separate keys with the required assignments and select the appropriate client in application code. Do not invent a request-body group parameter.

Failure handling

Before enabling retries or fallbacks:
  • Define eligible error conditions.
  • Set bounded retry counts and total deadlines.
  • Account for SDK retries as well as application retries.
  • Avoid restarting a partially streamed answer without an explicit recovery strategy.
  • Preserve tool and structured-output requirements.
  • Never relax data-handling requirements merely to obtain a successful response.

Step 5: Update model identifiers

Discover model IDs using the same key that will serve the migrated workload.
Use the returned data[].id exactly. The catalog reflects the key’s channel group. NexLLM’s documented examples include: See the Models API. Do not apply a universal transformation such as:
  • Removing every provider prefix.
  • Preserving every Merge alias.
  • Replacing aws/ with bedrock/.
  • Adding openai/ or anthropic/ to every model.
For example, migrate openai/gpt-4o to gpt-4o only after verifying that the latter is available for your key. Model listing alone is not a substitute for checking endpoint and feature compatibility. Review the model’s details in Model Square before testing Responses, tools, or multimodal requests. See Model details and pricing.

Step 6: Map projects, budgets, and compression

Treat access controls, accounting, and context management as separate migration tasks. NexLLM documents these key controls:
  • Expiration.
  • Remaining quota and unlimited-quota configuration.
  • Model restrictions.
  • IP allowlists.
  • Channel-group assignment.
The key documentation describes Remaining Quota as a token-consumption limit and states that a key is disabled after exceeding it. It does not establish equivalence to Merge’s monetary budget periods or HTTP 402 behavior. See API Keys.

Recalculate effective cost

Account for model pricing, applicable token categories, and the assigned group ratio.
Verify current prices rather than copying fixed discounts into production assumptions. See Pricing.

Preserve context behavior

If Merge previously compressed requests, compare the actual model inputs before and after migration. Test long conversations, system instructions, tool-call/result pairs, structured content, and output quality. Do not silently discard required history just to make a request fit.

Step 7: Reconnect observability

NexLLM Usage Logs document these per-request fields:
  • Request time.
  • API key name.
  • Model name.
  • Timing.
  • Input and output token consumption.
  • Cost or quota deducted.
The documentation also describes filtering by time, model, and group. See Usage & Logs.
Usage logs alone do not establish distributed tracing, full prompt logging, audit-log parity, geographic processing guarantees, zero data retention, prompt-injection protection, or DLP enforcement.

Preserve Merge history

Before deprovisioning Merge:
  1. Identify historical logs and configuration versions you must retain.
  2. Check available export mechanisms.
  3. Export permitted records into access-controlled storage.
  4. Verify completeness and readability.
  5. Record retention and deletion responsibilities.
Do not assume historical records will automatically appear in NexLLM.

Step 8: Optionally use the Responses API

NexLLM lists an OpenAI Responses-compatible endpoint at POST /v1/responses. For a native Merge Responses application, evaluate this path before converting to another API format. See the API Reference. Set a separate model variable after confirming Responses compatibility in the model’s details:
These examples use the standard OpenAI SDK Responses interface. See the official Python SDK and JavaScript SDK.
Endpoint availability does not guarantee every Responses feature for every model. Test streaming, tools, structured output, multimodal input, and stored conversation state independently.Do not assume Merge response IDs can be reused with NexLLM or that previous_response_id has identical persistence behavior.

Why migrate to NexLLM

Reuse familiar interfaces

Use documented OpenAI-compatible endpoints or native Claude Messages where appropriate. See API Reference.

Access multiple model families

Build against supported GPT, Claude, and Gemini models while selecting the appropriate model for each workload. See GPT models, Claude models, and Gemini models.

Choose access and pricing through channel groups

Review model availability and group ratios together instead of assuming one configuration fits every workload. See Channel Groups.

Review usage after cutover

Use request-level records and dashboard statistics to compare the migrated workload with your baseline. See Usage & Logs.

Troubleshooting

Model not found or unavailable

Query /v1/models with the failing application’s key. Compare the exact identifier with your configuration, then check its group and model restrictions. Do not assume an old Merge alias remains valid.

Authentication fails

Check for stale secrets, extra whitespace, expired credentials, and source-IP restrictions. Confirm that the request uses NexLLM credentials and the authentication format appropriate to its endpoint.

Requests succeed but no NexLLM logs appear

Inspect the actual outbound host and the dashboard account you are viewing. Check separately initialized workers and background jobs; changing one client may not update every caller.

Project attribution or tags disappeared

Inspect where Merge project IDs and tags were previously attached. Preserve those values in application telemetry and verify your project-to-key-name mapping.

A routing policy no longer applies

Removing routing_policy_id does not preserve the policy. Keep affected traffic on the previous integration until every required routing condition has a tested replacement.

Long conversations fail or cost more

Compare the context actually sent before and after migration. Check whether compression or trimming was lost, then review context limits and applicable pricing.

Response parsing fails

Confirm the endpoint before debugging the parser. Do not use Chat Completions-specific parsing for a Responses or native Messages result. Review streaming and tool-result handling as well.

Connection errors or incorrect paths

For the OpenAI SDK, use:
Remove Merge-specific suffixes such as /openai, /anthropic, or /ai-sdk. Test authenticated model discovery:
Then test the intended inference endpoint separately. Do not reuse Concentrate’s health-check path as a NexLLM endpoint.

Production migration checklist

  • Create and secure a dedicated NexLLM key.
  • Check account funding and key controls.
  • Confirm model IDs, group access, and endpoint compatibility.
  • Update every client, worker, and deployment environment.
  • Remove Merge-specific extensions without dropping required behavior.
  • Validate routing, retries, fallbacks, and deadlines.
  • Verify budget enforcement and project attribution.
  • Test long-context and compression requirements.
  • Test streaming, tools, structured output, and multimodal inputs where used.
  • Preserve telemetry and historical records.
  • Revalidate required security and data-handling controls.
  • Compare output quality, latency, and effective cost.
  • Roll out gradually with a tested rollback path.
  • Remove unused dependencies and secrets after acceptance.

Next steps

Feedback

When reporting a migration issue, include the client library and version, endpoint, model ID, approximate request time, sanitized error, and a minimal reproduction. Describe the Merge behavior you need to preserve. Remove credentials, private prompts, personal data, and sensitive configuration before sharing diagnostics.