> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nexllm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Merge

### Migrate from Merge Gateway to NexLLM: update your API client, review routing policies, and preserve project controls and observability.

Move your Merge Gateway integration to NexLLM while retaining the request format that best fits your application.

This guide covers client configuration, Merge-specific extensions, routing policies, model identifiers, project attribution, budgets, and observability.

<Warning>
  API compatibility does not transfer gateway configuration. Removing a Merge field does not recreate the routing, attribution, or budget behavior it previously controlled.
</Warning>

## Migration overview

Choose the migration path that matches your existing integration.

<table>
  <colgroup>
    <col width="257" />

    <col width="390" />
  </colgroup>

  <thead>
    <tr>
      <th>Existing integration</th>
      <th>NexLLM migration path</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>Native `merge-gateway-sdk`</td>
      <td>Replace the Merge client with the OpenAI SDK and evaluate the Responses endpoint.</td>
    </tr>

    <tr>
      <td>OpenAI SDK</td>
      <td>Update the base URL, API key, and model identifier.</td>
    </tr>

    <tr>
      <td>Anthropic Messages integration</td>
      <td>Keep the native Messages format and use NexLLM's documented authentication.</td>
    </tr>

    <tr>
      <td>Vercel AI SDK</td>
      <td>Replace the Merge provider with a configured OpenAI provider and explicitly select the API format.</td>
    </tr>

    <tr>
      <td>LangChain or another wrapper</td>
      <td>Update its underlying client and verify the final request URL and payload.</td>
    </tr>

    <tr>
      <td>Direct HTTP requests</td>
      <td>Update the endpoint, authentication, model, and Merge-specific request extensions.</td>
    </tr>
  </tbody>
</table>

NexLLM documents these endpoints:

<table>
  <colgroup>
    <col width="156" />

    <col width="414" />
  </colgroup>

  <thead>
    <tr>
      <th>Purpose</th>
      <th>Full endpoint</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>Chat Completions</td>
      <td>`https://www.nexllm.ai/v1/chat/completions`</td>
    </tr>

    <tr>
      <td>Responses</td>
      <td>`https://www.nexllm.ai/v1/responses`</td>
    </tr>

    <tr>
      <td>Claude Messages</td>
      <td>`https://www.nexllm.ai/v1/messages`</td>
    </tr>

    <tr>
      <td>Model discovery</td>
      <td>`https://www.nexllm.ai/v1/models`</td>
    </tr>
  </tbody>
</table>

See the [API Reference](/api-reference/overview).

<Note>
  The OpenAI SDK base URL is `https://www.nexllm.ai/v1`. Endpoint paths shown above already include `/v1`; do not append it twice.
</Note>

## Prerequisites

1. **Create a NexLLM API key.** Use a dedicated migration key so you can isolate testing from existing production traffic. See [API Keys](/get-started/getting-your-api-keys).
2. **Check account funding.** Review your available balance and top up if necessary. See [Wallet](/account/wallet-and-top-up).
3. **Choose a channel group.** Confirm that it includes the models you need. See [Channel Groups](/channel-groups/overview).
4. **Inventory the existing integration.** Include client code, deployment secrets, routing policies, project settings, budgets, compression, and telemetry.
5. **Prepare rollback.** Retain the previous configuration until the migration passes your acceptance tests.

## Quick start for Claude Code users

You can give Claude Code the following prompt to assist with application changes.

This is a migration prompt, not a downloadable NexLLM skill or a change to Claude Code's own provider configuration.

```text theme={null}
Migrate this project's Merge Gateway integration to NexLLM.

Before editing, inventory the integration and propose a migration plan.

Requirements:
- Use https://www.nexllm.ai/v1 for OpenAI-compatible clients.
- Read the NexLLM key from NEXLLM_API_KEY.
- Preserve the existing API format where practical.
- Use exact model IDs verified for the application's NexLLM key.
- Remove Merge-specific request fields and headers only after
  identifying the behavior they control.
- Do not invent NexLLM routing parameters or auto-routing aliases.
- Treat channel groups as model-access and pricing configuration.
- Preserve required routing, tagging, budgets, and telemetry in
  application code unless an equivalent feature is verified.
- Test Responses, streaming, tools, and conversation state separately.
- Update environment examples and tests.
- Report unresolved compatibility gaps.

Do not print secrets.
Ask before sending billable requests or changing production settings.
```

## Step 1: Update environment variables

Keep the new configuration separate during testing.

```bash theme={null}
# Before: example Merge configuration
export MERGE_GATEWAY_API_KEY="YOUR_MERGE_API_KEY"
export MERGE_BASE_URL="https://api-gateway.merge.dev/v1/openai"

# After: NexLLM configuration
export NEXLLM_API_KEY="YOUR_NEXLLM_API_KEY"
export NEXLLM_BASE_URL="https://www.nexllm.ai/v1"

# Example for an initial Chat Completions test.
# Confirm availability using the same API key before calling it.
export NEXLLM_MODEL="gpt-4o"
```

These environment variable names are conventions used by this guide; the examples pass them explicitly to each client.

For OpenAI-compatible requests, use the NexLLM key as a bearer token. See [Authentication](/api-reference/authentication).

Do not forward Merge or upstream provider credentials to NexLLM. Keep credentials needed for rollback or other services until those dependencies are retired.

## Step 2: Update your client

### OpenAI SDK and Chat Completions

Replace Merge's OpenAI compatibility URL:

```text theme={null}
https://api-gateway.merge.dev/v1/openai
```

with:

```text theme={null}
https://www.nexllm.ai/v1
```

The following examples use the environment variables from Step 1.

<Tabs>
  <Tab title="Python">
    ```bash theme={null}
    pip install openai
    ```

    ```python theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url=os.environ["NEXLLM_BASE_URL"],
        api_key=os.environ["NEXLLM_API_KEY"],
    )

    response = client.chat.completions.create(
        model=os.environ["NEXLLM_MODEL"],
        messages=[
            {"role": "user", "content": "Say hello in one word."}
        ],
        max_tokens=32,
    )

    print(response.choices[0].message.content)
    ```
  </Tab>

  <Tab title="JavaScript">
    ```bash theme={null}
    npm install openai
    ```

    ```javascript theme={null}
    import OpenAI from "openai";

    const client = new OpenAI({
      baseURL: process.env.NEXLLM_BASE_URL,
      apiKey: process.env.NEXLLM_API_KEY,
    });

    const response = await client.chat.completions.create({
      model: process.env.NEXLLM_MODEL,
      messages: [
        { role: "user", content: "Say hello in one word." },
      ],
      max_tokens: 32,
    });

    console.log(response.choices[0]?.message?.content);
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={null}
    curl "${NEXLLM_BASE_URL}/chat/completions" \
      -H "Authorization: Bearer $NEXLLM_API_KEY" \
      -H "Content-Type: application/json" \
      -d "{
        \"model\": \"${NEXLLM_MODEL}\",
        \"messages\": [
          {\"role\": \"user\", \"content\": \"Say hello in one word.\"}
        ],
        \"max_tokens\": 32
      }"
    ```
  </Tab>
</Tabs>

See [Chat Completions](/api-reference/chat-completions).

### Native Merge SDK

Replace `MergeGateway` with an OpenAI client configured for NexLLM.

If the existing application calls `client.responses.create(...)`, start with the Responses example in [Step 8](#step-8-optionally-use-the-responses-api).

Do not automatically rewrite a Responses application into Chat Completions. Such a change also requires reviewing input structure, output parsing, tool results, streaming events, and conversation state.

### Anthropic Messages integrations

You can retain the native Messages request format instead of converting every request to Chat Completions.

NexLLM's documented native example uses:

* `POST /v1/messages`
* `x-api-key` containing the NexLLM key
* `anthropic-version: 2023-06-01`

```bash theme={null}
curl https://www.nexllm.ai/v1/messages \
  -H "x-api-key: $NEXLLM_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aws/claude-haiku-4-5",
    "system": "You are a helpful assistant.",
    "max_tokens": 32,
    "messages": [
      {"role": "user", "content": "Say hello in one word."}
    ]
  }'
```

This model is a documented example, not a recommendation to replace your production model with a different capability tier.

For SDK wrappers, verify that the configured base URL produces exactly `/v1/messages`. Do not retain Merge's `/v1/anthropic` prefix.

See the [Quickstart](/get-started/quickstart).

### Vercel AI SDK

Replace `merge-gateway-ai-sdk-provider` with `@ai-sdk/openai`.

```bash theme={null}
npm install ai @ai-sdk/openai
```

```javascript theme={null}
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";

const nexllm = createOpenAI({
  baseURL: process.env.NEXLLM_BASE_URL,
  apiKey: process.env.NEXLLM_API_KEY,
});

const { text } = await generateText({
  model: nexllm.chat(process.env.NEXLLM_MODEL),
  prompt: "Say hello in one word.",
});

console.log(text);
```

The `.chat(...)` factory explicitly selects Chat Completions. Do not rely on the provider factory's default API format.

See the [AI SDK OpenAI provider documentation](https://ai-sdk.dev/providers/ai-sdk-providers/openai).

## Step 3: Remove Merge-specific fields and headers

Record the purpose of each extension before removing it. The table below describes migration actions, not guaranteed one-to-one feature replacements.

### Request-body fields

| Merge field | Migration action |
| - | - |
| `project_id` | Remove from the outbound request. Preserve the project identifier in application telemetry; consider a dedicated key per project. |
| `tags` | Keep required labels in your own telemetry or routing layer. |
| `routing_policy_id` | Recreate the required policy behavior explicitly before cutover. |
| `vendor`, `vendors` | Review exact model IDs and channel-group configuration. Do not mechanically translate vendor names into model prefixes. |
| `include_routing_metadata` | Update code that depends on this metadata; record application-side routing decisions yourself. |
| `model: "default_routing"` | Select an explicit, available model. Do not automatically replace it with `auto`. |

### Headers and SDK options

Search application code, shared clients, and deployment configuration for:

```text theme={null}
X-Project-Id
X-Merge-Tags
providerOptions.mergeGateway
requestOptions.extraBodyProperties
extra_headers
```

Remove the Merge-specific entries, including `project_id` or tags nested inside wrapper options. Preserve unrelated headers and supported request parameters.

Replace Merge authentication with the appropriate NexLLM authentication from Step 2.

### Response dependencies

Audit consumers of:

* Gateway request or trace IDs.
* Resolved-provider headers.
* Routing-decision metadata.
* Rate-limit headers.

Do not assume header names or semantics carry over. Retain your own correlation IDs and, for Chat Completions, record the returned completion `id`.

See [Chat Completions response fields](/api-reference/chat-completions).

## Step 4: Decompose your routing policy

Channel groups define model access and pricing. Each NexLLM API key belongs to one group.

A group is not a general-purpose replacement for Merge's named routing policies or tag conditions. See [Channel Groups](/channel-groups/overview).

Use the following migration design:

| Existing policy | Recommended implementation |
| - | - |
| Single target | Configure an exact model ID and verify access through the selected key. |
| Ordered fallback targets | Maintain an ordered candidate list in application code. |
| Least latency | Benchmark eligible model and group configurations using representative traffic. |
| Lowest cost / cost optimized | Compare effective pricing and measured token usage, subject to quality requirements. |
| Balanced / quality first | Evaluate candidate models against workload-specific acceptance criteria. |
| Tag-based routing | Evaluate conditions in application code and select an approved client/model configuration. |

<Warning>
  This guide does not assume support for `routing.model.fallbacks`, `routing.provider.fallbacks`, `routing.model.sort`, or `model: "auto"`.

  Do not copy these Concentrate-specific examples into NexLLM requests without separate confirmation.
</Warning>

### Selecting between channel groups

When a workload needs different groups, prepare separate keys with the required assignments and select the appropriate client in application code.

Do not invent a request-body `group` parameter.

### Failure handling

Before enabling retries or fallbacks:

* Define eligible error conditions.
* Set bounded retry counts and total deadlines.
* Account for SDK retries as well as application retries.
* Avoid restarting a partially streamed answer without an explicit recovery strategy.
* Preserve tool and structured-output requirements.
* Never relax data-handling requirements merely to obtain a successful response.

## Step 5: Update model identifiers

Discover model IDs using the same key that will serve the migrated workload.

```bash theme={null}
curl "${NEXLLM_BASE_URL}/models" \
  -H "Authorization: Bearer $NEXLLM_API_KEY"
```

Use the returned `data[].id` exactly. The catalog reflects the key's channel group.

NexLLM's documented examples include:

| Model family | Example identifier |
| - | - |
| GPT | `gpt-4o` |
| Claude | `aws/claude-haiku-4-5` |
| Gemini | `gemini-2.5-flash` |

See the [Models API](/api-reference/models).

Do not apply a universal transformation such as:

* Removing every provider prefix.
* Preserving every Merge alias.
* Replacing `aws/` with `bedrock/`.
* Adding `openai/` or `anthropic/` to every model.

For example, migrate `openai/gpt-4o` to `gpt-4o` only after verifying that the latter is available for your key.

Model listing alone is not a substitute for checking endpoint and feature compatibility. Review the model's details in Model Square before testing Responses, tools, or multimodal requests.

See [Model details and pricing](/channel-groups/pricing-overview).

## Step 6: Map projects, budgets, and compression

Treat access controls, accounting, and context management as separate migration tasks.

| Merge concept | NexLLM mapping or migration action |
| - | - |
| Project | Use a purpose-specific API key and retain a project-to-key-name mapping in your application records. |
| Soft budget alerts | Preserve existing alerting until replacement thresholds and notifications are verified. |
| Hard monetary budget | Review key quotas separately; do not equate them automatically with currency-denominated spending caps. |
| Daily, weekly, or monthly resets | Confirm reset semantics independently before relying on them. |
| Context compression | Keep required trimming or summarization in application code unless an equivalent service behavior is verified. |
| BYOK | Confirm support and terms separately. Do not assume provider accounts or free-BYOK pricing transfer. |
| Project membership and roles | Review account access separately from API-key creation. |

NexLLM documents these key controls:

* Expiration.
* Remaining quota and unlimited-quota configuration.
* Model restrictions.
* IP allowlists.
* Channel-group assignment.

The key documentation describes Remaining Quota as a token-consumption limit and states that a key is disabled after exceeding it. It does not establish equivalence to Merge's monetary budget periods or HTTP 402 behavior.

See [API Keys](/get-started/getting-your-api-keys).

### Recalculate effective cost

Account for model pricing, applicable token categories, and the assigned group ratio.

```text theme={null}
Effective usage cost = applicable model usage cost × group ratio
```

Verify current prices rather than copying fixed discounts into production assumptions.

See [Pricing](/channel-groups/pricing-overview).

### Preserve context behavior

If Merge previously compressed requests, compare the actual model inputs before and after migration.

Test long conversations, system instructions, tool-call/result pairs, structured content, and output quality. Do not silently discard required history just to make a request fit.

## Step 7: Reconnect observability

NexLLM Usage Logs document these per-request fields:

* Request time.
* API key name.
* Model name.
* Timing.
* Input and output token consumption.
* Cost or quota deducted.

The documentation also describes filtering by time, model, and group.

See [Usage & Logs](/account/usage-logs-and-statistics).

| Existing requirement | Migration approach |
| - | - |
| Request usage and timing | Validate new traffic in NexLLM Usage Logs. |
| Project attribution | Associate the non-secret key name with your project in application records. |
| Tag breakdowns | Keep tags in your telemetry system. |
| Application routing decisions | Log the selected policy branch, model, and key alias—not the secret key. |
| Distributed tracing | Preserve your existing tracing and correlation system. |
| Security and audit evidence | Verify required coverage and retention independently. |

<Note>
  Usage logs alone do not establish distributed tracing, full prompt logging, audit-log parity, geographic processing guarantees, zero data retention, prompt-injection protection, or DLP enforcement.
</Note>

### Preserve Merge history

Before deprovisioning Merge:

1. Identify historical logs and configuration versions you must retain.
2. Check available export mechanisms.
3. Export permitted records into access-controlled storage.
4. Verify completeness and readability.
5. Record retention and deletion responsibilities.

Do not assume historical records will automatically appear in NexLLM.

## Step 8: Optionally use the Responses API

NexLLM lists an OpenAI Responses-compatible endpoint at `POST /v1/responses`.

For a native Merge Responses application, evaluate this path before converting to another API format. See the [API Reference](/api-reference/overview).

Set a separate model variable after confirming Responses compatibility in the model's details:

```bash theme={null}
export NEXLLM_RESPONSES_MODEL="YOUR_VERIFIED_RESPONSES_MODEL_ID"
```

<Tabs>
  <Tab title="Python">
    ```python theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url=os.environ["NEXLLM_BASE_URL"],
        api_key=os.environ["NEXLLM_API_KEY"],
    )

    response = client.responses.create(
        model=os.environ["NEXLLM_RESPONSES_MODEL"],
        input="Say hello in one word.",
    )

    print(response.output_text)
    ```
  </Tab>

  <Tab title="JavaScript">
    ```javascript theme={null}
    import OpenAI from "openai";

    const client = new OpenAI({
      baseURL: process.env.NEXLLM_BASE_URL,
      apiKey: process.env.NEXLLM_API_KEY,
    });

    const response = await client.responses.create({
      model: process.env.NEXLLM_RESPONSES_MODEL,
      input: "Say hello in one word.",
    });

    console.log(response.output_text);
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={null}
    curl "${NEXLLM_BASE_URL}/responses" \
      -H "Authorization: Bearer $NEXLLM_API_KEY" \
      -H "Content-Type: application/json" \
      -d "{
        \"model\": \"${NEXLLM_RESPONSES_MODEL}\",
        \"input\": \"Say hello in one word.\"
      }"
    ```
  </Tab>
</Tabs>

These examples use the standard OpenAI SDK Responses interface. See the official [Python SDK](https://github.com/openai/openai-python) and [JavaScript SDK](https://github.com/openai/openai-node).

<Warning>
  Endpoint availability does not guarantee every Responses feature for every model. Test streaming, tools, structured output, multimodal input, and stored conversation state independently.

  Do not assume Merge response IDs can be reused with NexLLM or that `previous_response_id` has identical persistence behavior.
</Warning>

## Why migrate to NexLLM

### Reuse familiar interfaces

Use documented OpenAI-compatible endpoints or native Claude Messages where appropriate.

See [API Reference](/api-reference/overview).

### Access multiple model families

Build against supported GPT, Claude, and Gemini models while selecting the appropriate model for each workload.

See [GPT models](/models/gpt-series), [Claude models](/models/claude-series), and [Gemini models](/models/gemini-series).

### Choose access and pricing through channel groups

Review model availability and group ratios together instead of assuming one configuration fits every workload.

See [Channel Groups](/channel-groups/overview).

### Review usage after cutover

Use request-level records and dashboard statistics to compare the migrated workload with your baseline.

See [Usage & Logs](/account/usage-logs-and-statistics).

## Troubleshooting

### Model not found or unavailable

Query `/v1/models` with the failing application's key. Compare the exact identifier with your configuration, then check its group and model restrictions.

Do not assume an old Merge alias remains valid.

### Authentication fails

Check for stale secrets, extra whitespace, expired credentials, and source-IP restrictions.

Confirm that the request uses NexLLM credentials and the authentication format appropriate to its endpoint.

### Requests succeed but no NexLLM logs appear

Inspect the actual outbound host and the dashboard account you are viewing.

Check separately initialized workers and background jobs; changing one client may not update every caller.

### Project attribution or tags disappeared

Inspect where Merge project IDs and tags were previously attached.

Preserve those values in application telemetry and verify your project-to-key-name mapping.

### A routing policy no longer applies

Removing `routing_policy_id` does not preserve the policy.

Keep affected traffic on the previous integration until every required routing condition has a tested replacement.

### Long conversations fail or cost more

Compare the context actually sent before and after migration. Check whether compression or trimming was lost, then review context limits and applicable pricing.

### Response parsing fails

Confirm the endpoint before debugging the parser.

Do not use Chat Completions-specific parsing for a Responses or native Messages result. Review streaming and tool-result handling as well.

### Connection errors or incorrect paths

For the OpenAI SDK, use:

```text theme={null}
https://www.nexllm.ai/v1
```

Remove Merge-specific suffixes such as `/openai`, `/anthropic`, or `/ai-sdk`.

Test authenticated model discovery:

```bash theme={null}
curl -i https://www.nexllm.ai/v1/models \
  -H "Authorization: Bearer $NEXLLM_API_KEY"
```

Then test the intended inference endpoint separately. Do not reuse Concentrate's health-check path as a NexLLM endpoint.

## Production migration checklist

* [ ] Create and secure a dedicated NexLLM key.
* [ ] Check account funding and key controls.
* [ ] Confirm model IDs, group access, and endpoint compatibility.
* [ ] Update every client, worker, and deployment environment.
* [ ] Remove Merge-specific extensions without dropping required behavior.
* [ ] Validate routing, retries, fallbacks, and deadlines.
* [ ] Verify budget enforcement and project attribution.
* [ ] Test long-context and compression requirements.
* [ ] Test streaming, tools, structured output, and multimodal inputs where used.
* [ ] Preserve telemetry and historical records.
* [ ] Revalidate required security and data-handling controls.
* [ ] Compare output quality, latency, and effective cost.
* [ ] Roll out gradually with a tested rollback path.
* [ ] Remove unused dependencies and secrets after acceptance.

## Next steps

* [API Reference](/api-reference/overview)
* [Available Models](/api-reference/models)
* [Channel Groups](/channel-groups/overview)
* [Pricing](/channel-groups/pricing-overview)
* [Usage & Logs](/account/usage-logs-and-statistics)

## Feedback

When reporting a migration issue, include the client library and version, endpoint, model ID, approximate request time, sanitized error, and a minimal reproduction.

Describe the Merge behavior you need to preserve.

Remove credentials, private prompts, personal data, and sensitive configuration before sharing diagnostics.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.