> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nexllm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Portkey

### Migrate from Portkey to NexLLM: update your client, review gateway configuration, and reconnect usage monitoring.

Move your Portkey integration to NexLLM while keeping an OpenAI-compatible request format.

For basic Chat Completions calls, the migration involves changing the endpoint and credential, removing Portkey-specific settings, and selecting an accessible model. Applications using Portkey Configs or administrative features need an additional behavior-by-behavior review.

<Warning>
  API compatibility does not mean gateway feature parity. Removing a Portkey header does not recreate the routing, caching, tracing, or policy enforcement that header previously enabled.
</Warning>

## Migration at a glance

<Tabs>
  <Tab title="BYOK integration">
    Replace the provider credential and Portkey headers with a NexLLM credential.

    ```python Before theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url="https://api.portkey.ai/v1",
        api_key=os.environ["OPENAI_API_KEY"],
        default_headers={
            "x-portkey-api-key": os.environ["PORTKEY_API_KEY"],
            "x-portkey-provider": "openai",
        },
    )
    ```

    ```python After theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url="https://www.nexllm.ai/v1",
        api_key=os.environ["NEXLLM_API_KEY"],
    )
    ```
  </Tab>

  <Tab title="Config-driven integration">
    Remove the Config reference from the NexLLM client, but first record every behavior the Config enables.

    ```python Before theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url="https://api.portkey.ai/v1",
        api_key=os.environ["OPENAI_API_KEY"],
        default_headers={
            "x-portkey-api-key": os.environ["PORTKEY_API_KEY"],
            "x-portkey-config": "pc-production",
        },
    )
    ```

    ```python After theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        base_url="https://www.nexllm.ai/v1",
        api_key=os.environ["NEXLLM_API_KEY"],
    )

    # Client initialization does not migrate Config behavior.
    # Review routing, caching, retries, and policies separately.
    ```

    Complete [Step 4](#step-4-rebuild-required-config-behavior) before switching production traffic.
  </Tab>
</Tabs>

See the [Quickstart](/get-started/quickstart) for NexLLM client configuration.

## Prerequisites

<Steps>
  <Step title="Create a NexLLM API key">
    Create a dedicated migration key in the dashboard. Review its expiration, quota, model restrictions, IP allowlist, and channel group.

    See [API Keys](/get-started/getting-your-api-keys).
  </Step>

  <Step title="Check funding and model access">
    Review your [Wallet](/account/wallet-and-top-up). NexLLM supports wallet top-ups through Stripe.

    Confirm that the intended model appears in the [Models API](/api-reference/models) response for your key.
  </Step>

  <Step title="Inventory Portkey dependencies">
    Locate SDK initialization, headers, Configs, Virtual Keys, prompt templates, observability integrations, and administrative automation.

    Include settings stored in Portkey or deployment systems—not just code in your repository.
  </Step>

  <Step title="Prepare acceptance tests and rollback">
    Define checks for output quality, required features, errors, latency, cost, and policy enforcement. Keep the existing integration available until the replacement passes those checks.
  </Step>
</Steps>

## Quick start for Claude Code users

You can give a coding assistant the following migration brief. This is an application-editing prompt, not an installable NexLLM skill or a change to Claude Code's own API configuration.

```text Migration brief theme={null}
Migrate this application's Portkey integration to NexLLM.

Use https://docs.nexllm.ai/llms.txt to discover NexLLM documentation.

1. Inventory Portkey SDK calls, x-portkey-* headers, Configs,
   Virtual Keys, prompt management, and administrative integrations.
2. Record the behavior each dependency provides before removing it.
3. Use https://www.nexllm.ai/v1 for NexLLM OpenAI-compatible calls.
4. Load the NexLLM credential from NEXLLM_API_KEY.
5. Select exact model IDs returned by GET /v1/models.
6. Check the key's channel group and model restrictions.
7. Remove Portkey-specific settings from the NexLLM client.
8. Do not invent routing fields, cache controls, trace headers,
   governance APIs, or provider-prefix conventions.
9. Preserve required behavior in application code or flag a blocker.
10. Add a model-discovery check and a minimal inference smoke test.
11. Update deployment configuration and document rollback.

Never print or commit secrets.
Ask before sending billable requests or changing production traffic.
```

## Step 1: Update your environment variables

Use separate variables for the new integration so rollback remains explicit.

```bash Before theme={null}
export PORTKEY_API_KEY="YOUR_PORTKEY_KEY"
export OPENAI_API_KEY="YOUR_PROVIDER_KEY"
export PORTKEY_BASE_URL="https://api.portkey.ai/v1"
```

```bash After theme={null}
export NEXLLM_API_KEY="YOUR_NEXLLM_KEY"
export NEXLLM_BASE_URL="https://www.nexllm.ai/v1"

# Example only: confirm access in Step 5.
export NEXLLM_MODEL="gpt-4o"
```

`NEXLLM_BASE_URL` and `NEXLLM_MODEL` are application variables used by this guide. Pass them explicitly to your client.

For OpenAI-compatible requests, authenticate with:

```http theme={null}
Authorization: Bearer YOUR_NEXLLM_KEY
```

Use the NexLLM-issued key—not a Portkey key or an upstream provider key. See [Authentication](/api-reference/authentication).

<Note>
  Keep rollback credentials in your secret manager during verification. Do not send them to NexLLM. Remove or revoke them only after confirming that no other application still needs them.
</Note>

## Step 2: Update your client

If you use the `portkey-ai` SDK, replace its inference calls with an OpenAI-compatible client or direct HTTP requests. Review its administrative and prompt-management calls separately.

Install the SDK for your application language:

<CodeGroup>
  ```bash Python theme={null}
  pip install openai
  ```

  ```bash JavaScript theme={null}
  npm install openai
  ```
</CodeGroup>

The examples below use the environment variables from Step 1.

<CodeGroup>
  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url=os.environ["NEXLLM_BASE_URL"],
      api_key=os.environ["NEXLLM_API_KEY"],
  )

  result = client.chat.completions.create(
      model=os.environ["NEXLLM_MODEL"],
      messages=[
          {"role": "user", "content": "Reply with a short greeting."}
      ],
      max_tokens=32,
  )

  print(result.choices[0].message.content)
  ```

  ```javascript JavaScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: process.env.NEXLLM_BASE_URL,
    apiKey: process.env.NEXLLM_API_KEY,
  });

  const result = await client.chat.completions.create({
    model: process.env.NEXLLM_MODEL,
    messages: [
      { role: "user", content: "Reply with a short greeting." },
    ],
    max_tokens: 32,
  });

  console.log(result.choices[0]?.message?.content);
  ```

  ```bash cURL theme={null}
  curl --fail-with-body "${NEXLLM_BASE_URL}/chat/completions" \
    -H "Authorization: Bearer ${NEXLLM_API_KEY}" \
    -H "Content-Type: application/json" \
    -d "{
      \"model\": \"${NEXLLM_MODEL}\",
      \"messages\": [
        {\"role\": \"user\", \"content\": \"Reply with a short greeting.\"}
      ],
      \"max_tokens\": 32
    }"
  ```
</CodeGroup>

See [Chat Completions](/api-reference/chat-completions) for the documented request and response fields.

Keep credentials and API calls on your backend rather than embedding keys in browser code.

## Step 3: Remove Portkey-specific headers

Treat these as migration actions, not one-to-one feature mappings. Record the dependency first, then remove it from requests sent to NexLLM.

<AccordionGroup>
  <Accordion title="Request header checklist">
    | Portkey header | Migration action |
    | - | - |
    | `x-portkey-api-key` | Use NexLLM bearer authentication. |
    | `x-portkey-provider` | Select an exact NexLLM model ID and the appropriate channel group. |
    | `x-portkey-virtual-key` | Review provider-account dependencies before removing. |
    | `x-portkey-config` | Rebuild required behaviors using Step 4. |
    | `x-portkey-trace-id` | Keep an application-generated correlation ID in your telemetry. |
    | `x-portkey-metadata` | Preserve business tags in application telemetry. |
    | `x-portkey-cache-force-refresh` | Revisit freshness requirements; do not assume an equivalent control. |
    | `x-portkey-cache-namespace` | Revisit cache isolation; do not automatically convert it to `prompt_cache_key`. |
    | `x-portkey-request-timeout` | Configure the HTTP client's timeout. |
    | `x-portkey-forward-headers`, `x-portkey-sensitive-headers` | Audit forwarding and redaction dependencies. |
    | `x-portkey-custom-host` | Keep custom-host calls on a separately verified integration. |
    | `x-portkey-azure-*`, `x-portkey-vertex-*`, `x-portkey-aws-*` | Do not forward cloud credentials or deployment configuration. Revalidate model access. |
  </Accordion>

  <Accordion title="Portkey SDK settings to locate">
    Search for both Python and JavaScript naming styles:

    | Settings | Review area |
    | - | - |
    | `api_key`, `apiKey` | Authentication |
    | `virtual_key`, `virtualKey`, `provider` | Provider selection and credentials |
    | `config` | Gateway behavior |
    | `trace_id`, `traceID`, `metadata` | Observability |
    | `cache_force_refresh`, `cacheForceRefresh` | Cache freshness |
    | `cache_namespace`, `cacheNamespace` | Cache isolation |
    | `custom_host`, `customHost` | Endpoint dependencies |
    | `forward_headers`, `forwardHeaders` | Header forwarding |

    Do not copy these constructor options into the replacement client without reviewing what they do.
  </Accordion>

  <Accordion title="Response header dependencies">
    Update code that reads Portkey-specific response headers.

    * Associate your local correlation ID with the completion response `id`.
    * Rebuild cache metrics from verified response fields.
    * Remove dependencies on Portkey's fallback target index.
    * Inspect error responses rather than assuming rate-limit header parity.
    * Do not assume an equivalent provider-identification header.

    NexLLM's [Chat Completions reference](/api-reference/chat-completions) documents the completion `id` and input/output token usage fields.
  </Accordion>
</AccordionGroup>

## Step 4: Rebuild required Config behavior

Create a migration record for each saved or inline Portkey Config. Assign every behavior an owner and an acceptance test.

| Portkey behavior | Recommended migration plan |
| - | - |
| Fallback strategy and targets | Implement an explicit application fallback policy unless a suitable NexLLM control is verified. |
| Weighted load balancing | Preserve selection logic in your application or a separate orchestration layer. |
| Conditional routing by metadata | Choose the model and, where necessary, a separately configured key in application code. |
| Simple or semantic response cache | Retain a separate cache if the application requires response reuse. |
| Retry attempts and status filters | Define client-side retries, backoff, and an overall deadline. |
| Request timeout | Configure and test client timeout behavior. |
| Parameter overrides | Move only supported parameters into the inference request. |
| Virtual Keys in targets | Revalidate the target model, group, and provider-account requirements. |
| Custom hosts | Retain a separate client for those workloads. |
| Guardrails | Keep the existing enforcement path until a replacement passes tests. |

<Warning>
  Do not copy Concentrate-specific settings such as `routing.model.fallbacks`, `routing.provider.fallbacks`, `routing.model.sort`, or `model: "auto"` into NexLLM requests without explicit NexLLM documentation for that behavior.
</Warning>

### Channel groups are not Config strategies

A NexLLM key belongs to one channel group. Its group determines model access and applies a pricing ratio.

Groups do not, by themselves, reproduce your application's ordered fallback chain, weighted distribution, or metadata conditions.

For workflows requiring different groups, consider separate keys and an explicit application selection policy. Never allow untrusted user input to select arbitrary credentials.

See [Channel Groups](/channel-groups/overview).

### Set client timeout and retry behavior deliberately

For example, the OpenAI Python SDK accepts client-level settings:

```python theme={null}
client = OpenAI(
    base_url=os.environ["NEXLLM_BASE_URL"],
    api_key=os.environ["NEXLLM_API_KEY"],
    timeout=60.0,
    max_retries=0,
)
```

Here, SDK retries are disabled so the application can own its retry policy. This is an example configuration, not a universal production recommendation.

Check your installed SDK's retry behavior before adding another retry layer. See the [OpenAI Python SDK documentation](https://github.com/openai/openai-python).

### Reassess caching separately

NexLLM documents model-dependent cache pricing, including cache reads and writes. That does not establish equivalence with Portkey's gateway response cache, semantic matching, namespaces, or invalidation rules.

Test repeated prompts and inspect actual usage before estimating savings. Consult [Pricing and Caching](/channel-groups/pricing-overview) for the chosen model and group.

## Step 5: Verify model identifiers

Discover models using the same key that will serve production requests:

```bash theme={null}
curl --fail-with-body "${NEXLLM_BASE_URL}/models" \
  -H "Authorization: Bearer ${NEXLLM_API_KEY}"
```

Use a returned `data[].id` value exactly. NexLLM's documentation includes:

| Model family | Example identifier |
| - | - |
| GPT | `gpt-4o` |
| Claude | `aws/claude-haiku-4-5` |
| Gemini | `gemini-2.5-flash` |

These are examples, not a guarantee that every key can access them.

<Warning>
  Do not manufacture identifiers by prepending `openai/`, `anthropic/`, `bedrock/`, or another provider name. Use the exact identifier NexLLM returns.

  Preserve the intended model family and version. A change from Sonnet to Haiku, for example, is a model change—not just a provider-name translation.
</Warning>

If the expected model is missing, review the key's channel group and model restrictions. Also confirm endpoint compatibility in Model Square before using a listed model for a new API shape.

See [Models API](/api-reference/models) and [Model Details](/channel-groups/pricing-overview).

## Step 6: Reconnect observability

NexLLM's Usage Logs show request time, API key name, model, timing, token consumption, and cost. Documented filters include time range, model, and group.

Its Dashboard summarizes requests, quota consumption, tokens, average RPM, and average TPM.

See [Usage Logs and Statistics](/account/usage-logs-and-statistics).

| Existing dependency | Migration task |
| - | - |
| Per-call usage monitoring | Compare staging calls with NexLLM Usage Logs. |
| Portkey trace grouping | Retain application tracing and correlation IDs. |
| Custom metadata filters | Preserve tags in your telemetry system. |
| Feedback and evaluation records | Keep the existing evaluation store until a replacement is verified. |
| Full prompt/response inspection | Verify visibility and retention separately from usage logging. |
| Audit events and exports | Validate required events, access controls, and export procedures. |

### Preserve historical records

Before closing the old workspace:

1. Identify the logs, Config versions, templates, and evaluations to retain.
2. Use the export facilities available to your Portkey account.
3. Store the exports with appropriate access controls.
4. Verify that records are complete and readable.
5. Record the migration cutover time in both monitoring systems.

Plan for a separate historical archive; do not rely on an automatic import.

## Step 7: Review administration and governance

NexLLM documents key expiration, quota limits, model restrictions, IP allowlists, and channel-group assignment.

Configure these controls for the migrated integration using [API Keys](/get-started/getting-your-api-keys).

Review other requirements individually:

| Requirement | Acceptance check |
| - | - |
| Workspaces and membership | Confirm the required user isolation and permissions. |
| Virtual Keys and provider integrations | Confirm any dependency on a specific provider account or deployment. |
| Usage policies | Test quota behavior; validate rate limits and monetary budgets separately. |
| SCIM or identity provisioning | Verify an appropriate integration before replacing automation. |
| Prompt libraries and versioning | Preserve templates and version history in a verified system. |
| Guardrails and redaction | Test enforcement before serving production traffic. |
| Auditability | Confirm event coverage, retention, access, and exports. |
| Data handling and residency | Validate requirements for every intended route. |
| Secret references | Retain only secrets needed by active integrations or rollback. |

<Note>
  A channel group is a model-access and pricing setting. Do not treat it as evidence of organizational isolation or a replacement for a Portkey workspace.

  Likewise, a successful inference request does not verify governance parity.
</Note>

## Step 8: Optionally evaluate the Responses API

Keep Chat Completions for the initial migration unless there is a reason to change API shapes.

NexLLM's [API Reference](/api-reference/overview) also lists:

```http theme={null}
POST https://www.nexllm.ai/v1/responses
```

After confirming that the selected model supports this endpoint, try a minimal request:

```bash theme={null}
# Set this to a verified Responses-compatible model ID.
export NEXLLM_RESPONSES_MODEL="YOUR_RESPONSES_MODEL_ID"

curl --fail-with-body "${NEXLLM_BASE_URL}/responses" \
  -H "Authorization: Bearer ${NEXLLM_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "{
    \"model\": \"${NEXLLM_RESPONSES_MODEL}\",
    \"input\": \"Write one sentence welcoming a new teammate.\"
  }"
```

Validate streaming, tools, structured outputs, multimodal inputs, and state handling individually if your application needs them.

Do not assume `previous_response_id` support across every model, or use conversation state as a substitute for observability tracing.

### Optional: retain Claude-native request format

NexLLM also documents a Claude Messages endpoint. For an accessible Claude model, a minimal request is:

```bash theme={null}
curl --fail-with-body "${NEXLLM_BASE_URL}/messages" \
  -H "x-api-key: ${NEXLLM_API_KEY}" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aws/claude-haiku-4-5",
    "max_tokens": 32,
    "messages": [
      {"role": "user", "content": "Reply with a short greeting."}
    ]
  }'
```

The credential is still your NexLLM key. See [Native Authentication](/api-reference/authentication).

## Verify before cutover

The following smoke test checks model visibility and makes one inference request. The inference request may incur charges.

```python verify_nexllm.py theme={null}
import os
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["NEXLLM_BASE_URL"],
    api_key=os.environ["NEXLLM_API_KEY"],
    timeout=60.0,
    max_retries=0,
)

selected_model = os.environ["NEXLLM_MODEL"]
available_models = {item.id for item in client.models.list().data}

if selected_model not in available_models:
    raise SystemExit(
        "Selected model is not visible to this key. "
        "Review model selection, group, and restrictions."
    )

result = client.chat.completions.create(
    model=selected_model,
    messages=[
        {"role": "user", "content": "Reply with the word ready."}
    ],
    max_tokens=32,
)

if not result.choices:
    raise SystemExit("No completion choices were returned.")

print("Model discovery and basic inference succeeded.")
print("Completion ID:", result.id)
print("Output:", result.choices[0].message.content)
```

```bash theme={null}
python verify_nexllm.py
```

This test does not establish feature or policy parity. Complete the checks relevant to your application:

* [ ] Streaming and interrupted-stream handling.
* [ ] Tool calls and structured outputs.
* [ ] Images, audio, or other required modalities.
* [ ] Timeout, retry, and fallback behavior.
* [ ] Cache freshness and tenant isolation.
* [ ] Prompt templates and versioning.
* [ ] Guardrails and access restrictions.
* [ ] Usage visibility and correlation.
* [ ] Representative quality, latency, and cost.
* [ ] Rollback under realistic deployment conditions.

Start with limited traffic, compare results, and expand only after your acceptance criteria are met.

## Why migrate to NexLLM?

* **Keep an OpenAI-compatible integration:** reuse the documented [Chat Completions format](/api-reference/chat-completions).
* **Access multiple model families:** discover supported GPT, Claude, and Gemini options through the [Models API](/api-reference/models).
* **Choose access and pricing settings:** evaluate [Channel Groups](/channel-groups/overview) for your workload.
* **Separate application credentials:** configure dedicated [API Keys](/get-started/getting-your-api-keys).
* **Review consumption centrally:** use [Usage Logs and Statistics](/account/usage-logs-and-statistics).

Compare these benefits against any gateway features your application would need to retain elsewhere.

## Troubleshooting

<AccordionGroup>
  <Accordion title="Authentication fails">
    Confirm that the credential came from NexLLM and that the client sends it using the authentication format for the chosen endpoint.

    Check for stale deployment secrets, accidental whitespace, expiration, and IP restrictions. Do not print the key while debugging.

    See [Authentication](/api-reference/authentication).
  </Accordion>

  <Accordion title="The requested model is unavailable">
    Call `/models` with the exact deployed key. Compare the returned identifier with your configured model, then review the key's group and restrictions.

    Do not assume a model visible under a different key is also accessible here.

    See [Models API](/api-reference/models).
  </Accordion>

  <Accordion title="Requests fail after previously working">
    Inspect the error response and check the key's status and remaining quota, account balance, and model access.

    Do not automatically retry every failure.

    See [API Keys](/get-started/getting-your-api-keys) and [Wallet](/account/wallet-and-top-up).
  </Accordion>

  <Accordion title="Successful requests are missing from NexLLM logs">
    Inspect the effective client configuration in the deployed application. Check for a stale Portkey base URL or a different NexLLM account/key.

    Review the time range, model, and group filters in [Usage Logs](/account/usage-logs-and-statistics).
  </Accordion>

  <Accordion title="Fallbacks, metadata filters, or guardrails changed">
    Revisit the dependency inventory. A successful client replacement does not prove that the old gateway behavior was migrated.

    Restore the required control or stop the affected rollout until its replacement has passed acceptance tests.
  </Accordion>

  <Accordion title="Cache usage or cost differs">
    Compare equivalent workloads using the selected model's pricing, cache behavior, and channel-group ratio.

    Do not compare a Portkey response-cache hit directly with a provider prompt cache read; validate the actual behavior your application depends on.

    See [Pricing](/channel-groups/pricing-overview).
  </Accordion>

  <Accordion title="Connection or endpoint errors">
    Confirm the configured base URL:

    ```text theme={null}
    https://www.nexllm.ai/v1
    ```

    Append `/chat/completions`, `/models`, or another documented endpoint only once. Avoid accidentally producing `/v1/v1/...`.

    Use the authenticated `/models` request in Step 5 as a connectivity check. Do not assume a `/responses/health` endpoint exists.
  </Accordion>
</AccordionGroup>

## Next steps

<CardGroup cols={2}>
  <Card title="API Reference" href="/api-reference/overview">
    Review documented endpoints and authentication.
  </Card>

  <Card title="Available Models" href="/api-reference/models">
    Discover exact model identifiers for your key.
  </Card>

  <Card title="Channel Groups" href="/channel-groups/overview">
    Understand model access and pricing ratios.
  </Card>

  <Card title="Usage and Logs" href="/account/usage-logs-and-statistics">
    Check request activity after cutover.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.