Skip to main content

Migrate from Portkey to NexLLM: update your client, review gateway configuration, and reconnect usage monitoring.

Move your Portkey integration to NexLLM while keeping an OpenAI-compatible request format. For basic Chat Completions calls, the migration involves changing the endpoint and credential, removing Portkey-specific settings, and selecting an accessible model. Applications using Portkey Configs or administrative features need an additional behavior-by-behavior review.
API compatibility does not mean gateway feature parity. Removing a Portkey header does not recreate the routing, caching, tracing, or policy enforcement that header previously enabled.

Migration at a glance

Replace the provider credential and Portkey headers with a NexLLM credential.
Before
After
See the Quickstart for NexLLM client configuration.

Prerequisites

1

Create a NexLLM API key

Create a dedicated migration key in the dashboard. Review its expiration, quota, model restrictions, IP allowlist, and channel group.See API Keys.
2

Check funding and model access

Review your Wallet. NexLLM supports wallet top-ups through Stripe.Confirm that the intended model appears in the Models API response for your key.
3

Inventory Portkey dependencies

Locate SDK initialization, headers, Configs, Virtual Keys, prompt templates, observability integrations, and administrative automation.Include settings stored in Portkey or deployment systems—not just code in your repository.
4

Prepare acceptance tests and rollback

Define checks for output quality, required features, errors, latency, cost, and policy enforcement. Keep the existing integration available until the replacement passes those checks.

Quick start for Claude Code users

You can give a coding assistant the following migration brief. This is an application-editing prompt, not an installable NexLLM skill or a change to Claude Code’s own API configuration.
Migration brief

Step 1: Update your environment variables

Use separate variables for the new integration so rollback remains explicit.
Before
After
NEXLLM_BASE_URL and NEXLLM_MODEL are application variables used by this guide. Pass them explicitly to your client. For OpenAI-compatible requests, authenticate with:
Use the NexLLM-issued key—not a Portkey key or an upstream provider key. See Authentication.
Keep rollback credentials in your secret manager during verification. Do not send them to NexLLM. Remove or revoke them only after confirming that no other application still needs them.

Step 2: Update your client

If you use the portkey-ai SDK, replace its inference calls with an OpenAI-compatible client or direct HTTP requests. Review its administrative and prompt-management calls separately. Install the SDK for your application language:
The examples below use the environment variables from Step 1.
See Chat Completions for the documented request and response fields. Keep credentials and API calls on your backend rather than embedding keys in browser code.

Step 3: Remove Portkey-specific headers

Treat these as migration actions, not one-to-one feature mappings. Record the dependency first, then remove it from requests sent to NexLLM.
Search for both Python and JavaScript naming styles:Do not copy these constructor options into the replacement client without reviewing what they do.
Update code that reads Portkey-specific response headers.
  • Associate your local correlation ID with the completion response id.
  • Rebuild cache metrics from verified response fields.
  • Remove dependencies on Portkey’s fallback target index.
  • Inspect error responses rather than assuming rate-limit header parity.
  • Do not assume an equivalent provider-identification header.
NexLLM’s Chat Completions reference documents the completion id and input/output token usage fields.

Step 4: Rebuild required Config behavior

Create a migration record for each saved or inline Portkey Config. Assign every behavior an owner and an acceptance test.
Do not copy Concentrate-specific settings such as routing.model.fallbacks, routing.provider.fallbacks, routing.model.sort, or model: "auto" into NexLLM requests without explicit NexLLM documentation for that behavior.

Channel groups are not Config strategies

A NexLLM key belongs to one channel group. Its group determines model access and applies a pricing ratio. Groups do not, by themselves, reproduce your application’s ordered fallback chain, weighted distribution, or metadata conditions. For workflows requiring different groups, consider separate keys and an explicit application selection policy. Never allow untrusted user input to select arbitrary credentials. See Channel Groups.

Set client timeout and retry behavior deliberately

For example, the OpenAI Python SDK accepts client-level settings:
Here, SDK retries are disabled so the application can own its retry policy. This is an example configuration, not a universal production recommendation. Check your installed SDK’s retry behavior before adding another retry layer. See the OpenAI Python SDK documentation.

Reassess caching separately

NexLLM documents model-dependent cache pricing, including cache reads and writes. That does not establish equivalence with Portkey’s gateway response cache, semantic matching, namespaces, or invalidation rules. Test repeated prompts and inspect actual usage before estimating savings. Consult Pricing and Caching for the chosen model and group.

Step 5: Verify model identifiers

Discover models using the same key that will serve production requests:
Use a returned data[].id value exactly. NexLLM’s documentation includes: These are examples, not a guarantee that every key can access them.
Do not manufacture identifiers by prepending openai/, anthropic/, bedrock/, or another provider name. Use the exact identifier NexLLM returns.Preserve the intended model family and version. A change from Sonnet to Haiku, for example, is a model change—not just a provider-name translation.
If the expected model is missing, review the key’s channel group and model restrictions. Also confirm endpoint compatibility in Model Square before using a listed model for a new API shape. See Models API and Model Details.

Step 6: Reconnect observability

NexLLM’s Usage Logs show request time, API key name, model, timing, token consumption, and cost. Documented filters include time range, model, and group. Its Dashboard summarizes requests, quota consumption, tokens, average RPM, and average TPM. See Usage Logs and Statistics.

Preserve historical records

Before closing the old workspace:
  1. Identify the logs, Config versions, templates, and evaluations to retain.
  2. Use the export facilities available to your Portkey account.
  3. Store the exports with appropriate access controls.
  4. Verify that records are complete and readable.
  5. Record the migration cutover time in both monitoring systems.
Plan for a separate historical archive; do not rely on an automatic import.

Step 7: Review administration and governance

NexLLM documents key expiration, quota limits, model restrictions, IP allowlists, and channel-group assignment. Configure these controls for the migrated integration using API Keys. Review other requirements individually:
A channel group is a model-access and pricing setting. Do not treat it as evidence of organizational isolation or a replacement for a Portkey workspace.Likewise, a successful inference request does not verify governance parity.

Step 8: Optionally evaluate the Responses API

Keep Chat Completions for the initial migration unless there is a reason to change API shapes. NexLLM’s API Reference also lists:
After confirming that the selected model supports this endpoint, try a minimal request:
Validate streaming, tools, structured outputs, multimodal inputs, and state handling individually if your application needs them. Do not assume previous_response_id support across every model, or use conversation state as a substitute for observability tracing.

Optional: retain Claude-native request format

NexLLM also documents a Claude Messages endpoint. For an accessible Claude model, a minimal request is:
The credential is still your NexLLM key. See Native Authentication.

Verify before cutover

The following smoke test checks model visibility and makes one inference request. The inference request may incur charges.
verify_nexllm.py
This test does not establish feature or policy parity. Complete the checks relevant to your application:
  • Streaming and interrupted-stream handling.
  • Tool calls and structured outputs.
  • Images, audio, or other required modalities.
  • Timeout, retry, and fallback behavior.
  • Cache freshness and tenant isolation.
  • Prompt templates and versioning.
  • Guardrails and access restrictions.
  • Usage visibility and correlation.
  • Representative quality, latency, and cost.
  • Rollback under realistic deployment conditions.
Start with limited traffic, compare results, and expand only after your acceptance criteria are met.

Why migrate to NexLLM?

Compare these benefits against any gateway features your application would need to retain elsewhere.

Troubleshooting

Confirm that the credential came from NexLLM and that the client sends it using the authentication format for the chosen endpoint.Check for stale deployment secrets, accidental whitespace, expiration, and IP restrictions. Do not print the key while debugging.See Authentication.
Call /models with the exact deployed key. Compare the returned identifier with your configured model, then review the key’s group and restrictions.Do not assume a model visible under a different key is also accessible here.See Models API.
Inspect the error response and check the key’s status and remaining quota, account balance, and model access.Do not automatically retry every failure.See API Keys and Wallet.
Inspect the effective client configuration in the deployed application. Check for a stale Portkey base URL or a different NexLLM account/key.Review the time range, model, and group filters in Usage Logs.
Revisit the dependency inventory. A successful client replacement does not prove that the old gateway behavior was migrated.Restore the required control or stop the affected rollout until its replacement has passed acceptance tests.
Compare equivalent workloads using the selected model’s pricing, cache behavior, and channel-group ratio.Do not compare a Portkey response-cache hit directly with a provider prompt cache read; validate the actual behavior your application depends on.See Pricing.
Confirm the configured base URL:
Append /chat/completions, /models, or another documented endpoint only once. Avoid accidentally producing /v1/v1/....Use the authenticated /models request in Step 5 as a connectivity check. Do not assume a /responses/health endpoint exists.

Next steps

API Reference

Review documented endpoints and authentication.

Available Models

Discover exact model identifiers for your key.

Channel Groups

Understand model access and pricing ratios.

Usage and Logs

Check request activity after cutover.