Migrate from Portkey to NexLLM: update your client, review gateway configuration, and reconnect usage monitoring.
Move your Portkey integration to NexLLM while keeping an OpenAI-compatible request format. For basic Chat Completions calls, the migration involves changing the endpoint and credential, removing Portkey-specific settings, and selecting an accessible model. Applications using Portkey Configs or administrative features need an additional behavior-by-behavior review.Migration at a glance
- BYOK integration
- Config-driven integration
Replace the provider credential and Portkey headers with a NexLLM credential.
Before
After
Prerequisites
1
Create a NexLLM API key
Create a dedicated migration key in the dashboard. Review its expiration, quota, model restrictions, IP allowlist, and channel group.See API Keys.
2
Check funding and model access
Review your Wallet. NexLLM supports wallet top-ups through Stripe.Confirm that the intended model appears in the Models API response for your key.
3
Inventory Portkey dependencies
Locate SDK initialization, headers, Configs, Virtual Keys, prompt templates, observability integrations, and administrative automation.Include settings stored in Portkey or deployment systems—not just code in your repository.
4
Prepare acceptance tests and rollback
Define checks for output quality, required features, errors, latency, cost, and policy enforcement. Keep the existing integration available until the replacement passes those checks.
Quick start for Claude Code users
You can give a coding assistant the following migration brief. This is an application-editing prompt, not an installable NexLLM skill or a change to Claude Code’s own API configuration.Migration brief
Step 1: Update your environment variables
Use separate variables for the new integration so rollback remains explicit.Before
After
NEXLLM_BASE_URL and NEXLLM_MODEL are application variables used by this guide. Pass them explicitly to your client.
For OpenAI-compatible requests, authenticate with:
Keep rollback credentials in your secret manager during verification. Do not send them to NexLLM. Remove or revoke them only after confirming that no other application still needs them.
Step 2: Update your client
If you use theportkey-ai SDK, replace its inference calls with an OpenAI-compatible client or direct HTTP requests. Review its administrative and prompt-management calls separately.
Install the SDK for your application language:
Step 3: Remove Portkey-specific headers
Treat these as migration actions, not one-to-one feature mappings. Record the dependency first, then remove it from requests sent to NexLLM.Request header checklist
Request header checklist
Portkey SDK settings to locate
Portkey SDK settings to locate
Search for both Python and JavaScript naming styles:
Do not copy these constructor options into the replacement client without reviewing what they do.
Response header dependencies
Response header dependencies
Update code that reads Portkey-specific response headers.
- Associate your local correlation ID with the completion response
id. - Rebuild cache metrics from verified response fields.
- Remove dependencies on Portkey’s fallback target index.
- Inspect error responses rather than assuming rate-limit header parity.
- Do not assume an equivalent provider-identification header.
id and input/output token usage fields.Step 4: Rebuild required Config behavior
Create a migration record for each saved or inline Portkey Config. Assign every behavior an owner and an acceptance test.Channel groups are not Config strategies
A NexLLM key belongs to one channel group. Its group determines model access and applies a pricing ratio. Groups do not, by themselves, reproduce your application’s ordered fallback chain, weighted distribution, or metadata conditions. For workflows requiring different groups, consider separate keys and an explicit application selection policy. Never allow untrusted user input to select arbitrary credentials. See Channel Groups.Set client timeout and retry behavior deliberately
For example, the OpenAI Python SDK accepts client-level settings:Reassess caching separately
NexLLM documents model-dependent cache pricing, including cache reads and writes. That does not establish equivalence with Portkey’s gateway response cache, semantic matching, namespaces, or invalidation rules. Test repeated prompts and inspect actual usage before estimating savings. Consult Pricing and Caching for the chosen model and group.Step 5: Verify model identifiers
Discover models using the same key that will serve production requests:data[].id value exactly. NexLLM’s documentation includes:
These are examples, not a guarantee that every key can access them.
If the expected model is missing, review the key’s channel group and model restrictions. Also confirm endpoint compatibility in Model Square before using a listed model for a new API shape.
See Models API and Model Details.
Step 6: Reconnect observability
NexLLM’s Usage Logs show request time, API key name, model, timing, token consumption, and cost. Documented filters include time range, model, and group. Its Dashboard summarizes requests, quota consumption, tokens, average RPM, and average TPM. See Usage Logs and Statistics.Preserve historical records
Before closing the old workspace:- Identify the logs, Config versions, templates, and evaluations to retain.
- Use the export facilities available to your Portkey account.
- Store the exports with appropriate access controls.
- Verify that records are complete and readable.
- Record the migration cutover time in both monitoring systems.
Step 7: Review administration and governance
NexLLM documents key expiration, quota limits, model restrictions, IP allowlists, and channel-group assignment. Configure these controls for the migrated integration using API Keys. Review other requirements individually:A channel group is a model-access and pricing setting. Do not treat it as evidence of organizational isolation or a replacement for a Portkey workspace.Likewise, a successful inference request does not verify governance parity.
Step 8: Optionally evaluate the Responses API
Keep Chat Completions for the initial migration unless there is a reason to change API shapes. NexLLM’s API Reference also lists:previous_response_id support across every model, or use conversation state as a substitute for observability tracing.
Optional: retain Claude-native request format
NexLLM also documents a Claude Messages endpoint. For an accessible Claude model, a minimal request is:Verify before cutover
The following smoke test checks model visibility and makes one inference request. The inference request may incur charges.verify_nexllm.py
- Streaming and interrupted-stream handling.
- Tool calls and structured outputs.
- Images, audio, or other required modalities.
- Timeout, retry, and fallback behavior.
- Cache freshness and tenant isolation.
- Prompt templates and versioning.
- Guardrails and access restrictions.
- Usage visibility and correlation.
- Representative quality, latency, and cost.
- Rollback under realistic deployment conditions.
Why migrate to NexLLM?
- Keep an OpenAI-compatible integration: reuse the documented Chat Completions format.
- Access multiple model families: discover supported GPT, Claude, and Gemini options through the Models API.
- Choose access and pricing settings: evaluate Channel Groups for your workload.
- Separate application credentials: configure dedicated API Keys.
- Review consumption centrally: use Usage Logs and Statistics.
Troubleshooting
Authentication fails
Authentication fails
Confirm that the credential came from NexLLM and that the client sends it using the authentication format for the chosen endpoint.Check for stale deployment secrets, accidental whitespace, expiration, and IP restrictions. Do not print the key while debugging.See Authentication.
Requests fail after previously working
Requests fail after previously working
Successful requests are missing from NexLLM logs
Successful requests are missing from NexLLM logs
Inspect the effective client configuration in the deployed application. Check for a stale Portkey base URL or a different NexLLM account/key.Review the time range, model, and group filters in Usage Logs.
Fallbacks, metadata filters, or guardrails changed
Fallbacks, metadata filters, or guardrails changed
Revisit the dependency inventory. A successful client replacement does not prove that the old gateway behavior was migrated.Restore the required control or stop the affected rollout until its replacement has passed acceptance tests.
Cache usage or cost differs
Cache usage or cost differs
Compare equivalent workloads using the selected model’s pricing, cache behavior, and channel-group ratio.Do not compare a Portkey response-cache hit directly with a provider prompt cache read; validate the actual behavior your application depends on.See Pricing.
Connection or endpoint errors
Connection or endpoint errors
Confirm the configured base URL:Append
/chat/completions, /models, or another documented endpoint only once. Avoid accidentally producing /v1/v1/....Use the authenticated /models request in Step 5 as a connectivity check. Do not assume a /responses/health endpoint exists.Next steps
API Reference
Review documented endpoints and authentication.
Available Models
Discover exact model identifiers for your key.
Channel Groups
Understand model access and pricing ratios.
Usage and Logs
Check request activity after cutover.