Move from LiteLLM to NexLLM: update authentication, clients, model identifiers, routing requirements, and observability.
Migrate an existing LiteLLM integration to NexLLM while preserving the application behavior you depend on. For a basic OpenAI-compatible Chat Completions integration, begin with three changes:- Point the client to
https://www.nexllm.ai/v1. - Authenticate with a NexLLM API key.
- Use a model ID available to that key.
Prerequisites
- A NexLLM account and an API key created in the dashboard.
- Sufficient account balance and available key quota.
- A channel group that includes the models you intend to use.
- Access to your application’s LiteLLM configuration and deployment settings.
- A staging environment and a rollback plan.
Identify your integration type
Quick Start for Claude Code Users
You can ask Claude Code to prepare the migration using the prompt below. This is a project-review prompt, not an official downloadable NexLLM migration skill.Step 1: Update Your Environment Variables
Create a dedicated NexLLM key for the workload. Configure its channel group, model restrictions, quota, expiration, and IP allowlist as needed in the dashboard.Before
After
NEXLLM_BASE_URL and NEXLLM_MODEL are application configuration variables used in this guide. Your code must read them explicitly.
If your application uses a .env file, ensure its existing environment loader loads these values. Configure deployment secrets separately.
Never put real API keys in source control or browser-delivered code.
See Authentication.
Step 2: Update Your Client
OpenAI-compatible Chat Completions
A typical LiteLLM Proxy client looks like this:Before
The OpenAI SDK base URL includes
/v1. The complete HTTP endpoint is https://www.nexllm.ai/v1/chat/completions.Do not append another /v1, retain the LiteLLM host, or substitute an undocumented NexLLM API hostname.Claude-native Messages requests
You do not need to convert an existing Messages integration to Chat Completions just to migrate. NexLLM documents:- Endpoint:
POST https://www.nexllm.ai/v1/messages - Authentication:
x-api-key - Version header:
anthropic-version: 2023-06-01
Claude-native HTTP
/v1/messages, not /v1/v1/messages.
Preserve Messages-specific request and response handling. Do not send an Anthropic Messages payload to /v1/chat/completions.
See Native Authentication.
Step 3: Review and Remove LiteLLM-Specific Headers
Inventory the behavior behind each header before removing it from NexLLM-bound requests. This guide does not establish NexLLM equivalents for LiteLLM’s custom gateway headers.Request headers
Response headers
Remove hard dependencies on LiteLLM-specific response headers.
A completion-body
id is not a promise of a particular HTTP request-ID header.
See Chat Completions and Usage & Logs.
Step 4: Rebuild the Behavior Behind config.yaml
Review the effective LiteLLM configuration, including settings maintained outside the YAML file.
Separate configuration into:
- Model and request defaults.
- Routing and resilience policies.
- Access and usage controls.
- Observability and privacy requirements.
Configuration mapping
Channel groups are not router strategies
Each NexLLM key belongs to one channel group. The group determines accessible models and the pricing multiplier applied to usage. A group is not a documented substitute for an ordered fallback chain, dynamic cost optimizer, regional policy, or least-busy router. For workload-specific routing:- Select a suitable group when creating the key.
- Check the models available to that key.
- Apply application-side selection only where needed and approved.
- Validate any provider, regional, or privacy constraints independently.
Preserve required failure behavior
Before replacing a fallback policy, define:- Which errors permit another attempt.
- Which models are acceptable alternatives.
- Maximum attempts and total request duration.
- Handling of context-window failures.
- Behavior after partial streaming output.
- How retries and fallback events are recorded.
Step 5: Update Model Identifiers
Inspect each LiteLLM alias and determine the model it actually represents. For example:Example LiteLLM configuration
my-chat-model, but that alias is local to the LiteLLM configuration.
For NexLLM, use an exact available model ID. Documented examples include:
These examples are not a guarantee that every key can access every model.
Discover models with the target key
data[].id values.
The catalog depends on the key’s channel group. Also review any key-level model restrictions.
Do not apply a blanket prefix conversion.An Azure deployment name may not identify the underlying model directly. A LiteLLM
bedrock/ or vertex_ai/ identifier does not establish the corresponding NexLLM ID. Some NexLLM IDs include a prefix; others do not.Application-owned alias
Step 6: Reconnect Observability and Accounting
NexLLM’s Usage Logs document request time, API key name, model name, timing, token consumption, and cost. Its dashboard also provides aggregate request and usage statistics. Use those surfaces for the information they expose, and preserve additional application instrumentation.
See Usage & Logs.
Reconcile costs using NexLLM pricing
Do not carry over LiteLLM’s local pricing assumptions. Review the selected model, token categories, applicable cache charges, and channel-group ratio. Compare representative requests against NexLLM’s recorded usage. See Pricing.Archive LiteLLM history
Before retiring the old deployment:- Identify the logs and configuration records you need to retain.
- Export them from the actual database or telemetry destination in use.
- Store the archive with appropriate access controls.
- Validate its completeness and readability.
- Keep the old system available until archival and rollback requirements are satisfied.
Step 7 — Optional: Evaluate the Responses API
NexLLM listsPOST /v1/responses as an OpenAI Responses API-compatible endpoint.
Adopting it is optional. Keep Chat Completions for the initial migration if that is what your application already uses.
Before changing API families, confirm the selected model supports the endpoint and the features your application requires.
Minimal Responses verification
NEXLLM_RESPONSES_MODEL is an application variable used by this example.
Update response parsing when changing API families. Do not reuse Chat Completions parsing without checking the Responses schema.
See API Overview.
Verify Before Production
Use a short synthetic prompt before testing real application data. The following smoke test checks catalog membership and a basic text completion. It does not establish feature parity.verify_nexllm.py
Production checklist
- The deployed process loads the intended NexLLM key.
- The endpoint and authentication match the API format.
- Every alias has a reviewed model mapping.
- Channel-group access and key restrictions are correct.
- No old gateway or upstream credentials are forwarded.
- Required routing, fallback, quota, and rate controls are preserved.
- Privacy and logging requirements have been confirmed.
- Streaming, cancellation, tools, and structured output pass relevant tests.
- Retry and timeout behavior is intentional.
- New usage appears in the expected NexLLM account.
- Historical logs are archived where required.
- Rollback restores the old endpoint, credentials, models, and options together.
What About the LiteLLM Python SDK?
If your application callslitellm.completion(...) directly, there may be no proxy client URL to change.
For a basic Chat Completions integration, one migration option is to replace the call with the OpenAI SDK configured for NexLLM.
Why Migrate to NexLLM?
Keep an OpenAI-compatible interface
Use a familiar client and request format for supported Chat Completions workflows.Access multiple model families
Discover available GPT, Claude, and Gemini models through the Models API.Configure workload-specific access
Use named API keys with documented quota, expiration, model, IP, and channel-group settings.Select access and pricing through channel groups
Choose an appropriate model-access category and review its pricing multiplier.Troubleshooting
Authentication fails
Authentication fails
Confirm that the process is loading a NexLLM key, not a LiteLLM virtual key or master key.Use Bearer authentication for OpenAI-compatible endpoints. For Claude-native Messages requests, use
x-api-key and the documented anthropic-version.Check key expiration, IP restrictions, quota, and account balance.Connection or endpoint errors
Connection or endpoint errors
Check the final outgoing URL.Chat Completions uses
https://www.nexllm.ai/v1/chat/completions.Claude Messages uses https://www.nexllm.ai/v1/messages.Remove duplicated path segments and stale proxy configuration. Use the documented Models endpoint for an authenticated connectivity check rather than inventing a health-check endpoint.Requests succeed but NexLLM logs appear empty
Requests succeed but NexLLM logs appear empty
Check the runtime client configuration, account, and selected log time range. Verify that the request did not still go through the old LiteLLM endpoint.If you retained a separate telemetry integration, inspect that independently.
Routing or fallbacks stopped working
Routing or fallbacks stopped working
Review the behavior previously implemented by LiteLLM Router,
config.yaml, or request headers.A direct NexLLM request does not execute your old local routing policy. Restore the existing path until a verified replacement is ready.Quota or cost differs from expectations
Quota or cost differs from expectations
Check account balance, key quota, selected model, token consumption, and channel-group pricing.Do not treat an old recurring budget, an RPM limit, and a NexLLM key quota as interchangeable controls.
An optional parameter or advanced feature fails
An optional parameter or advanced feature fails
Isolate the issue with a minimal request, then restore required features one at a time.Verify endpoint and model compatibility. Do not permanently remove a required feature just to make a smoke test pass.
Next Steps
- API Keys: Configure credentials and restrictions.
- Authentication: Review endpoint-specific authentication.
- Models API: Discover exact available model IDs.
- Channel Groups: Review access and pricing categories.
- Chat Completions: Review requests and streaming.
- Usage & Logs: Inspect migrated traffic.
- API Overview: Explore additional endpoints.
Feedback
If a requirement does not translate cleanly, prepare a sanitized report containing:- Your integration type.
- The affected LiteLLM setting or behavior.
- The target endpoint and model ID.
- Expected and observed behavior.
- A minimal reproduction with synthetic data.