Skip to main content

Move from LiteLLM to NexLLM: update authentication, clients, model identifiers, routing requirements, and observability.

Migrate an existing LiteLLM integration to NexLLM while preserving the application behavior you depend on. For a basic OpenAI-compatible Chat Completions integration, begin with three changes:
  1. Point the client to https://www.nexllm.ai/v1.
  2. Authenticate with a NexLLM API key.
  3. Use a model ID available to that key.
Then review your LiteLLM-specific configuration, headers, routing, accounting, and logging before switching production traffic.
API compatibility does not automatically migrate gateway behavior. Treat routing policies, fallbacks, budgets, rate limits, privacy controls, and observability integrations as separate migration requirements.

Prerequisites

  • A NexLLM account and an API key created in the dashboard.
  • Sufficient account balance and available key quota.
  • A channel group that includes the models you intend to use.
  • Access to your application’s LiteLLM configuration and deployment settings.
  • A staging environment and a rollback plan.
See API Keys, Wallet, and Channel Groups.

Identify your integration type

Quick Start for Claude Code Users

You can ask Claude Code to prepare the migration using the prompt below. This is a project-review prompt, not an official downloadable NexLLM migration skill.

Step 1: Update Your Environment Variables

Create a dedicated NexLLM key for the workload. Configure its channel group, model restrictions, quota, expiration, and IP allowlist as needed in the dashboard.
Before
After
NEXLLM_BASE_URL and NEXLLM_MODEL are application configuration variables used in this guide. Your code must read them explicitly. If your application uses a .env file, ensure its existing environment loader loads these values. Configure deployment secrets separately.
Keep the LiteLLM endpoint, virtual keys, master key, database, and upstream credentials available for rollback until the migration is accepted.Do not delete shared provider credentials or database settings merely because one workload has moved.
Never put real API keys in source control or browser-delivered code. See Authentication.

Step 2: Update Your Client

OpenAI-compatible Chat Completions

A typical LiteLLM Proxy client looks like this:
Before
Replace the proxy configuration and alias with NexLLM settings:
These SDK examples disable client retries to make initial testing easier to interpret. Configure and test an intentional retry policy before production rollout.
The OpenAI SDK base URL includes /v1. The complete HTTP endpoint is https://www.nexllm.ai/v1/chat/completions.Do not append another /v1, retain the LiteLLM host, or substitute an undocumented NexLLM API hostname.
See Chat Completions.

Claude-native Messages requests

You do not need to convert an existing Messages integration to Chat Completions just to migrate. NexLLM documents:
  • Endpoint: POST https://www.nexllm.ai/v1/messages
  • Authentication: x-api-key
  • Version header: anthropic-version: 2023-06-01
Claude-native HTTP
When configuring an Anthropic SDK, check how your installed version combines its base URL with endpoint paths. The final request must reach /v1/messages, not /v1/v1/messages. Preserve Messages-specific request and response handling. Do not send an Anthropic Messages payload to /v1/chat/completions. See Native Authentication.

Step 3: Review and Remove LiteLLM-Specific Headers

Inventory the behavior behind each header before removing it from NexLLM-bound requests. This guide does not establish NexLLM equivalents for LiteLLM’s custom gateway headers.

Request headers

Do not translate a redaction header into a claimed Zero Data Retention guarantee. If a workload requires specific logging or retention behavior, obtain confirmation before moving that workload.

Response headers

Remove hard dependencies on LiteLLM-specific response headers. A completion-body id is not a promise of a particular HTTP request-ID header. See Chat Completions and Usage & Logs.

Step 4: Rebuild the Behavior Behind config.yaml

Review the effective LiteLLM configuration, including settings maintained outside the YAML file. Separate configuration into:
  1. Model and request defaults.
  2. Routing and resilience policies.
  3. Access and usage controls.
  4. Observability and privacy requirements.
Moving the inference URL does not transfer these settings.

Configuration mapping

Channel groups are not router strategies

Each NexLLM key belongs to one channel group. The group determines accessible models and the pricing multiplier applied to usage. A group is not a documented substitute for an ordered fallback chain, dynamic cost optimizer, regional policy, or least-busy router. For workload-specific routing:
  1. Select a suitable group when creating the key.
  2. Check the models available to that key.
  3. Apply application-side selection only where needed and approved.
  4. Validate any provider, regional, or privacy constraints independently.
If different workloads require different groups, use separately configured keys and clients rather than inventing a per-request group parameter. See Channel Groups.
Do not copy another gateway’s model: "auto", routing.model.sort, routing.model.fallbacks, or routing.provider.fallbacks into NexLLM requests without NexLLM documentation confirming support.

Preserve required failure behavior

Before replacing a fallback policy, define:
  • Which errors permit another attempt.
  • Which models are acceptable alternatives.
  • Maximum attempts and total request duration.
  • Handling of context-window failures.
  • Behavior after partial streaming output.
  • How retries and fallback events are recorded.
Do not silently replace a required structured response with plain text merely to make a request succeed.

Step 5: Update Model Identifiers

Inspect each LiteLLM alias and determine the model it actually represents. For example:
Example LiteLLM configuration
The application calls my-chat-model, but that alias is local to the LiteLLM configuration. For NexLLM, use an exact available model ID. Documented examples include: These examples are not a guarantee that every key can access every model.

Discover models with the target key

Use the returned data[].id values. The catalog depends on the key’s channel group. Also review any key-level model restrictions.
Do not apply a blanket prefix conversion.An Azure deployment name may not identify the underlying model directly. A LiteLLM bedrock/ or vertex_ai/ identifier does not establish the corresponding NexLLM ID. Some NexLLM IDs include a prefix; others do not.
If you want to keep application-facing aliases, resolve them in your own configuration before sending the request:
Application-owned alias
See Models API.

Step 6: Reconnect Observability and Accounting

NexLLM’s Usage Logs document request time, API key name, model name, timing, token consumption, and cost. Its dashboard also provides aggregate request and usage statistics. Use those surfaces for the information they expose, and preserve additional application instrumentation. See Usage & Logs.

Reconcile costs using NexLLM pricing

Do not carry over LiteLLM’s local pricing assumptions. Review the selected model, token categories, applicable cache charges, and channel-group ratio. Compare representative requests against NexLLM’s recorded usage. See Pricing.

Archive LiteLLM history

Before retiring the old deployment:
  1. Identify the logs and configuration records you need to retain.
  2. Export them from the actual database or telemetry destination in use.
  3. Store the archive with appropriate access controls.
  4. Validate its completeness and readability.
  5. Keep the old system available until archival and rollback requirements are satisfied.
Do not assume that changing the inference endpoint imports historical records into NexLLM.

Step 7 — Optional: Evaluate the Responses API

NexLLM lists POST /v1/responses as an OpenAI Responses API-compatible endpoint. Adopting it is optional. Keep Chat Completions for the initial migration if that is what your application already uses. Before changing API families, confirm the selected model supports the endpoint and the features your application requires.
Minimal Responses verification
NEXLLM_RESPONSES_MODEL is an application variable used by this example.
Catalog availability does not prove support for every endpoint or feature.Verify streaming, tools, structured output, multimodal input, storage, and conversation state individually. Do not assume universal support for previous_response_id, web search, or automatic feature degradation.
Update response parsing when changing API families. Do not reuse Chat Completions parsing without checking the Responses schema. See API Overview.

Verify Before Production

Use a short synthetic prompt before testing real application data. The following smoke test checks catalog membership and a basic text completion. It does not establish feature parity.
verify_nexllm.py
Run live inference only with approval; it may incur charges.

Production checklist

  • The deployed process loads the intended NexLLM key.
  • The endpoint and authentication match the API format.
  • Every alias has a reviewed model mapping.
  • Channel-group access and key restrictions are correct.
  • No old gateway or upstream credentials are forwarded.
  • Required routing, fallback, quota, and rate controls are preserved.
  • Privacy and logging requirements have been confirmed.
  • Streaming, cancellation, tools, and structured output pass relevant tests.
  • Retry and timeout behavior is intentional.
  • New usage appears in the expected NexLLM account.
  • Historical logs are archived where required.
  • Rollback restores the old endpoint, credentials, models, and options together.
Start in staging, then move a small portion of production traffic. Retire LiteLLM infrastructure only after acceptance and after confirming that no other workload depends on it.

What About the LiteLLM Python SDK?

If your application calls litellm.completion(...) directly, there may be no proxy client URL to change. For a basic Chat Completions integration, one migration option is to replace the call with the OpenAI SDK configured for NexLLM.
Review asynchronous calls, streaming, exception handling, response access, callbacks, token accounting, and Router usage separately. Do not remove LiteLLM from your dependencies until all remaining imports and integrations have been reviewed.

Why Migrate to NexLLM?

Keep an OpenAI-compatible interface

Use a familiar client and request format for supported Chat Completions workflows.

Access multiple model families

Discover available GPT, Claude, and Gemini models through the Models API.

Configure workload-specific access

Use named API keys with documented quota, expiration, model, IP, and channel-group settings.

Select access and pricing through channel groups

Choose an appropriate model-access category and review its pricing multiplier.

Troubleshooting

Call GET /v1/models with the same key used by the application.Check the exact model ID, the assigned channel group, and any key-level model restrictions. A LiteLLM alias or cloud deployment name is not automatically a NexLLM model ID.
Confirm that the process is loading a NexLLM key, not a LiteLLM virtual key or master key.Use Bearer authentication for OpenAI-compatible endpoints. For Claude-native Messages requests, use x-api-key and the documented anthropic-version.Check key expiration, IP restrictions, quota, and account balance.
Check the final outgoing URL.Chat Completions uses https://www.nexllm.ai/v1/chat/completions.Claude Messages uses https://www.nexllm.ai/v1/messages.Remove duplicated path segments and stale proxy configuration. Use the documented Models endpoint for an authenticated connectivity check rather than inventing a health-check endpoint.
Check the runtime client configuration, account, and selected log time range. Verify that the request did not still go through the old LiteLLM endpoint.If you retained a separate telemetry integration, inspect that independently.
Review the behavior previously implemented by LiteLLM Router, config.yaml, or request headers.A direct NexLLM request does not execute your old local routing policy. Restore the existing path until a verified replacement is ready.
Inspect where those values were previously recorded.Preserve required application telemetry rather than assuming it is recreated by a new API key or inference endpoint.
Check account balance, key quota, selected model, token consumption, and channel-group pricing.Do not treat an old recurring budget, an RPM limit, and a NexLLM key quota as interchangeable controls.
Isolate the issue with a minimal request, then restore required features one at a time.Verify endpoint and model compatibility. Do not permanently remove a required feature just to make a smoke test pass.

Next Steps

Feedback

If a requirement does not translate cleanly, prepare a sanitized report containing:
  • Your integration type.
  • The affected LiteLLM setting or behavior.
  • The target endpoint and model ID.
  • Expected and observed behavior.
  • A minimal reproduction with synthetic data.
Share it through your existing NexLLM support channel. Do not include API keys, authorization headers, private prompts, or sensitive historical logs.