> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nexllm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# LiteLLM

### Move from LiteLLM to NexLLM: update authentication, clients, model identifiers, routing requirements, and observability.

Migrate an existing LiteLLM integration to NexLLM while preserving the application behavior you depend on.

For a basic OpenAI-compatible Chat Completions integration, begin with three changes:

1. Point the client to `https://www.nexllm.ai/v1`.
2. Authenticate with a NexLLM API key.
3. Use a model ID available to that key.

Then review your LiteLLM-specific configuration, headers, routing, accounting, and logging before switching production traffic.

<Warning>
  API compatibility does not automatically migrate gateway behavior. Treat routing policies, fallbacks, budgets, rate limits, privacy controls, and observability integrations as separate migration requirements.
</Warning>

## Prerequisites

* A NexLLM account and an API key created in the dashboard.
* Sufficient account balance and available key quota.
* A channel group that includes the models you intend to use.
* Access to your application's LiteLLM configuration and deployment settings.
* A staging environment and a rollback plan.

See [API Keys](/get-started/getting-your-api-keys), [Wallet](/account/wallet-and-top-up), and [Channel Groups](/channel-groups/overview).

### Identify your integration type

| Integration | Migration approach |
| - | - |
| OpenAI SDK calling LiteLLM Proxy | Replace the base URL, credentials, and model IDs; review proxy-specific behavior. |
| HTTP client calling LiteLLM Proxy | Update the complete endpoint URL, authentication, and request configuration. |
| Anthropic client calling LiteLLM's Messages endpoint | Preserve the Messages format and use NexLLM's documented native authentication. |
| Direct `litellm.completion(...)` calls | Replace the library call with an appropriate client, or evaluate a separate compatibility approach. See the SDK section below. |
| Application using LiteLLM Router | Inventory routing and fallback behavior before removing the router. A client URL change does not replace it. |

## Quick Start for Claude Code Users

You can ask Claude Code to prepare the migration using the prompt below.

This is a project-review prompt, not an official downloadable NexLLM migration skill.

```text theme={null}
Prepare a migration from LiteLLM to NexLLM in this repository.

First produce an inventory and migration plan. Do not deploy changes.

1. Find LiteLLM Proxy clients, direct litellm SDK calls, Router usage,
   config.yaml files, model aliases, custom headers, callbacks, budgets,
   rate limits, caching, redaction, and fallback policies.

2. Inspect code and configuration structure without printing secrets,
   environment-file contents, private prompts, or historical request logs.

3. For OpenAI-compatible calls, use:
   https://www.nexllm.ai/v1
   with NEXLLM_API_KEY loaded from the environment.

4. For Claude-native HTTP calls, use:
   POST https://www.nexllm.ai/v1/messages
   with x-api-key and anthropic-version: 2023-06-01.

5. Map aliases to exact NexLLM model IDs confirmed for the target key.
   Do not mechanically remove or replace provider prefixes.

6. Review every x-litellm-* header and preserve its required behavior
   through a verified replacement or application implementation.

7. Do not invent auto-routing, fallback parameters, retention controls,
   BYOK support, budget hierarchies, or header mappings.

8. Preserve required telemetry and identify how to archive old logs.

9. Propose a small patch, a synthetic-data verification test, and rollback
   instructions. Ask for approval before live API calls or infrastructure
   changes. Live inference may incur charges.

Keep unresolved requirements as migration blockers.
```

## Step 1: Update Your Environment Variables

Create a dedicated NexLLM key for the workload.

Configure its channel group, model restrictions, quota, expiration, and IP allowlist as needed in the dashboard.

```bash Before theme={null}
export LITELLM_API_KEY="your-litellm-virtual-key"
export BASE_URL="http://localhost:4000"
```

```bash After theme={null}
export NEXLLM_API_KEY="your-nexllm-api-key"
export NEXLLM_BASE_URL="https://www.nexllm.ai/v1"

# Example only: confirm availability using your target API key.
export NEXLLM_MODEL="aws/claude-haiku-4-5"
```

`NEXLLM_BASE_URL` and `NEXLLM_MODEL` are application configuration variables used in this guide. Your code must read them explicitly.

If your application uses a `.env` file, ensure its existing environment loader loads these values. Configure deployment secrets separately.

<Warning>
  Keep the LiteLLM endpoint, virtual keys, master key, database, and upstream credentials available for rollback until the migration is accepted.

  Do not delete shared provider credentials or database settings merely because one workload has moved.
</Warning>

Never put real API keys in source control or browser-delivered code.

See [Authentication](/api-reference/authentication).

## Step 2: Update Your Client

### OpenAI-compatible Chat Completions

A typical LiteLLM Proxy client looks like this:

```python Before theme={null}
import os
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:4000",
    api_key=os.environ["LITELLM_API_KEY"],
    default_headers={
        "x-litellm-tags": "summarizer,production",
    },
)

response = client.chat.completions.create(
    model="my-chat-model",
    messages=[{"role": "user", "content": "Say hello."}],
)
```

Replace the proxy configuration and alias with NexLLM settings:

<CodeGroup>
  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url=os.environ["NEXLLM_BASE_URL"],
      api_key=os.environ["NEXLLM_API_KEY"],
      timeout=60.0,
      max_retries=0,
  )

  response = client.chat.completions.create(
      model=os.environ["NEXLLM_MODEL"],
      messages=[{"role": "user", "content": "Say hello in one word."}],
      max_tokens=100,
  )

  print(response.choices[0].message.content)
  ```

  ```javascript JavaScript theme={null}
  import OpenAI from "openai";

  const apiKey = process.env.NEXLLM_API_KEY;
  const baseURL = process.env.NEXLLM_BASE_URL;
  const model = process.env.NEXLLM_MODEL;

  if (!apiKey || !baseURL || !model) {
    throw new Error("Configure the NEXLLM environment variables first.");
  }

  const client = new OpenAI({
    baseURL,
    apiKey,
    timeout: 60_000,
    maxRetries: 0,
  });

  const response = await client.chat.completions.create({
    model,
    messages: [{ role: "user", content: "Say hello in one word." }],
    max_tokens: 100,
  });

  console.log(response.choices[0].message.content);
  ```

  ```bash cURL theme={null}
  # Requires jq.
  : "${NEXLLM_API_KEY:?Configure NEXLLM_API_KEY}"
  : "${NEXLLM_BASE_URL:?Configure NEXLLM_BASE_URL}"
  : "${NEXLLM_MODEL:?Configure NEXLLM_MODEL}"

  payload="$(jq -n --arg model "$NEXLLM_MODEL" '{
    model: $model,
    messages: [
      {role: "user", content: "Say hello in one word."}
    ],
    max_tokens: 100
  }')" && \
  curl --fail-with-body --silent --show-error \
    --max-time 60 \
    "${NEXLLM_BASE_URL}/chat/completions" \
    -H "Authorization: Bearer $NEXLLM_API_KEY" \
    -H "Content-Type: application/json" \
    --data-binary "$payload"
  ```
</CodeGroup>

These SDK examples disable client retries to make initial testing easier to interpret. Configure and test an intentional retry policy before production rollout.

<Note>
  The OpenAI SDK base URL includes `/v1`. The complete HTTP endpoint is `https://www.nexllm.ai/v1/chat/completions`.

  Do not append another `/v1`, retain the LiteLLM host, or substitute an undocumented NexLLM API hostname.
</Note>

See [Chat Completions](/api-reference/chat-completions).

### Claude-native Messages requests

You do not need to convert an existing Messages integration to Chat Completions just to migrate.

NexLLM documents:

* Endpoint: `POST https://www.nexllm.ai/v1/messages`
* Authentication: `x-api-key`
* Version header: `anthropic-version: 2023-06-01`

```bash Claude-native HTTP theme={null}
# Confirm this model is available to your key.
curl --fail-with-body --silent --show-error \
  --max-time 60 \
  https://www.nexllm.ai/v1/messages \
  -H "x-api-key: $NEXLLM_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aws/claude-haiku-4-5",
    "max_tokens": 100,
    "messages": [
      {"role": "user", "content": "Say hello in one word."}
    ]
  }'
```

When configuring an Anthropic SDK, check how your installed version combines its base URL with endpoint paths. The final request must reach `/v1/messages`, not `/v1/v1/messages`.

Preserve Messages-specific request and response handling. Do not send an Anthropic Messages payload to `/v1/chat/completions`.

See [Native Authentication](/api-reference/authentication).

## Step 3: Review and Remove LiteLLM-Specific Headers

Inventory the behavior behind each header before removing it from NexLLM-bound requests.

This guide does not establish NexLLM equivalents for LiteLLM's custom gateway headers.

### Request headers

| Existing header or behavior | Migration action |
| - | - |
| `Authorization: Bearer` with a LiteLLM virtual key | Use a NexLLM key for OpenAI-compatible requests. |
| Custom key header such as `X-Litellm-Key` | Replace it with the authentication required by the selected NexLLM endpoint. |
| `x-litellm-timeout`, `x-litellm-stream-timeout` | Configure and test client-side request and streaming timeouts. |
| `x-litellm-num-retries` | Move the required retry policy into a verified SDK or application configuration. |
| `x-litellm-tags` | Preserve required tags in application telemetry; review any routing dependency separately. |
| `x-litellm-spend-logs-metadata` | Preserve custom dimensions in your analytics system. |
| `x-litellm-customer-id`, `x-litellm-end-user-id` | Retain customer attribution in your application. Do not assume automatic dashboard equivalence. |
| `x-litellm-enable-message-redaction` | Verify a replacement privacy control before migrating affected data. Removing the header does not preserve redaction. |
| `openai-organization` used for an upstream account | Remove the old account association from the NexLLM client unless a documented requirement establishes otherwise. |
| `anthropic-version` | Keep it for NexLLM's Claude-native Messages requests. |
| `anthropic-beta` | Validate the specific beta feature and route before retaining it. |

<Warning>
  Do not translate a redaction header into a claimed Zero Data Retention guarantee. If a workload requires specific logging or retention behavior, obtain confirmation before moving that workload.
</Warning>

### Response headers

Remove hard dependencies on LiteLLM-specific response headers.

| Existing dependency | Migration approach |
| - | - |
| `x-litellm-call-id` | Keep an application correlation ID. For Chat Completions, record the documented response-body `id` when present. |
| `x-litellm-response-cost`, `x-litellm-key-spend` | Reconcile usage with NexLLM Usage Logs rather than assuming a replacement cost header. |
| `x-litellm-model-id`, `x-litellm-model-group`, `x-litellm-model-api-base` | Record the requested model and key configuration; inspect available NexLLM log fields. |
| Retry or fallback counters | Instrument the retry and fallback policy actually used after migration. |
| Duration and overhead headers | Measure application latency and review NexLLM's logged timing. |
| `x-litellm-version` | Remove proxy-version dependencies from the migrated path. |
| `llm_provider-*` | Do not assume provider metadata is forwarded unchanged. |
| Rate-limit headers | Verify the returned headers before relying on their names or semantics. |

A completion-body `id` is not a promise of a particular HTTP request-ID header.

See [Chat Completions](/api-reference/chat-completions) and [Usage & Logs](/account/usage-logs-and-statistics).

## Step 4: Rebuild the Behavior Behind `config.yaml`

Review the effective LiteLLM configuration, including settings maintained outside the YAML file.

Separate configuration into:

1. Model and request defaults.
2. Routing and resilience policies.
3. Access and usage controls.
4. Observability and privacy requirements.

Moving the inference URL does not transfer these settings.

### Configuration mapping

| LiteLLM concept | NexLLM migration approach |
| - | - |
| `model_list[].model_name` | Map each application alias to an exact NexLLM model ID. |
| `litellm_params.model` | Identify the actual model behind the deployment, then verify its NexLLM identifier. |
| Model-level request defaults | Move required supported defaults into application configuration or request construction. |
| Latency-based or cost-based routing | No equivalent strategy is established here. Choose an explicit model initially, then implement or verify the required selection policy. |
| Shuffle, least-busy, or usage-based routing | Reassess load-balancing requirements rather than assuming a channel group reproduces the algorithm. |
| `fallbacks`, `context_window_fallbacks` | Preserve an approved fallback policy in the application or retain the existing routing layer until a replacement is verified. |
| `num_retries`, `retry_policy`, `cooldown_time` | Review retry limits, backoff, and circuit-breaking separately. |
| `timeout`, `stream_timeout` | Configure the client with appropriate request and stream limits. |
| `max_budget`, `budget_duration` | Review NexLLM key quotas, but do not equate them with recurring currency budgets. Preserve required financial controls separately. |
| `rpm`, `tpm`, `max_parallel_requests` | Preserve required rate and concurrency enforcement; do not substitute a usage quota. |
| Upstream `api_key`, `api_base`, cloud credentials | Do not forward them in the direct NexLLM request. Keep credentials required by remaining workloads or rollback. |
| `general_settings.master_key` | Keep it out of the NexLLM client. |
| Callbacks, caching, guardrails, or redaction | Inventory and validate each required replacement before removal. |

### Channel groups are not router strategies

Each NexLLM key belongs to one channel group. The group determines accessible models and the pricing multiplier applied to usage.

A group is not a documented substitute for an ordered fallback chain, dynamic cost optimizer, regional policy, or least-busy router.

For workload-specific routing:

1. Select a suitable group when creating the key.
2. Check the models available to that key.
3. Apply application-side selection only where needed and approved.
4. Validate any provider, regional, or privacy constraints independently.

If different workloads require different groups, use separately configured keys and clients rather than inventing a per-request group parameter.

See [Channel Groups](/channel-groups/overview).

<Warning>
  Do not copy another gateway's `model: "auto"`, `routing.model.sort`, `routing.model.fallbacks`, or `routing.provider.fallbacks` into NexLLM requests without NexLLM documentation confirming support.
</Warning>

### Preserve required failure behavior

Before replacing a fallback policy, define:

* Which errors permit another attempt.
* Which models are acceptable alternatives.
* Maximum attempts and total request duration.
* Handling of context-window failures.
* Behavior after partial streaming output.
* How retries and fallback events are recorded.

Do not silently replace a required structured response with plain text merely to make a request succeed.

## Step 5: Update Model Identifiers

Inspect each LiteLLM alias and determine the model it actually represents.

For example:

```yaml Example LiteLLM configuration theme={null}
model_list:
  - model_name: my-chat-model
    litellm_params:
      model: openai/gpt-4o
```

The application calls `my-chat-model`, but that alias is local to the LiteLLM configuration.

For NexLLM, use an exact available model ID. Documented examples include:

| Model | NexLLM identifier |
| - | - |
| GPT-4o | `gpt-4o` |
| Claude Haiku 4.5 through the documented AWS route | `aws/claude-haiku-4-5` |
| Gemini 2.5 Flash | `gemini-2.5-flash` |

These examples are not a guarantee that every key can access every model.

### Discover models with the target key

```bash theme={null}
curl --fail-with-body --silent --show-error \
  --max-time 30 \
  https://www.nexllm.ai/v1/models \
  -H "Authorization: Bearer $NEXLLM_API_KEY"
```

Use the returned `data[].id` values.

The catalog depends on the key's channel group. Also review any key-level model restrictions.

<Note>
  Do not apply a blanket prefix conversion.

  An Azure deployment name may not identify the underlying model directly. A LiteLLM `bedrock/` or `vertex_ai/` identifier does not establish the corresponding NexLLM ID. Some NexLLM IDs include a prefix; others do not.
</Note>

If you want to keep application-facing aliases, resolve them in your own configuration before sending the request:

```python Application-owned alias theme={null}
MODEL_ALIASES = {
    "my-chat-model": "gpt-4o",
}

# Validate this mapping against the target NexLLM key before deployment.
nexllm_model = MODEL_ALIASES["my-chat-model"]
```

See [Models API](/api-reference/models).

## Step 6: Reconnect Observability and Accounting

NexLLM's Usage Logs document request time, API key name, model name, timing, token consumption, and cost.

Its dashboard also provides aggregate request and usage statistics.

Use those surfaces for the information they expose, and preserve additional application instrumentation.

| Existing requirement | Migration approach |
| - | - |
| Per-request usage history | Review new requests in NexLLM Usage Logs. |
| Model, token, timing, and cost inspection | Use the documented log fields. |
| Customer, feature, or environment tags | Preserve them in application telemetry. Named keys can help distinguish workloads, but are not a full analytics replacement. |
| Distributed traces and session relationships | Keep the existing tracing implementation or validate an alternative. |
| `success_callback` / `failure_callback` integrations | Reconnect required telemetry outside the removed LiteLLM execution path. |
| Virtual key provisioning | Create and configure NexLLM keys through the documented dashboard workflow. Do not assume LiteLLM's key-management API carries over. |
| Budget alerts or recurring spend limits | Preserve required enforcement and alerting until a verified replacement exists. |

See [Usage & Logs](/account/usage-logs-and-statistics).

### Reconcile costs using NexLLM pricing

Do not carry over LiteLLM's local pricing assumptions.

Review the selected model, token categories, applicable cache charges, and channel-group ratio. Compare representative requests against NexLLM's recorded usage.

See [Pricing](/channel-groups/pricing-overview).

### Archive LiteLLM history

Before retiring the old deployment:

1. Identify the logs and configuration records you need to retain.
2. Export them from the actual database or telemetry destination in use.
3. Store the archive with appropriate access controls.
4. Validate its completeness and readability.
5. Keep the old system available until archival and rollback requirements are satisfied.

Do not assume that changing the inference endpoint imports historical records into NexLLM.

## Step 7 — Optional: Evaluate the Responses API

NexLLM lists `POST /v1/responses` as an OpenAI Responses API-compatible endpoint.

Adopting it is optional. Keep Chat Completions for the initial migration if that is what your application already uses.

Before changing API families, confirm the selected model supports the endpoint and the features your application requires.

```python Minimal Responses verification theme={null}
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://www.nexllm.ai/v1",
    api_key=os.environ["NEXLLM_API_KEY"],
    timeout=60.0,
    max_retries=0,
)

# Configure this separately after confirming Responses compatibility.
response = client.responses.create(
    model=os.environ["NEXLLM_RESPONSES_MODEL"],
    input="Reply with a short greeting.",
)

# Synthetic-data diagnostic only; review the returned schema.
print(response.model_dump_json())
```

`NEXLLM_RESPONSES_MODEL` is an application variable used by this example.

<Warning>
  Catalog availability does not prove support for every endpoint or feature.

  Verify streaming, tools, structured output, multimodal input, storage, and conversation state individually. Do not assume universal support for `previous_response_id`, web search, or automatic feature degradation.
</Warning>

Update response parsing when changing API families. Do not reuse Chat Completions parsing without checking the Responses schema.

See [API Overview](/api-reference/overview).

## Verify Before Production

Use a short synthetic prompt before testing real application data.

The following smoke test checks catalog membership and a basic text completion. It does not establish feature parity.

```python verify_nexllm.py theme={null}
import os
import sys
from openai import OpenAI

def main():
    api_key = os.environ.get("NEXLLM_API_KEY")
    model = os.environ.get("NEXLLM_MODEL")

    if not api_key or not model:
        print("FAIL: configure NEXLLM_API_KEY and NEXLLM_MODEL.")
        return 1

    client = OpenAI(
        base_url="https://www.nexllm.ai/v1",
        api_key=api_key,
        timeout=30.0,
        max_retries=0,
    )

    try:
        available = {item.id for item in client.models.list().data}
        if model not in available:
            print("FAIL: selected model was not returned for this key.")
            return 1

        response = client.chat.completions.create(
            model=model,
            messages=[
                {"role": "user", "content": "Reply with a short greeting."}
            ],
            max_tokens=100,
        )

        if not response.choices:
            print("FAIL: response contains no choices.")
            return 1

        content = response.choices[0].message.content
        if not isinstance(content, str) or not content.strip():
            print("FAIL: expected nonempty text.")
            return 1

    except Exception:
        print("FAIL: request failed. Inspect sanitized diagnostics.")
        return 1

    print("PASS: catalog lookup and basic text completion.")
    return 0

if __name__ == "__main__":
    sys.exit(main())
```

Run live inference only with approval; it may incur charges.

### Production checklist

* [ ] The deployed process loads the intended NexLLM key.
* [ ] The endpoint and authentication match the API format.
* [ ] Every alias has a reviewed model mapping.
* [ ] Channel-group access and key restrictions are correct.
* [ ] No old gateway or upstream credentials are forwarded.
* [ ] Required routing, fallback, quota, and rate controls are preserved.
* [ ] Privacy and logging requirements have been confirmed.
* [ ] Streaming, cancellation, tools, and structured output pass relevant tests.
* [ ] Retry and timeout behavior is intentional.
* [ ] New usage appears in the expected NexLLM account.
* [ ] Historical logs are archived where required.
* [ ] Rollback restores the old endpoint, credentials, models, and options together.

Start in staging, then move a small portion of production traffic.

Retire LiteLLM infrastructure only after acceptance and after confirming that no other workload depends on it.

## What About the LiteLLM Python SDK?

If your application calls `litellm.completion(...)` directly, there may be no proxy client URL to change.

For a basic Chat Completions integration, one migration option is to replace the call with the OpenAI SDK configured for NexLLM.

<CodeGroup>
  ```python Before theme={null}
  import litellm

  response = litellm.completion(
      model="openai/gpt-4o",
      messages=[{"role": "user", "content": "Say hello."}],
  )
  ```

  ```python After theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://www.nexllm.ai/v1",
      api_key=os.environ["NEXLLM_API_KEY"],
  )

  response = client.chat.completions.create(
      model=os.environ["NEXLLM_MODEL"],
      messages=[{"role": "user", "content": "Say hello."}],
  )

  print(response.choices[0].message.content)
  ```
</CodeGroup>

Review asynchronous calls, streaming, exception handling, response access, callbacks, token accounting, and Router usage separately.

Do not remove LiteLLM from your dependencies until all remaining imports and integrations have been reviewed.

## Why Migrate to NexLLM?

### Keep an OpenAI-compatible interface

Use a familiar client and request format for supported Chat Completions workflows.

### Access multiple model families

Discover available GPT, Claude, and Gemini models through the Models API.

### Configure workload-specific access

Use named API keys with documented quota, expiration, model, IP, and channel-group settings.

### Select access and pricing through channel groups

Choose an appropriate model-access category and review its pricing multiplier.

## Troubleshooting

<AccordionGroup>
  <Accordion title="Model not found or unavailable">
    Call `GET /v1/models` with the same key used by the application.

    Check the exact model ID, the assigned channel group, and any key-level model restrictions. A LiteLLM alias or cloud deployment name is not automatically a NexLLM model ID.
  </Accordion>

  <Accordion title="Authentication fails">
    Confirm that the process is loading a NexLLM key, not a LiteLLM virtual key or master key.

    Use Bearer authentication for OpenAI-compatible endpoints. For Claude-native Messages requests, use `x-api-key` and the documented `anthropic-version`.

    Check key expiration, IP restrictions, quota, and account balance.
  </Accordion>

  <Accordion title="Connection or endpoint errors">
    Check the final outgoing URL.

    Chat Completions uses `https://www.nexllm.ai/v1/chat/completions`.

    Claude Messages uses `https://www.nexllm.ai/v1/messages`.

    Remove duplicated path segments and stale proxy configuration. Use the documented Models endpoint for an authenticated connectivity check rather than inventing a health-check endpoint.
  </Accordion>

  <Accordion title="Requests succeed but NexLLM logs appear empty">
    Check the runtime client configuration, account, and selected log time range. Verify that the request did not still go through the old LiteLLM endpoint.

    If you retained a separate telemetry integration, inspect that independently.
  </Accordion>

  <Accordion title="Routing or fallbacks stopped working">
    Review the behavior previously implemented by LiteLLM Router, `config.yaml`, or request headers.

    A direct NexLLM request does not execute your old local routing policy. Restore the existing path until a verified replacement is ready.
  </Accordion>

  <Accordion title="Tags, customer attribution, or callbacks disappeared">
    Inspect where those values were previously recorded.

    Preserve required application telemetry rather than assuming it is recreated by a new API key or inference endpoint.
  </Accordion>

  <Accordion title="Quota or cost differs from expectations">
    Check account balance, key quota, selected model, token consumption, and channel-group pricing.

    Do not treat an old recurring budget, an RPM limit, and a NexLLM key quota as interchangeable controls.
  </Accordion>

  <Accordion title="An optional parameter or advanced feature fails">
    Isolate the issue with a minimal request, then restore required features one at a time.

    Verify endpoint and model compatibility. Do not permanently remove a required feature just to make a smoke test pass.
  </Accordion>
</AccordionGroup>

## Next Steps

* [API Keys](/get-started/getting-your-api-keys): Configure credentials and restrictions.
* [Authentication](/api-reference/authentication): Review endpoint-specific authentication.
* [Models API](/api-reference/models): Discover exact available model IDs.
* [Channel Groups](/channel-groups/overview): Review access and pricing categories.
* [Chat Completions](/api-reference/chat-completions): Review requests and streaming.
* [Usage & Logs](/account/usage-logs-and-statistics): Inspect migrated traffic.
* [API Overview](/api-reference/overview): Explore additional endpoints.

## Feedback

If a requirement does not translate cleanly, prepare a sanitized report containing:

* Your integration type.
* The affected LiteLLM setting or behavior.
* The target endpoint and model ID.
* Expected and observed behavior.
* A minimal reproduction with synthetic data.

Share it through your existing NexLLM support channel.

Do not include API keys, authorization headers, private prompts, or sensitive historical logs.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.