> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nexllm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Helicone

## Migrate from Helicone to NexLLM

Move your Helicone Chat Completions integration to NexLLM by updating your base URL, API key, and model identifiers.

NexLLM supports the OpenAI Chat Completions format, so you can retain your existing OpenAI SDK or HTTP client. Helicone-specific headers, observability, caching, and routing policies need separate review.

Choose your existing integration mode below.

<Note>
  These examples cover OpenAI-compatible Chat Completions. If you use a provider-native API, an asynchronous logging integration, or a Helicone-specific SDK, review that integration separately rather than applying a base-URL replacement blindly.
</Note>

<Tabs>
  <Tab title="Proxy mode" id="proxy-mode">
    In provider-proxy mode, your existing integration sends a provider API key to Helicone's provider proxy and a separate `Helicone-Auth` header for Helicone authentication.

    For migrated requests, use your NexLLM key instead and remove the Helicone authentication header.

    <CodeGroup>
      ```python Python (OpenAI SDK) theme={null}
      # Before: Helicone's OpenAI proxy
      import os
      from openai import OpenAI

      client = OpenAI(
          base_url="https://oai.helicone.ai/v1",
          api_key=os.environ["OPENAI_API_KEY"],
          default_headers={
              "Helicone-Auth": (
                  f"Bearer {os.environ['HELICONE_API_KEY']}"
              ),
          },
      )

      # After: NexLLM
      client = OpenAI(
          base_url="https://www.nexllm.ai/v1",
          api_key=os.environ["NEXLLM_API_KEY"],
      )
      ```

      ```javascript JavaScript (OpenAI SDK) theme={null}
      import OpenAI from "openai";

      // Before: Helicone's OpenAI proxy
      const heliconeClient = new OpenAI({
        baseURL: "https://oai.helicone.ai/v1",
        apiKey: process.env.OPENAI_API_KEY,
        defaultHeaders: {
          "Helicone-Auth": `Bearer ${process.env.HELICONE_API_KEY}`,
        },
      });

      // After: NexLLM
      const apiKey = process.env.NEXLLM_API_KEY;

      if (!apiKey) {
        throw new Error("NEXLLM_API_KEY is not configured");
      }

      const nexllmClient = new OpenAI({
        baseURL: "https://www.nexllm.ai/v1",
        apiKey,
      });
      ```

      ```bash cURL theme={null}
      # Before: Helicone's OpenAI proxy
      curl --fail --silent --show-error \
        https://oai.helicone.ai/v1/chat/completions \
        -H "Authorization: Bearer $OPENAI_API_KEY" \
        -H "Helicone-Auth: Bearer $HELICONE_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
          "model": "gpt-4o",
          "messages": [
            {"role": "user", "content": "Say hello in one word."}
          ],
          "max_tokens": 100
        }'

      # After: NexLLM
      # Confirm gpt-4o is available to your NexLLM key first.
      curl --fail --silent --show-error \
        https://www.nexllm.ai/v1/chat/completions \
        -H "Authorization: Bearer $NEXLLM_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
          "model": "gpt-4o",
          "messages": [
            {"role": "user", "content": "Say hello in one word."}
          ],
          "max_tokens": 100
        }'
      ```
    </CodeGroup>
  </Tab>

  <Tab title="AI Gateway" id="ai-gateway">
    For a Helicone AI Gateway integration authenticated with a Helicone key, replace the gateway URL and credentials with NexLLM's configuration.

    Model identifiers and gateway-specific behavior still require review.

    <CodeGroup>
      ```python Python (OpenAI SDK) theme={null}
      import os
      from openai import OpenAI

      # Before: Helicone AI Gateway
      client = OpenAI(
          base_url="https://ai-gateway.helicone.ai",
          api_key=os.environ["HELICONE_API_KEY"],
      )

      # After: NexLLM
      client = OpenAI(
          base_url="https://www.nexllm.ai/v1",
          api_key=os.environ["NEXLLM_API_KEY"],
      )
      ```

      ```javascript JavaScript (OpenAI SDK) theme={null}
      import OpenAI from "openai";

      // Before: Helicone AI Gateway
      const heliconeClient = new OpenAI({
        baseURL: "https://ai-gateway.helicone.ai",
        apiKey: process.env.HELICONE_API_KEY,
      });

      // After: NexLLM
      const apiKey = process.env.NEXLLM_API_KEY;

      if (!apiKey) {
        throw new Error("NEXLLM_API_KEY is not configured");
      }

      const nexllmClient = new OpenAI({
        baseURL: "https://www.nexllm.ai/v1",
        apiKey,
      });
      ```

      ```bash cURL theme={null}
      # Before: Helicone AI Gateway
      # Use the model ID from your existing integration.
      curl --fail --silent --show-error \
        https://ai-gateway.helicone.ai/chat/completions \
        -H "Authorization: Bearer $HELICONE_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
          "model": "openai/gpt-4o",
          "messages": [
            {"role": "user", "content": "Say hello in one word."}
          ],
          "max_tokens": 100
        }'

      # After: NexLLM
      # This mapping is illustrative. Verify the exact target ID first.
      curl --fail --silent --show-error \
        https://www.nexllm.ai/v1/chat/completions \
        -H "Authorization: Bearer $NEXLLM_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
          "model": "gpt-4o",
          "messages": [
            {"role": "user", "content": "Say hello in one word."}
          ],
          "max_tokens": 100
        }'
      ```
    </CodeGroup>
  </Tab>
</Tabs>

<Warning>
  The model examples are not a promise of access for every key. Discover models using your NexLLM credentials before running inference. Requests may incur usage charges.
</Warning>

## Prerequisites

You will need:

* **A NexLLM account and API key.** Create a key in the NexLLM dashboard.
* **Access to your intended models.** Review the key's channel group and model restrictions.
* **An existing Helicone integration.** Identify whether it uses provider-proxy mode, AI Gateway mode, or a separate logging integration.
* **An inventory of required behavior.** Include caching, retries, fallbacks, moderation, privacy controls, request tagging, sessions, and downstream analytics.
* **A rollback plan.** Keep the existing configuration and credentials secure until the migrated paths have been validated.

`NEXLLM_API_KEY` is the environment-variable name used in this guide. Your application must read that same name.

## Quick Start for Claude Code Users

The NexLLM migration skill instructs Claude Code to inspect your integration, propose model mappings, review Helicone-specific behavior, and generate verification tests.

It asks for approval before editing and does not silently remove privacy, reliability, or observability requirements.

### 1. Copy and review the migration skill

Expand the section below and copy the complete code block into a local file named `migrate-helicone.md`.

Include the YAML frontmatter at the beginning. Review the instructions before installing.

<Accordion title="View and copy the complete migration skill">
  ```markdown migrate-helicone.md theme={null}
  ---
  name: migrate-helicone
  description: Migrate an application's Helicone provider-proxy or AI Gateway integration to NexLLM. Use when the user wants to replace Helicone with NexLLM.
  ---

  # Migrate from Helicone to NexLLM

  Migrate approved application call sites with minimal, reviewed changes.
  This skill changes application code, not Claude Code's inference provider.

  ## Target configuration

  - SDK base URL: `https://www.nexllm.ai/v1`
  - Chat Completions: `POST https://www.nexllm.ai/v1/chat/completions`
  - Model discovery: `GET https://www.nexllm.ai/v1/models`
  - Authentication: `Authorization: Bearer <NEXLLM_API_KEY>`
  - API key environment variable: `NEXLLM_API_KEY`
  - Optional application base URL variable: `NEXLLM_BASE_URL`

  Use exact discovered model IDs. Do not invent model prefixes, routing
  parameters, health endpoints, or credential formats.

  NexLLM documents a Responses API endpoint, but adopting it is a separate
  change requiring feature-specific verification. Do not change API
  families as part of the basic migration without approval.

  ## Safety

  - Inspect first and obtain approval before editing.
  - Never ask the user to paste credentials into the conversation.
  - Do not read or print secret-bearing environment files or secret values.
  - Search source code and placeholder-only examples, excluding secrets,
    dependency directories, generated output, and irrelevant files.
  - Do not log authentication headers or use verbose HTTP tracing.
  - Preserve unrelated edits and unmigrated integrations.
  - Do not globally replace API keys, provider names, or headers.
  - Ask before installing dependencies, making authenticated network calls,
    running inference tests, or deploying.
  - Keep rollback credentials in the existing secret store.
  - Do not revoke keys or remove required behavior without permission.

  Use normal session tool permissions. This skill does not pre-authorize
  unrestricted commands or modifications.

  ## 1. Inspect the integration

  Identify the language, SDK, HTTP client, environment-loading mechanism,
  tests, deployment configuration, and centralized client factories.

  Search non-secret files for:

  - `helicone.ai` and configured self-hosted Helicone endpoints.
  - `HELICONE_API_KEY` and Helicone base URL variables.
  - Case-insensitive `Helicone-` header names.
  - Helicone SDKs, adapters, wrappers, and asynchronous logging calls.
  - Consumers of Helicone response headers.
  - Model identifiers, fallback lists, and auto-routing configuration.
  - Custom properties, sessions, prompts, alerts, and export jobs.

  Classify each integration:

  1. Provider proxy: provider credentials plus Helicone authentication.
  2. AI Gateway: gateway credentials and gateway model identifiers.
  3. Other: asynchronous logging, native provider APIs, custom gateways,
     self-hosted deployments, or mixed configurations.

  Do not apply a Chat Completions URL swap to incompatible native payloads
  or logging-only integrations.

  Report relevant files, models, required features, and unresolved questions.

  ## 2. Check for a partial migration

  Look for `NEXLLM_API_KEY`, `NEXLLM_BASE_URL`, and NexLLM endpoints.

  Inspect each call site separately. A file containing NexLLM configuration
  may still contain an active Helicone integration.

  Preserve correctly migrated code.

  ## 3. Confirm scope and environment

  Ask whether this migration covers local development, deployment, or both.
  Ask which call sites should migrate and which must remain unchanged.

  STOP and wait for the user's response.

  Have the user configure credentials through their existing environment
  or secret manager. Never request the key itself.

  A local `.env` file must be loaded by the application or tooling.
  Do not create a local file as a substitute for deployment secrets.

  ## 4. Discover models and approve mappings

  With permission, query NexLLM's Models endpoint using credentials from
  the process environment. Print only model IDs or sanitized status.

  Use the same key intended for inference. Review its channel group and
  model restrictions when a required model is inaccessible.

  If network access or credentials are unavailable, ask the user to run
  discovery locally and share only non-sensitive model IDs.

  Build a mapping table with:

  - Existing model ID.
  - Exact proposed NexLLM ID.
  - Required endpoint and features.
  - Mapping status and unresolved differences.

  Do not assume a Helicone bare slug works unchanged.
  Do not strip author prefixes or invent provider prefixes.
  Do not map `auto` to an unverified target strategy.
  Do not silently upgrade or replace models.

  STOP and obtain approval for mappings and behavior changes.

  ## 5. Review headers and dependent features

  Inventory every Helicone-specific request header and response consumer.

  For each, identify its purpose and choose one of:

  - A verified NexLLM equivalent.
  - An approved application implementation.
  - An explicitly approved removal.
  - A blocker requiring the call site to remain on Helicone.

  Review authentication, target URLs, user tags, sessions, properties,
  prompt IDs, model overrides, fallback chains, response caches, cache
  seeds, retry controls, rate limits, context truncation, moderation,
  security filters, logging omissions, PostHog delivery, stream formats,
  and request IDs.

  Never treat prompt caching as an automatic substitute for response
  caching. Never assume channel groups replace fallback or privacy policies.

  Do not remove a generic `Cache-Control` header globally.
  Do not promise that Responses API state replaces observability traces.

  ## 6. Present the edit plan

  List affected files, credential changes, model mappings, header changes,
  feature gaps, tests, remaining integrations, and rollback steps.

  STOP and wait for approval.

  ## 7. Apply approved edits

  For compatible OpenAI SDK clients, keep the SDK and update centralized
  configuration where possible.

  Use the NexLLM key for migrated requests. Never forward provider or
  Helicone credentials to NexLLM.

  Remove Helicone authentication from migrated clients. Remove other
  Helicone-specific headers only after resolving their required behavior.

  Keep separate clients if some traffic remains on Helicone.

  Preserve asynchronous patterns, appropriate timeouts, transport settings,
  unrelated headers, and approved retry behavior.

  Review SDK wrappers and adapters before changing them. Ask before
  replacing dependencies or changing API formats.

  Update placeholder-only environment examples and setup documentation.
  Do not edit secret-bearing files or delete rollback credentials.

  ## 8. Reconnect observability

  Use NexLLM's documented usage logs for request time, API key name, model,
  timing, token consumption, and cost.

  Treat custom properties, user breakdowns, session traces, prompt
  versioning, cache alerts, and external analytics as separate requirements.

  Propose application instrumentation or a retained observability service
  where needed. Do not claim automatic historical log migration.

  Ask the user to arrange any required Helicone exports before
  deprovisioning. Do not export sensitive history without approval.

  ## 9. Generate and run verification tests

  Use the project's existing test framework where possible.

  Generate a smoke test that:

  - Reads credentials from the process environment.
  - Uses one approved model.
  - Sends only a short synthetic prompt.
  - Has a finite timeout and limited retries.
  - Checks response structure and nonempty text.
  - Exits unsuccessfully on failure.
  - Prints no credentials or sensitive request data.

  Test required streaming, tools, structured output, multimodal input,
  timeouts, error handling, retries, fallback behavior, privacy controls,
  logging, and usage accounting separately.

  Ask permission before live requests and explain that inference tests
  may incur charges.

  STOP and wait before running them.

  ## 10. Report and preserve rollback

  Summarize changes, model mappings, tests passed, tests not run,
  unresolved gaps, deployment tasks, and remaining Helicone integrations.

  A successful text completion proves only the tested path.

  Do not deploy, revoke credentials, delete historical logs, or remove
  rollback configuration without explicit approval.

  ## Troubleshooting

  - Authentication: confirm the intended NexLLM key is loaded, without
    printing it. Review expiration, quota, and IP restrictions.
  - Missing model: rediscover exact IDs with the inference key and inspect
    channel-group access and model restrictions.
  - Wrong endpoint: use `/v1` once, with no inherited Helicone target URL.
  - Missing telemetry: inspect application instrumentation and NexLLM logs;
    do not assume Helicone headers are interpreted.
  - Cache or retry differences: identify the old guarantee and test the
    approved replacement rather than silently deleting it.
  ```
</Accordion>

### 2. Install the skill

From the directory containing `migrate-helicone.md`, run these commands in a macOS, Linux, or WSL terminal:

```bash theme={null}
mkdir -p ~/.claude/skills/migrate-helicone

cp -i migrate-helicone.md \
  ~/.claude/skills/migrate-helicone/SKILL.md
```

`SKILL.md` is the installed filename for the same contents. You do not need a second script.

If a skill already exists under that name, the command asks before overwriting it. Review the existing file, especially if it targets a different migration provider.

### 3. Run the migration

Start Claude Code in your application's project directory and invoke:

```text theme={null}
/migrate-helicone
```

Or ask:

```text theme={null}
Use the migrate-helicone skill to migrate this application's
Helicone integration to NexLLM.
```

<Warning>
  Never paste API keys into the conversation. Configure credentials through your existing local environment or deployment secret store. Live inference tests may incur usage charges.
</Warning>

<Note>
  This skill migrates application code. It does not configure Claude Code itself to use NexLLM as its inference provider.
</Note>

You can also migrate manually using the steps below.

## Step 1: Update Your Environment Variables

Add NexLLM configuration for the clients you intend to migrate.

The following values are placeholders, not real credentials.

<CodeGroup>
  ```dotenv Before: provider proxy theme={null}
  OPENAI_API_KEY=your-existing-provider-key
  HELICONE_API_KEY=your-existing-helicone-key
  BASE_URL=https://oai.helicone.ai/v1
  ```

  ```dotenv Before: AI Gateway theme={null}
  HELICONE_API_KEY=your-existing-helicone-key
  BASE_URL=https://ai-gateway.helicone.ai
  ```

  ```dotenv After: NexLLM theme={null}
  NEXLLM_API_KEY=your-nexllm-key
  NEXLLM_BASE_URL=https://www.nexllm.ai/v1
  NEXLLM_MODEL=replace-with-an-approved-model-id
  ```
</CodeGroup>

`NEXLLM_BASE_URL` and `NEXLLM_MODEL` are application configuration conventions used in this guide. Your code must explicitly read them.

For local development, use your existing environment-loading mechanism. Creating a `.env` file does not automatically load its values.

For deployed applications, use your hosting platform or secret manager.

Migrated NexLLM requests authenticate with the NexLLM key, not your Helicone key or upstream provider key.

<Warning>
  Do not delete provider or Helicone keys globally. They may still be needed by unmigrated requests, logging integrations, or rollback. Keep them in the existing secret store, not commented into source code.
</Warning>

## Step 2: Update Your Client

Update the base URL and credentials for migrated call sites.

Before running these examples, complete the model discovery in Step 4 and set `NEXLLM_MODEL` to one approved Chat Completions model.

Preserve appropriate timeouts, retry settings, custom transports, and unrelated headers. Review existing gateway retry behavior before adding new retry layers.

<CodeGroup>
  ```python Python (OpenAI SDK) theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url=os.environ.get(
          "NEXLLM_BASE_URL",
          "https://www.nexllm.ai/v1",
      ),
      api_key=os.environ["NEXLLM_API_KEY"],
  )

  response = client.chat.completions.create(
      model=os.environ["NEXLLM_MODEL"],
      messages=[
          {"role": "user", "content": "Say hello in one word."}
      ],
  )

  print(response.choices[0].message.content)
  ```

  ```javascript JavaScript (OpenAI SDK) theme={null}
  import OpenAI from "openai";

  const apiKey = process.env.NEXLLM_API_KEY;
  const model = process.env.NEXLLM_MODEL;

  if (!apiKey || !model) {
    throw new Error(
      "Configure NEXLLM_API_KEY and an approved NEXLLM_MODEL"
    );
  }

  const client = new OpenAI({
    baseURL:
      process.env.NEXLLM_BASE_URL ??
      "https://www.nexllm.ai/v1",
    apiKey,
  });

  const response = await client.chat.completions.create({
    model,
    messages: [
      { role: "user", content: "Say hello in one word." },
    ],
  });

  console.log(response.choices[0].message.content);
  ```

  ```bash cURL theme={null}
  # Requires jq to construct JSON safely.
  : "${NEXLLM_API_KEY:?Configure NEXLLM_API_KEY first}"
  : "${NEXLLM_MODEL:?Choose an approved model first}"

  payload="$(jq -n \
    --arg model "$NEXLLM_MODEL" \
    '{
      model: $model,
      messages: [
        {role: "user", content: "Say hello in one word."}
      ]
    }')" && \
  curl --fail --silent --show-error \
    --max-time 60 \
    https://www.nexllm.ai/v1/chat/completions \
    -H "Authorization: Bearer $NEXLLM_API_KEY" \
    -H "Content-Type: application/json" \
    --data-binary "$payload"
  ```
</CodeGroup>

Remove `Helicone-Auth` from migrated clients. Resolve the behavior behind other Helicone headers using the next section.

If a shared client still serves Helicone requests, create separate client configurations instead of changing it globally.

## Step 3: Review and Remove Helicone Headers

Do not forward Helicone credentials to NexLLM.

For other Helicone-specific headers, first identify the feature they control. Remove them from migrated requests only after you have a verified replacement or have explicitly approved the behavior change.

The tables below describe migration actions. They do not assert that NexLLM has a one-to-one equivalent for every Helicone feature.

<Accordion title="Request header mapping">
  | Helicone configuration | NexLLM migration action |
  | - | - |
  | `Helicone-Auth` | Remove. Authenticate the migrated request with `Authorization: Bearer <NEXLLM_API_KEY>`. |
  | `Helicone-Target-URL`, `Helicone-OpenAI-Api-Base` | Remove from migrated requests. Use the NexLLM endpoint and an exact available model ID. Verify any required host or region constraints separately. |
  | `Helicone-User-Id` | Preserve user attribution in application telemetry. Separate named NexLLM keys can help distinguish workloads, but do not assume equivalent per-user analytics. |
  | `Helicone-Session-Id`, `Helicone-Session-Path`, `Helicone-Session-Name` | Preserve session and trace relationships in your application or observability system. API conversation state is not automatically an observability trace replacement. |
  | `Helicone-Property-[Name]` | Preserve required custom properties outside the request unless a documented NexLLM mapping is confirmed. |
  | `Helicone-Prompt-Id` | Retain prompt identifiers and version tracking in your application or prompt-management system. |
  | `Helicone-Model-Override` | Select an actual available model through `model`. Check any accounting behavior that depended on the old override. |
  | `Helicone-Fallbacks` | Verify a supported replacement or implement an approved application-level fallback policy. Do not assume the header configures NexLLM routing. |
  | `Helicone-Cache-Enabled` | Review response-cache requirements separately from model prompt caching. Do not assume equivalent cache hits or output reuse. |
  | `Cache-Control` used with Helicone caching | Review in context. This is a general HTTP header; do not delete unrelated uses globally. |
  | `Helicone-Cache-Seed` | Preserve cache isolation requirements in the approved caching implementation. Do not invent a NexLLM seed parameter. |
  | `Helicone-Cache-Bucket-Max-Size` | Recreate required cache-capacity controls in your chosen cache implementation. |
  | `Helicone-Retry-Enabled`, `helicone-retry-num`, `helicone-retry-factor`, `helicone-retry-min-timeout`, `helicone-retry-max-timeout` | Configure and test the SDK or application retry policy. Do not assume gateway retries or automatic failover replace it. |
  | `Helicone-RateLimit-Policy` | Review request-rate controls separately from key quota and access restrictions. Enforce application-side rate limits if needed. |
  | `Helicone-Token-Limit-Exception-Handler` | Preserve approved truncation, summarization, or context-overflow handling in application code. Do not silently shorten prompts. |
  | `Helicone-Moderations-Enabled`, `Helicone-LLM-Security-Enabled` | Verify moderation and security replacements before moving affected traffic. Do not remove required safeguards merely to make requests succeed. |
  | `Helicone-Omit-Request`, `Helicone-Omit-Response` | Confirm logging and retention requirements with NexLLM before migration. Removing these headers does not establish equivalent privacy behavior. |
  | `Helicone-Posthog-Key`, `Helicone-Posthog-Host` | Reconnect analytics through application instrumentation or a verified integration. Do not forward analytics credentials unnecessarily. |
  | `Helicone-Stream-Force-Format` | Test NexLLM's Chat Completions SSE stream with your parser or SDK. |
  | `Helicone-Request-Id` | Preserve an application-generated correlation ID locally. Verify any supported request-header mapping rather than assuming it is echoed. |
</Accordion>

<Accordion title="Response header mapping">
  | Helicone response header | Migration action |
  | - | - |
  | `Helicone-Id` | Review consumers of this header. NexLLM's Chat Completions body includes an `id`, but do not assume it has identical correlation semantics. |
  | `Helicone-Cache` | Replace header-based cache metrics with signals from your verified cache implementation. |
  | `Helicone-Cache-Bucket-Idx` | Remove dependent parsing only after replacing any required cache diagnostics. |
  | `Helicone-Fallback-Index` | Instrument the approved fallback implementation. Do not assume NexLLM returns an equivalent header. |
  | `Helicone-RateLimit-Limit`, `Helicone-RateLimit-Remaining`, `Helicone-RateLimit-Policy` | Verify actual response headers and rate-limit behavior before changing parsers or alerts. |
</Accordion>

If alerts depend on a Helicone response header, update and test those alerts before rollout.

<Warning>
  Keep affected requests on the existing integration if required privacy, moderation, provider-selection, or reliability behavior has not been replaced and approved.
</Warning>

## Step 4: Update Model Identifiers

Discover the models available to the NexLLM key you will use for inference:

```bash theme={null}
curl --fail --silent --show-error \
  --max-time 30 \
  https://www.nexllm.ai/v1/models \
  -H "Authorization: Bearer $NEXLLM_API_KEY"
```

The response contains model objects in `data`. Use the exact `data[].id` value as the request's `model`.

NexLLM's documentation includes these example identifiers:

* `gpt-4o`
* `aws/claude-haiku-4-5`
* `gemini-2.5-flash`

These examples do not guarantee that all three are available to your key.

### Provider-proxy models

A model ID used through a provider proxy might also appear in NexLLM's catalog. Keep it unchanged only after confirming its availability and required behavior.

### AI Gateway model identifiers

Do not assume an `author/model` identifier works unchanged. Do not automatically remove its prefix or replace it with a guessed provider prefix.

Build a mapping for each model actually used by your application:

| Existing Helicone model | Proposed NexLLM model | Required features | Status |
| - | - | - | - |
| Existing configured ID | Exact discovered ID | Streaming, tools, or other requirements | Pending verification |

Changing a model identifier does not by itself establish equivalent hosting, regional processing, retention, or contractual guarantees.

### Channel groups and model restrictions

Each NexLLM key belongs to a channel group that controls model access and pricing. Keys may also have explicit model restrictions.

Use the same key for discovery and inference. If a model is inaccessible, check both its channel-group availability and the key's restrictions.

Do not silently switch model versions or select a replacement.

### What about `auto`?

This guide does not establish a NexLLM equivalent to Helicone's automatic model-selection behavior.

Choose an explicit available model, implement an approved selection policy, or keep that call site on Helicone until a suitable alternative has been verified.

## Step 5: Reconnect Observability

NexLLM documents per-request usage logs and dashboard statistics. Use these for the information they provide, and preserve additional application telemetry where needed.

| Existing observability requirement | NexLLM migration approach |
| - | - |
| Request history | Review NexLLM Usage Logs for new requests. |
| Request timing | Use the documented timing field and application-side measurements where needed. |
| Token and cost tracking | Review token consumption and cost in Usage Logs. |
| Workload attribution | Use descriptive API key names and retain application-level workload tags. |
| User-level breakdown | Preserve user attribution in your application or analytics system unless a verified equivalent is available. |
| Custom properties | Keep required properties in application telemetry. |
| Session traces | Preserve trace and span relationships in your observability system. |
| Prompt versions | Retain the prompt registry or application version metadata. |
| Cache hit rate | Measure the cache implementation actually used after migration. |
| Fallback events | Instrument your approved fallback policy. |
| Rate-limit monitoring | Monitor actual API errors and application-side rate limits. |

NexLLM Usage Logs document request time, API key name, model name, timing, input/output token consumption, and cost. Dashboard statistics include requests, quota consumption, token usage, RPM, and TPM.

Do not assume those statistics replace every Helicone dashboard view.

### Exporting Your Helicone History

Changing the inference endpoint does not migrate historical logs.

Before deprovisioning Helicone:

1. Identify the history needed for your audit, debugging, or reporting requirements.
2. Review Helicone's current export options for your account or deployment.
3. Export the required records to approved storage.
4. Check completeness, access controls, and retention requirements.
5. Keep access to the original system until the archive is validated.

This guide does not provide a Helicone-to-NexLLM historical import workflow.

## Step 6 (Optional): Evaluate the Responses API

NexLLM's API overview lists `POST /v1/responses` as an OpenAI Responses API-compatible endpoint.

You do not need to adopt it to migrate an existing Chat Completions integration. Treat changing API families as a separate project.

Before using Responses in production, verify the chosen model's support for the features you need, including streaming, tools, structured output, multimodal input, state, storage, and response parsing.

<Note>
  Model discovery establishes catalog availability, not support for every endpoint or feature. Do not assume `previous_response_id`, web search, or other advanced features work for every available model.
</Note>

Set `NEXLLM_RESPONSES_MODEL` to a model whose Responses support you have confirmed. This is a suggested application variable, not a provider-defined setting.

The following are minimal starting requests for that verification.

<CodeGroup>
  ```bash cURL theme={null}
  # Requires jq.
  : "${NEXLLM_API_KEY:?Configure NEXLLM_API_KEY first}"
  : "${NEXLLM_RESPONSES_MODEL:?Choose a verified Responses model}"

  payload="$(jq -n \
    --arg model "$NEXLLM_RESPONSES_MODEL" \
    '{
      model: $model,
      input: "What is the capital of France?"
    }')" && \
  curl --fail --silent --show-error \
    --max-time 60 \
    https://www.nexllm.ai/v1/responses \
    -H "Authorization: Bearer $NEXLLM_API_KEY" \
    -H "Content-Type: application/json" \
    --data-binary "$payload"
  ```

  ```python Python theme={null}
  import os
  import requests

  response = requests.post(
      "https://www.nexllm.ai/v1/responses",
      headers={
          "Authorization": f"Bearer {os.environ['NEXLLM_API_KEY']}",
          "Content-Type": "application/json",
      },
      json={
          "model": os.environ["NEXLLM_RESPONSES_MODEL"],
          "input": "What is the capital of France?",
      },
      timeout=60,
  )

  response.raise_for_status()
  print(response.json())
  ```

  ```javascript JavaScript theme={null}
  const apiKey = process.env.NEXLLM_API_KEY;
  const model = process.env.NEXLLM_RESPONSES_MODEL;

  if (!apiKey || !model) {
    throw new Error(
      "Configure NEXLLM_API_KEY and NEXLLM_RESPONSES_MODEL"
    );
  }

  // Requires a runtime supporting fetch and AbortSignal.timeout.
  const response = await fetch(
    "https://www.nexllm.ai/v1/responses",
    {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
        model,
        input: "What is the capital of France?",
      }),
      signal: AbortSignal.timeout(60000),
    }
  );

  if (!response.ok) {
    throw new Error(`Responses request failed: ${response.status}`);
  }

  console.log(await response.json());
  ```
</CodeGroup>

Do not reuse Chat Completions response parsing without adapting it to the Responses format.

## Step 7: Verify Before Switching Production Traffic

Start with one approved model and a short synthetic prompt.

Check that:

* The process loads the intended NexLLM credentials.
* The selected model is available to that key.
* The request reaches NexLLM rather than a Helicone endpoint.
* No migrated request forwards a Helicone or provider key.
* The response has the expected structure.
* Application errors and timeouts are handled correctly.
* New requests appear in the expected usage logs.
* Required privacy and security behavior remains intact.
* Required application features pass their own tests.

A successful text completion verifies basic connectivity, not full feature parity.

### Minimal Python smoke test

This test uses the OpenAI Python SDK. Configure `NEXLLM_API_KEY` and `NEXLLM_MODEL` through your existing environment.

```python theme={null}
import os
import sys
from openai import OpenAI

def main():
    api_key = os.environ.get("NEXLLM_API_KEY")
    model = os.environ.get("NEXLLM_MODEL")

    if not api_key or not model:
        print("FAIL: configure NEXLLM_API_KEY and NEXLLM_MODEL")
        return 1

    client = OpenAI(
        base_url="https://www.nexllm.ai/v1",
        api_key=api_key,
        timeout=30.0,
        max_retries=0,
    )

    try:
        response = client.chat.completions.create(
            model=model,
            messages=[
                {"role": "user", "content": "Reply with one short word."}
            ],
        )

        if not response.choices:
            print("FAIL: no completion choices")
            return 1

        content = response.choices[0].message.content
        if not isinstance(content, str) or not content.strip():
            print("FAIL: expected nonempty text")
            return 1

    except Exception:
        # Do not print request headers, credentials, or raw error bodies.
        print("FAIL: request failed; inspect sanitized diagnostics")
        return 1

    print("PASS: basic text completion")
    return 0

if __name__ == "__main__":
    sys.exit(main())
```

### Test streaming

NexLLM Chat Completions supports Server-Sent Events when `stream=True`.

Using your configured client and an approved streaming-capable model:

```python theme={null}
stream = client.chat.completions.create(
    model=os.environ["NEXLLM_MODEL"],
    messages=[
        {"role": "user", "content": "Write a one-sentence greeting."}
    ],
    stream=True,
)

for chunk in stream:
    if not chunk.choices:
        continue

    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)

print()
```

Also test cancellation, stream errors, and completion handling if your application depends on them.

### Test required behavior separately

Create focused tests for:

* Tool calls and tool-result round trips.
* Structured output and schema validation.
* Images or other multimodal inputs.
* Retry limits and timeout handling.
* Fallback selection and failure cases.
* Cache isolation and expiry.
* Context-overflow handling.
* Moderation and logging restrictions.
* Telemetry, alerts, and usage accounting.

Run live inference tests only with approval; they may incur charges.

## Why Migrate to NexLLM?

### Keep your existing Chat Completions client

For compatible integrations, continue using the OpenAI SDK or your existing HTTP client with NexLLM configuration.

### Access multiple model families

NexLLM provides access to GPT, Claude, and Gemini models through a single API base URL. Your actual access depends on your credentials and channel group.

### Choose model access and pricing through channel groups

Channel groups determine model access and the pricing ratio applied to usage.

Review the current group and pricing documentation rather than assuming a universal markup, discount, or price match.

### Configure scoped API keys

NexLLM documents key expiration, quota settings, model restrictions, and IP allowlists.

These controls should be reviewed against your application's access requirements. They are not automatically equivalent to a hierarchical organization budgeting system or a request-rate policy.

### Monitor request usage

Use NexLLM Usage Logs and dashboard statistics to inspect model activity, token consumption, timing, and cost.

Keep separate instrumentation for application-specific traces and tags.

### Evaluate additional API formats when needed

NexLLM's API overview also lists Responses and Claude Messages endpoints.

Verify endpoint and model compatibility before changing request formats. Their presence does not establish feature parity across all models.

### Understand the trade-offs

This migration does not promise automatic replacements for:

* Helicone's custom properties and session tracing.
* Gateway response caching.
* Prompt management.
* Specific automatic routing or fallback policies.
* Existing moderation and security rules.
* Provider-specific retention or regional processing guarantees.

Resolve these requirements before migrating affected production traffic.

## Troubleshooting

### Model not found or inaccessible

Query `GET /v1/models` with the same key used for inference.

Check the exact model ID, channel group, and any explicit model restrictions. Do not guess a prefix or silently substitute a model.

A model missing from one key's catalog is not necessarily unavailable platform-wide.

### Authentication failure

Confirm that the process loads a NexLLM key and sends Bearer authentication to the NexLLM endpoint.

Check for stale secrets, accidental whitespace, expiration, exhausted quota, and IP restrictions where applicable.

Never print the key while debugging.

### Requests succeed but do not appear in NexLLM logs

Confirm that the actual runtime client points to NexLLM rather than a Helicone endpoint or another configured proxy.

Check that you are viewing the correct NexLLM account and time range.

If Helicone remains configured only for asynchronous logging, inspect that separate integration rather than assuming the inference URL controls all telemetry.

### Cache hit rate changed

Determine whether the old metric measured a gateway response cache, provider prompt caching, or an application cache.

Do not compare those as though they were interchangeable.

Rebuild required cache metrics around the implementation actually used after migration.

### Connection or endpoint errors

Use this SDK base URL:

```text theme={null}
https://www.nexllm.ai/v1
```

The Chat Completions request URL is:

```text theme={null}
https://www.nexllm.ai/v1/chat/completions
```

Do not duplicate `/v1`, retain an old target URL, or invent a health endpoint. Use the documented Models endpoint to test authenticated API access without generating a completion.

### Retry or rate-limit settings stopped applying

Review the old Helicone request headers and the behavior they controlled.

Configure an approved SDK or application policy and test failure cases. Do not assume key quota settings replace requests-per-minute controls.

### Custom properties or sessions disappeared

Review the application instrumentation that previously populated Helicone headers.

Preserve those values in your tracing or analytics system unless you have confirmed a supported NexLLM mapping.

### Unsupported parameter

Reduce the request to an approved `model` and a simple `messages` array to isolate the problem.

Reintroduce required options individually. Permanently removing a required feature needs a separate approval.

### Environment variable not found

Confirm the application process receives `NEXLLM_API_KEY` and `NEXLLM_MODEL`.

If using a `.env` file, confirm the application or development tooling loads it. Deployment secrets must be configured in the deployed environment separately.

### A provider-native client stopped working

Check whether the original integration used Chat Completions, Claude Messages, or another native request format.

Do not send native payloads to a Chat Completions endpoint. Use the corresponding documented endpoint and verify authentication, model support, and response parsing separately.

## Next Steps

* [Authentication](/api-reference/authentication): Review credential handling.
* [Models](/api-reference/models): Discover exact model identifiers.
* [Channel Groups](/channel-groups/overview): Review model access and pricing categories.
* [API Keys](/get-started/getting-your-api-keys): Configure key restrictions.
* [Usage & Logs](/account/usage-logs-and-statistics): Check request activity and consumption.
* [Chat Completions](/api-reference/chat-completions): Review request fields and streaming.
* [API Overview](/api-reference/overview): Review other documented endpoints.

Deploy through your normal rollout process after validation. Retire unused Helicone or provider credentials only when no remaining integration or rollback plan depends on them.

## Feedback

If a migration requirement does not translate cleanly, record:

* The affected integration mode.
* The feature or header involved.
* The endpoint and model identifier.
* The expected behavior.
* A minimal reproduction using synthetic data.
* Sanitized status codes and relevant diagnostic details.

Share the report through your existing NexLLM support contact.

Do not include API keys, authorization headers, private prompts, or sensitive historical logs.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.