Skip to main content

Migrate from TensorZero to NexLLM: update authentication, clients, model identifiers, application workflows, and observability.

ove your TensorZero inference integration to NexLLM using an OpenAI-compatible API. This guide covers both TensorZero’s OpenAI-compatible interface and its native inference client. It also explains how to preserve the behavior behind functions, variants, episodes, and feedback.
Migrating inference requests is separate from migrating experimentation, evaluation, and workflow state. Keep those behaviors in your application or existing tooling unless a suitable NexLLM replacement has been verified.

At a Glance

See API Reference and Authentication for NexLLM connection details.

Prerequisites

1

Create a NexLLM account and API key

Create a dedicated migration key. Review its quota, expiration, model restrictions, IP allowlist, and channel group.See API Keys.
2

Confirm model access

Use the Models API to check which model IDs are available to your key. Availability depends on its assigned channel group.See Models API.
3

Review your TensorZero integration

Locate gateway URLs, native clients, model references, prompt templates, schemas, variants, episodes, feedback calls, caching, and tracing.
4

Keep a rollback path

Preserve your existing deployment, credentials, and configuration until the migrated integration passes end-to-end tests.

Quick Start for Claude Code Users

You can use the following prompt in a Claude Code session opened in your repository. This is a migration prompt, not an official NexLLM migration skill.

Step 1: Update Your Environment Variables

Configure your NexLLM API key and the OpenAI-compatible base URL:
BASE_URL is an application configuration variable. Your client must explicitly read it, or use the NexLLM base URL directly as shown below. NexLLM uses bearer authentication for the OpenAI-compatible requests in this guide:
See Authentication.
Replace placeholder keys such as not-used with a real NexLLM key. Do not reuse a TensorZero gateway key or an upstream provider key as your NexLLM credential.
Keep old provider credentials securely available for rollback. Remove them only after confirming that no remaining workload needs them. Never commit credentials to source control or expose them in browser code.

Step 2: Update Your Client

For an OpenAI-compatible integration, update the base URL, API key, and model identifier. For a native TensorZero integration, replace the client call and translate the payload into the selected NexLLM API format. The examples use gpt-4o, which appears in NexLLM’s documentation. Confirm availability with your key before running them. See Quickstart, Chat Completions, and Models API.

OpenAI-Compatible Integration

Native TensorZero Client

Replace native inference calls rather than simply changing the gateway URL.
For more complex native requests:
  • Render templated inputs before constructing messages.
  • Preserve system instructions and conversation history.
  • Translate generation parameters individually.
  • Update native response parsers.
  • Test tool calls, structured output, multimodal content, and streaming separately.
A successful text request does not establish compatibility for every native inference feature. Preserve the original behavior in regression tests.

Step 3: Remove TensorZero-Specific Headers and Body Params

Review fields sent through extra_body, native inference arguments, and custom headers. The mappings below are recommended migration actions. They do not claim that undocumented NexLLM equivalents are unavailable; they identify behavior that must be preserved or verified separately.
See Authentication.
On TensorZero’s OpenAI-compatible interface, these fields may use the tensorzero:: prefix. Native inference calls may use the corresponding unprefixed names.See Chat Completions and Pricing and Caching.
See Chat Completions and Usage & Logs.
Removing dryrun, tracing, or metadata fields does not preserve their original behavior. Resolve privacy, retention, and audit requirements before sending production data.

Step 4: Decompose Functions, Variants, and Episodes

A function reference can represent more than a model choice. Review the prompts, schemas, parameters, experiment assignment, and output handling behind each call.

Move Prompt Templates into Application Code

For example, replace a named summarization function with explicit prompt construction:
Use the resulting messages with the NexLLM client from Step 2. Preserve any original input schemas, output validation, tools, and prompt-version tracking.

Preserve Experiment Assignment

For multi-step workflows, keep the experiment assignment stable across related calls. Record:
  • Experiment and variant names.
  • Prompt version.
  • Requested model.
  • Generation parameters.
  • Workflow ID.
  • Evaluation outcomes.
Changing only the model does not recreate a variant that also changed prompts or validation.

Keep Episodes Separate from Conversation History

Use your application’s workflow ID to group related requests. For Chat Completions, supply the conversation messages needed for each call. Do not assume that an episode ID retrieves prior messages, or that a response ID reproduces episode-level feedback relationships. See Chat Completions.

Configure Channel Groups Separately

NexLLM assigns each API key to one channel group. That group determines model access and the pricing ratio applied to usage. Choose the appropriate group for your workload, but do not treat a channel group as a replacement for experiment assignment or function configuration. See Channel Groups.

Step 5: Update Model Identifiers

Use exact model IDs returned by NexLLM rather than mechanically translating TensorZero namespaces.
The response contains model identifiers in data[].id. The returned models depend on the channel group assigned to your key. See Models API.

Example Mappings

Do not replace every :: with /.NexLLM documents both unprefixed and prefixed identifiers. Use the exact returned ID; do not invent a provider prefix or assume that another gateway’s naming conventions apply.
If provider hosting, region, or data handling is important to your application, verify those requirements independently of the model name.

Step 6: Reconnect Observability

NexLLM Usage Logs report request time, API key name, model, timing, token consumption, and cost or quota consumption. Use those records alongside your application telemetry. See Usage & Logs. See API Keys.

Exporting Your TensorZero History

Before retiring the gateway:
  1. Back up the databases used by your TensorZero deployment, including ClickHouse where applicable.
  2. Archive tensorzero.toml, prompt templates, schemas, and evaluation configuration.
  3. Preserve identifiers that join feedback to historical inferences.
  4. Include any associated object storage.
  5. Test restoration or read access.
  6. Document the required retention period.
Do not assume historical records will automatically appear in NexLLM.

Step 7 (Optional): Evaluate the Responses API

NexLLM lists an OpenAI Responses-compatible endpoint:
See API Reference. Check a model’s API compatibility in its details before selecting it for Responses requests. The model browser includes endpoint-type filtering. See Model Details and Pricing. The following is a minimal request pattern. Replace the model placeholder with a Responses-compatible model available to your key:
Endpoint availability does not guarantee feature parity with TensorZero’s native inference API.Verify streaming, tools, structured output, multimodal input, storage, and response chaining for your selected model. Do not assume previous_response_id replaces episode grouping or feedback attribution.
Update response parsing for the chosen API. Do not reuse Chat Completions-specific field access without checking the response format.

Why Migrate to NexLLM

One Integration for Multiple Model Families

NexLLM provides access to GPT, Claude, and Gemini through an OpenAI-compatible interface. See Introduction.

Channel-Based Model Access and Pricing

Choose a channel group that includes the models you need, then review its pricing ratio and connectivity characteristics. See Channel Groups.

Configurable API Keys

Create purpose-specific credentials with documented access and quota controls. See API Keys.

Usage Visibility

Review request-level logs and dashboard statistics after migration. See Usage & Logs.

A Direct Inference Path

You can point migrated clients directly at NexLLM. Retire the old gateway only after moving or preserving every workflow that still depends on it.
This migration does not assume built-in replacements for TensorZero’s function configuration, A/B testing, feedback optimization, or tracing customization. Treat these as explicit migration work.

Troubleshooting

Query GET /v1/models with the same key used by the application.Check:
  • The exact model ID.
  • The key’s channel group.
  • Model restrictions.
  • Remaining TensorZero namespaces or aliases.
Do not assume changing separators produces a valid NexLLM identifier.See Models API.
Confirm that the application sends a NexLLM key rather than a placeholder, TensorZero gateway key, or upstream provider credential.Check expiration, IP restrictions, stale deployment secrets, and accidental whitespace.See Authentication.
Inspect the actual destination URL and credential used by the request.Check background workers, scheduled jobs, and independently initialized clients. Some may still target the old gateway.See Usage & Logs.
Compare the entire original configuration:
  • Prompt rendering.
  • System instructions.
  • Model selection.
  • Generation parameters.
  • Tool definitions.
  • Input and output validation.
  • Experiment assignment.
  • Retry and fallback policy.
Do not send a TensorZero function reference as the NexLLM model ID.
Keep workflow IDs and feedback relationships in your application.Record each returned completion ID alongside that context. Do not assume an individual response ID replaces the complete workflow grouping.Keep the original feedback pipeline available until its replacement has been validated.
Review the selected model, prompt structure, request format, cache rules, and channel-group pricing ratio.Do not assume TensorZero cache entries or expiration settings transfer to another gateway.See Pricing and Caching.
Check whether the application still sends a TensorZero-native body.Translate it into the selected NexLLM endpoint’s format. Render templates before sending messages, and test tools or multimodal content separately.See Chat Completions.
Use this OpenAI SDK base URL:
Remove the previous /openai/v1 or /inference path. Avoid adding /v1 twice when constructing HTTP URLs.Test the documented Models endpoint:
Then test inference separately. Do not assume another gateway’s health-check path exists on NexLLM.See API Reference.

Migration Checklist

  • Create a dedicated NexLLM key.
  • Confirm model access and channel-group settings.
  • Update authentication and client URLs.
  • Replace TensorZero model and function references.
  • Translate native inference payloads and response parsing.
  • Review custom headers and body parameters.
  • Preserve prompts, schemas, and validation.
  • Preserve experiment assignment and fallbacks.
  • Retain workflow IDs and feedback relationships.
  • Verify caching, logging, and retention requirements.
  • Test streaming, tools, and multimodal requests where used.
  • Reconnect logs, traces, and consumption monitoring.
  • Back up historical data and configuration.
  • Compare response quality, latency, reliability, and cost.
  • Roll out gradually with a tested rollback path.
  • Remove unused infrastructure and credentials after acceptance.

Next Steps

API Reference

Review NexLLM endpoints and API formats.

Available Models

Discover model IDs available to your key.

Channel Groups

Configure model access and pricing.

Usage & Logs

Monitor requests after migration.

Feedback

When reporting a migration issue, include:
  • The behavior you need to preserve.
  • The client library and version.
  • The endpoint and model ID.
  • The approximate request time.
  • The HTTP status and sanitized error.
  • A completion or request identifier, if available.
  • A minimal reproduction.
Remove API keys, provider credentials, private prompts, and personal data before sharing diagnostics.