Migrate from TensorZero to NexLLM: update authentication, clients, model identifiers, application workflows, and observability.
ove your TensorZero inference integration to NexLLM using an OpenAI-compatible API. This guide covers both TensorZero’s OpenAI-compatible interface and its native inference client. It also explains how to preserve the behavior behind functions, variants, episodes, and feedback.Migrating inference requests is separate from migrating experimentation, evaluation, and workflow state. Keep those behaviors in your application or existing tooling unless a suitable NexLLM replacement has been verified.
At a Glance
See API Reference and Authentication for NexLLM connection details.
Prerequisites
1
Create a NexLLM account and API key
Create a dedicated migration key. Review its quota, expiration, model restrictions, IP allowlist, and channel group.See API Keys.
2
Confirm model access
Use the Models API to check which model IDs are available to your key. Availability depends on its assigned channel group.See Models API.
3
Review your TensorZero integration
Locate gateway URLs, native clients, model references, prompt templates, schemas, variants, episodes, feedback calls, caching, and tracing.
4
Keep a rollback path
Preserve your existing deployment, credentials, and configuration until the migrated integration passes end-to-end tests.
Quick Start for Claude Code Users
You can use the following prompt in a Claude Code session opened in your repository. This is a migration prompt, not an official NexLLM migration skill.Step 1: Update Your Environment Variables
Configure your NexLLM API key and the OpenAI-compatible base URL:BASE_URL is an application configuration variable. Your client must explicitly read it, or use the NexLLM base URL directly as shown below.
NexLLM uses bearer authentication for the OpenAI-compatible requests in this guide:
Step 2: Update Your Client
For an OpenAI-compatible integration, update the base URL, API key, and model identifier. For a native TensorZero integration, replace the client call and translate the payload into the selected NexLLM API format. The examples usegpt-4o, which appears in NexLLM’s documentation. Confirm availability with your key before running them.
See Quickstart, Chat Completions, and Models API.
OpenAI-Compatible Integration
- Python
- JavaScript
- cURL
Native TensorZero Client
Replace native inference calls rather than simply changing the gateway URL.- Render templated inputs before constructing messages.
- Preserve system instructions and conversation history.
- Translate generation parameters individually.
- Update native response parsers.
- Test tool calls, structured output, multimodal content, and streaming separately.
Step 3: Remove TensorZero-Specific Headers and Body Params
Review fields sent throughextra_body, native inference arguments, and custom headers.
The mappings below are recommended migration actions. They do not claim that undocumented NexLLM equivalents are unavailable; they identify behavior that must be preserved or verified separately.
Request header mapping
Request header mapping
See Authentication.
Request body parameter mapping
Request body parameter mapping
On TensorZero’s OpenAI-compatible interface, these fields may use the
tensorzero:: prefix. Native inference calls may use the corresponding unprefixed names.See Chat Completions and Pricing and Caching.
Response field mapping
Response field mapping
See Chat Completions and Usage & Logs.
Step 4: Decompose Functions, Variants, and Episodes
A function reference can represent more than a model choice. Review the prompts, schemas, parameters, experiment assignment, and output handling behind each call.Move Prompt Templates into Application Code
For example, replace a named summarization function with explicit prompt construction:Preserve Experiment Assignment
For multi-step workflows, keep the experiment assignment stable across related calls. Record:- Experiment and variant names.
- Prompt version.
- Requested model.
- Generation parameters.
- Workflow ID.
- Evaluation outcomes.
Keep Episodes Separate from Conversation History
Use your application’s workflow ID to group related requests. For Chat Completions, supply the conversation messages needed for each call. Do not assume that an episode ID retrieves prior messages, or that a response ID reproduces episode-level feedback relationships. See Chat Completions.Configure Channel Groups Separately
NexLLM assigns each API key to one channel group. That group determines model access and the pricing ratio applied to usage. Choose the appropriate group for your workload, but do not treat a channel group as a replacement for experiment assignment or function configuration. See Channel Groups.Step 5: Update Model Identifiers
Use exact model IDs returned by NexLLM rather than mechanically translating TensorZero namespaces.data[].id. The returned models depend on the channel group assigned to your key.
See Models API.
Example Mappings
If provider hosting, region, or data handling is important to your application, verify those requirements independently of the model name.
Step 6: Reconnect Observability
NexLLM Usage Logs report request time, API key name, model, timing, token consumption, and cost or quota consumption. Use those records alongside your application telemetry. See Usage & Logs.
See API Keys.
Exporting Your TensorZero History
Before retiring the gateway:- Back up the databases used by your TensorZero deployment, including ClickHouse where applicable.
- Archive
tensorzero.toml, prompt templates, schemas, and evaluation configuration. - Preserve identifiers that join feedback to historical inferences.
- Include any associated object storage.
- Test restoration or read access.
- Document the required retention period.
Step 7 (Optional): Evaluate the Responses API
NexLLM lists an OpenAI Responses-compatible endpoint:Why Migrate to NexLLM
One Integration for Multiple Model Families
NexLLM provides access to GPT, Claude, and Gemini through an OpenAI-compatible interface. See Introduction.Channel-Based Model Access and Pricing
Choose a channel group that includes the models you need, then review its pricing ratio and connectivity characteristics. See Channel Groups.Configurable API Keys
Create purpose-specific credentials with documented access and quota controls. See API Keys.Usage Visibility
Review request-level logs and dashboard statistics after migration. See Usage & Logs.A Direct Inference Path
You can point migrated clients directly at NexLLM. Retire the old gateway only after moving or preserving every workflow that still depends on it.This migration does not assume built-in replacements for TensorZero’s function configuration, A/B testing, feedback optimization, or tracing customization. Treat these as explicit migration work.
Troubleshooting
Model not found
Model not found
Query
GET /v1/models with the same key used by the application.Check:- The exact model ID.
- The key’s channel group.
- Model restrictions.
- Remaining TensorZero namespaces or aliases.
Invalid API key or authentication failure
Invalid API key or authentication failure
Confirm that the application sends a NexLLM key rather than a placeholder, TensorZero gateway key, or upstream provider credential.Check expiration, IP restrictions, stale deployment secrets, and accidental whitespace.See Authentication.
Requests succeed but do not appear in NexLLM logs
Requests succeed but do not appear in NexLLM logs
Inspect the actual destination URL and credential used by the request.Check background workers, scheduled jobs, and independently initialized clients. Some may still target the old gateway.See Usage & Logs.
Functions or variants behave differently
Functions or variants behave differently
Compare the entire original configuration:
- Prompt rendering.
- System instructions.
- Model selection.
- Generation parameters.
- Tool definitions.
- Input and output validation.
- Experiment assignment.
- Retry and fallback policy.
Episodes or feedback are missing
Episodes or feedback are missing
Keep workflow IDs and feedback relationships in your application.Record each returned completion ID alongside that context. Do not assume an individual response ID replaces the complete workflow grouping.Keep the original feedback pipeline available until its replacement has been validated.
Cache hit rate or cost changed
Cache hit rate or cost changed
Review the selected model, prompt structure, request format, cache rules, and channel-group pricing ratio.Do not assume TensorZero cache entries or expiration settings transfer to another gateway.See Pricing and Caching.
Native inference payloads are rejected
Native inference payloads are rejected
Check whether the application still sends a TensorZero-native body.Translate it into the selected NexLLM endpoint’s format. Render templates before sending messages, and test tools or multimodal content separately.See Chat Completions.
Connection errors or unexpected 404 responses
Connection errors or unexpected 404 responses
Use this OpenAI SDK base URL:Remove the previous Then test inference separately. Do not assume another gateway’s health-check path exists on NexLLM.See API Reference.
/openai/v1 or /inference path. Avoid adding /v1 twice when constructing HTTP URLs.Test the documented Models endpoint:Migration Checklist
- Create a dedicated NexLLM key.
- Confirm model access and channel-group settings.
- Update authentication and client URLs.
- Replace TensorZero model and function references.
- Translate native inference payloads and response parsing.
- Review custom headers and body parameters.
- Preserve prompts, schemas, and validation.
- Preserve experiment assignment and fallbacks.
- Retain workflow IDs and feedback relationships.
- Verify caching, logging, and retention requirements.
- Test streaming, tools, and multimodal requests where used.
- Reconnect logs, traces, and consumption monitoring.
- Back up historical data and configuration.
- Compare response quality, latency, reliability, and cost.
- Roll out gradually with a tested rollback path.
- Remove unused infrastructure and credentials after acceptance.
Next Steps
API Reference
Review NexLLM endpoints and API formats.
Available Models
Discover model IDs available to your key.
Channel Groups
Configure model access and pricing.
Usage & Logs
Monitor requests after migration.
Feedback
When reporting a migration issue, include:- The behavior you need to preserve.
- The client library and version.
- The endpoint and model ID.
- The approximate request time.
- The HTTP status and sanitized error.
- A completion or request identifier, if available.
- A minimal reproduction.