Migrate from Helicone to NexLLM
Move your Helicone Chat Completions integration to NexLLM by updating your base URL, API key, and model identifiers. NexLLM supports the OpenAI Chat Completions format, so you can retain your existing OpenAI SDK or HTTP client. Helicone-specific headers, observability, caching, and routing policies need separate review. Choose your existing integration mode below.These examples cover OpenAI-compatible Chat Completions. If you use a provider-native API, an asynchronous logging integration, or a Helicone-specific SDK, review that integration separately rather than applying a base-URL replacement blindly.
- Proxy mode
- AI Gateway
In provider-proxy mode, your existing integration sends a provider API key to Helicone’s provider proxy and a separate
Helicone-Auth header for Helicone authentication.For migrated requests, use your NexLLM key instead and remove the Helicone authentication header.Prerequisites
You will need:- A NexLLM account and API key. Create a key in the NexLLM dashboard.
- Access to your intended models. Review the key’s channel group and model restrictions.
- An existing Helicone integration. Identify whether it uses provider-proxy mode, AI Gateway mode, or a separate logging integration.
- An inventory of required behavior. Include caching, retries, fallbacks, moderation, privacy controls, request tagging, sessions, and downstream analytics.
- A rollback plan. Keep the existing configuration and credentials secure until the migrated paths have been validated.
NEXLLM_API_KEY is the environment-variable name used in this guide. Your application must read that same name.
Quick Start for Claude Code Users
The NexLLM migration skill instructs Claude Code to inspect your integration, propose model mappings, review Helicone-specific behavior, and generate verification tests. It asks for approval before editing and does not silently remove privacy, reliability, or observability requirements.1. Copy and review the migration skill
Expand the section below and copy the complete code block into a local file namedmigrate-helicone.md.
Include the YAML frontmatter at the beginning. Review the instructions before installing.
View and copy the complete migration skill
View and copy the complete migration skill
migrate-helicone.md
2. Install the skill
From the directory containingmigrate-helicone.md, run these commands in a macOS, Linux, or WSL terminal:
SKILL.md is the installed filename for the same contents. You do not need a second script.
If a skill already exists under that name, the command asks before overwriting it. Review the existing file, especially if it targets a different migration provider.
3. Run the migration
Start Claude Code in your application’s project directory and invoke:This skill migrates application code. It does not configure Claude Code itself to use NexLLM as its inference provider.
Step 1: Update Your Environment Variables
Add NexLLM configuration for the clients you intend to migrate. The following values are placeholders, not real credentials.NEXLLM_BASE_URL and NEXLLM_MODEL are application configuration conventions used in this guide. Your code must explicitly read them.
For local development, use your existing environment-loading mechanism. Creating a .env file does not automatically load its values.
For deployed applications, use your hosting platform or secret manager.
Migrated NexLLM requests authenticate with the NexLLM key, not your Helicone key or upstream provider key.
Step 2: Update Your Client
Update the base URL and credentials for migrated call sites. Before running these examples, complete the model discovery in Step 4 and setNEXLLM_MODEL to one approved Chat Completions model.
Preserve appropriate timeouts, retry settings, custom transports, and unrelated headers. Review existing gateway retry behavior before adding new retry layers.
Helicone-Auth from migrated clients. Resolve the behavior behind other Helicone headers using the next section.
If a shared client still serves Helicone requests, create separate client configurations instead of changing it globally.
Step 3: Review and Remove Helicone Headers
Do not forward Helicone credentials to NexLLM. For other Helicone-specific headers, first identify the feature they control. Remove them from migrated requests only after you have a verified replacement or have explicitly approved the behavior change. The tables below describe migration actions. They do not assert that NexLLM has a one-to-one equivalent for every Helicone feature.Request header mapping
Request header mapping
Response header mapping
Response header mapping
Step 4: Update Model Identifiers
Discover the models available to the NexLLM key you will use for inference:data. Use the exact data[].id value as the request’s model.
NexLLM’s documentation includes these example identifiers:
gpt-4oaws/claude-haiku-4-5gemini-2.5-flash
Provider-proxy models
A model ID used through a provider proxy might also appear in NexLLM’s catalog. Keep it unchanged only after confirming its availability and required behavior.AI Gateway model identifiers
Do not assume anauthor/model identifier works unchanged. Do not automatically remove its prefix or replace it with a guessed provider prefix.
Build a mapping for each model actually used by your application:
Changing a model identifier does not by itself establish equivalent hosting, regional processing, retention, or contractual guarantees.
Channel groups and model restrictions
Each NexLLM key belongs to a channel group that controls model access and pricing. Keys may also have explicit model restrictions. Use the same key for discovery and inference. If a model is inaccessible, check both its channel-group availability and the key’s restrictions. Do not silently switch model versions or select a replacement.What about auto?
This guide does not establish a NexLLM equivalent to Helicone’s automatic model-selection behavior.
Choose an explicit available model, implement an approved selection policy, or keep that call site on Helicone until a suitable alternative has been verified.
Step 5: Reconnect Observability
NexLLM documents per-request usage logs and dashboard statistics. Use these for the information they provide, and preserve additional application telemetry where needed.
NexLLM Usage Logs document request time, API key name, model name, timing, input/output token consumption, and cost. Dashboard statistics include requests, quota consumption, token usage, RPM, and TPM.
Do not assume those statistics replace every Helicone dashboard view.
Exporting Your Helicone History
Changing the inference endpoint does not migrate historical logs. Before deprovisioning Helicone:- Identify the history needed for your audit, debugging, or reporting requirements.
- Review Helicone’s current export options for your account or deployment.
- Export the required records to approved storage.
- Check completeness, access controls, and retention requirements.
- Keep access to the original system until the archive is validated.
Step 6 (Optional): Evaluate the Responses API
NexLLM’s API overview listsPOST /v1/responses as an OpenAI Responses API-compatible endpoint.
You do not need to adopt it to migrate an existing Chat Completions integration. Treat changing API families as a separate project.
Before using Responses in production, verify the chosen model’s support for the features you need, including streaming, tools, structured output, multimodal input, state, storage, and response parsing.
Model discovery establishes catalog availability, not support for every endpoint or feature. Do not assume
previous_response_id, web search, or other advanced features work for every available model.NEXLLM_RESPONSES_MODEL to a model whose Responses support you have confirmed. This is a suggested application variable, not a provider-defined setting.
The following are minimal starting requests for that verification.
Step 7: Verify Before Switching Production Traffic
Start with one approved model and a short synthetic prompt. Check that:- The process loads the intended NexLLM credentials.
- The selected model is available to that key.
- The request reaches NexLLM rather than a Helicone endpoint.
- No migrated request forwards a Helicone or provider key.
- The response has the expected structure.
- Application errors and timeouts are handled correctly.
- New requests appear in the expected usage logs.
- Required privacy and security behavior remains intact.
- Required application features pass their own tests.
Minimal Python smoke test
This test uses the OpenAI Python SDK. ConfigureNEXLLM_API_KEY and NEXLLM_MODEL through your existing environment.
Test streaming
NexLLM Chat Completions supports Server-Sent Events whenstream=True.
Using your configured client and an approved streaming-capable model:
Test required behavior separately
Create focused tests for:- Tool calls and tool-result round trips.
- Structured output and schema validation.
- Images or other multimodal inputs.
- Retry limits and timeout handling.
- Fallback selection and failure cases.
- Cache isolation and expiry.
- Context-overflow handling.
- Moderation and logging restrictions.
- Telemetry, alerts, and usage accounting.
Why Migrate to NexLLM?
Keep your existing Chat Completions client
For compatible integrations, continue using the OpenAI SDK or your existing HTTP client with NexLLM configuration.Access multiple model families
NexLLM provides access to GPT, Claude, and Gemini models through a single API base URL. Your actual access depends on your credentials and channel group.Choose model access and pricing through channel groups
Channel groups determine model access and the pricing ratio applied to usage. Review the current group and pricing documentation rather than assuming a universal markup, discount, or price match.Configure scoped API keys
NexLLM documents key expiration, quota settings, model restrictions, and IP allowlists. These controls should be reviewed against your application’s access requirements. They are not automatically equivalent to a hierarchical organization budgeting system or a request-rate policy.Monitor request usage
Use NexLLM Usage Logs and dashboard statistics to inspect model activity, token consumption, timing, and cost. Keep separate instrumentation for application-specific traces and tags.Evaluate additional API formats when needed
NexLLM’s API overview also lists Responses and Claude Messages endpoints. Verify endpoint and model compatibility before changing request formats. Their presence does not establish feature parity across all models.Understand the trade-offs
This migration does not promise automatic replacements for:- Helicone’s custom properties and session tracing.
- Gateway response caching.
- Prompt management.
- Specific automatic routing or fallback policies.
- Existing moderation and security rules.
- Provider-specific retention or regional processing guarantees.
Troubleshooting
Model not found or inaccessible
QueryGET /v1/models with the same key used for inference.
Check the exact model ID, channel group, and any explicit model restrictions. Do not guess a prefix or silently substitute a model.
A model missing from one key’s catalog is not necessarily unavailable platform-wide.
Authentication failure
Confirm that the process loads a NexLLM key and sends Bearer authentication to the NexLLM endpoint. Check for stale secrets, accidental whitespace, expiration, exhausted quota, and IP restrictions where applicable. Never print the key while debugging.Requests succeed but do not appear in NexLLM logs
Confirm that the actual runtime client points to NexLLM rather than a Helicone endpoint or another configured proxy. Check that you are viewing the correct NexLLM account and time range. If Helicone remains configured only for asynchronous logging, inspect that separate integration rather than assuming the inference URL controls all telemetry.Cache hit rate changed
Determine whether the old metric measured a gateway response cache, provider prompt caching, or an application cache. Do not compare those as though they were interchangeable. Rebuild required cache metrics around the implementation actually used after migration.Connection or endpoint errors
Use this SDK base URL:/v1, retain an old target URL, or invent a health endpoint. Use the documented Models endpoint to test authenticated API access without generating a completion.
Retry or rate-limit settings stopped applying
Review the old Helicone request headers and the behavior they controlled. Configure an approved SDK or application policy and test failure cases. Do not assume key quota settings replace requests-per-minute controls.Custom properties or sessions disappeared
Review the application instrumentation that previously populated Helicone headers. Preserve those values in your tracing or analytics system unless you have confirmed a supported NexLLM mapping.Unsupported parameter
Reduce the request to an approvedmodel and a simple messages array to isolate the problem.
Reintroduce required options individually. Permanently removing a required feature needs a separate approval.
Environment variable not found
Confirm the application process receivesNEXLLM_API_KEY and NEXLLM_MODEL.
If using a .env file, confirm the application or development tooling loads it. Deployment secrets must be configured in the deployed environment separately.
A provider-native client stopped working
Check whether the original integration used Chat Completions, Claude Messages, or another native request format. Do not send native payloads to a Chat Completions endpoint. Use the corresponding documented endpoint and verify authentication, model support, and response parsing separately.Next Steps
- Authentication: Review credential handling.
- Models: Discover exact model identifiers.
- Channel Groups: Review model access and pricing categories.
- API Keys: Configure key restrictions.
- Usage & Logs: Check request activity and consumption.
- Chat Completions: Review request fields and streaming.
- API Overview: Review other documented endpoints.
Feedback
If a migration requirement does not translate cleanly, record:- The affected integration mode.
- The feature or header involved.
- The endpoint and model identifier.
- The expected behavior.
- A minimal reproduction using synthetic data.
- Sanitized status codes and relevant diagnostic details.