We are an OpenAI-compatible endpoint. In almost every client that means a base URL and a key — no SDK to install, no code to rewrite.
base_url | https://api.jetinfer.com/v1 |
model | qwen3.8-27b |
api_key | sk-... — create one in your account |
Nothing to install:
curl https://api.jetinfer.com/v1/chat/completions \
-H "Authorization: Bearer $JETINFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Hello"}]
}'
The same call from the OpenAI Python SDK — two lines differ from OpenAI:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.jetinfer.com/v1",
api_key=os.environ["JETINFER_API_KEY"],
)
r = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Hello"}],
)
And from TypeScript:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.jetinfer.com/v1",
apiKey: process.env.JETINFER_API_KEY,
});
const r = await client.chat.completions.create({
model: "qwen3.8-27b",
messages: [{ role: "user", content: "Hello" }],
});
Streaming, tool calling, JSON-schema structured output, logprobs, seed and stop
sequences all work the way the OpenAI SDK expects. The legacy
/v1/completions text endpoint is served too.
Use the OpenAI-compatible provider package rather than @ai-sdk/openai,
so the SDK does not assume OpenAI-only fields:
npm i @ai-sdk/openai-compatible
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";
const jetinfer = createOpenAICompatible({
name: "jetinfer",
baseURL: "https://api.jetinfer.com/v1",
apiKey: process.env.JETINFER_API_KEY,
});
const { text } = await generateText({
model: jetinfer("qwen3.8-27b"),
prompt: "Hello",
});
Every one of these takes a custom OpenAI-compatible endpoint. Long agent sessions are what our prefix cache is built for: the file context you resend every turn bills at the cached rate.
Cline and Roo Code. Settings → API Provider:
| Field | Value |
|---|---|
| API Provider | OpenAI Compatible |
| Base URL | https://api.jetinfer.com/v1 |
| API Key | sk-... |
| Model ID | qwen3.8-27b |
Continue. In ~/.continue/config.yaml:
models:
- name: JetInfer Qwen3.8-27B
provider: openai
model: qwen3.8-27b
apiBase: https://api.jetinfer.com/v1
apiKey: sk-...
Aider. Two environment variables and a model prefix:
export OPENAI_API_BASE=https://api.jetinfer.com/v1
export OPENAI_API_KEY=sk-...
aider --model openai/qwen3.8-27b
If your agent speaks Anthropic's shape instead, we serve
POST /v1/messages at the same host on the same key, so
ANTHROPIC_BASE_URL=https://api.jetinfer.com reaches us with no
translating proxy in between. The exact requests the Anthropic SDKs send -
x-api-key auth, non-streaming and the streamed event sequence - are
verified against the live endpoint. We have not run Claude Code itself against it
end to end, so we are not going to tell you it works. Tell us what happens.
All of these accept an OpenAI-compatible endpoint somewhere in their settings.
Point them at https://api.jetinfer.com/v1 with model
qwen3.8-27b:
| Tool | Where it goes |
|---|---|
| LangChain | ChatOpenAI(base_url=..., api_key=..., model=...) |
| LlamaIndex | OpenAILike(api_base=..., api_key=..., model=...) |
| LiteLLM | model="openai/qwen3.8-27b" with api_base=... |
| Cloudflare AI Gateway | add us as a Custom Provider; header Authorization: Bearer sk-... |
| Portkey, Helicone, LangDB | custom OpenAI-compatible provider, same two fields |
| Open WebUI, LM Studio | add an OpenAI-compatible connection |
Put a caching gateway in front of us if you like: a hit inside it never reaches us and costs nothing here. Or reach us through a router you already use, once our listing there is live, and keep your existing billing. Same endpoint either way.
Error codes mean specific things here, so handle them distinctly rather than retrying everything:
| Code | Meaning |
|---|---|
401 | Key is missing, wrong, or revoked. |
402 | Out of credit. Retrying will not help — top up. Deliberately not a 429. |
429 | We are at capacity. Carries
Retry-After; retry is the right response. We shed load rather
than queue you. |
400 | Your request — context overflow, bad schema. The message names the real limit. |
5xx | Ours. You are never charged for one. |
Is it us? https://api.jetinfer.com/status/public
answers with no key and no account: current state, uptime over 24 hours, 7 days
and 30 days, and time-to-first-token percentiles. It is served from the edge, so
it still answers when the API does not.
Check your balance any time with your own key:
GET https://api.jetinfer.com/v1/credits. Rates are published in
/v1/models and on the specs page. We
don't repeat them here, because they move with the market.