JETINFER

Integrations

We are an OpenAI-compatible endpoint. In almost every client that means a base URL and a key — no SDK to install, no code to rewrite.

base_urlhttps://api.jetinfer.com/v1
modelqwen3.8-27b
api_keysk-... — create one in your account
Get an API key API docs sign in with Google · prepaid credit, nothing renews

Your first request

Nothing to install:

curl https://api.jetinfer.com/v1/chat/completions \
  -H "Authorization: Bearer $JETINFER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

The same call from the OpenAI Python SDK — two lines differ from OpenAI:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.jetinfer.com/v1",
    api_key=os.environ["JETINFER_API_KEY"],
)

r = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Hello"}],
)

And from TypeScript:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.jetinfer.com/v1",
  apiKey: process.env.JETINFER_API_KEY,
});

const r = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Hello" }],
});

Streaming, tool calling, JSON-schema structured output, logprobs, seed and stop sequences all work the way the OpenAI SDK expects. The legacy /v1/completions text endpoint is served too.

Vercel AI SDK

Use the OpenAI-compatible provider package rather than @ai-sdk/openai, so the SDK does not assume OpenAI-only fields:

npm i @ai-sdk/openai-compatible
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";

const jetinfer = createOpenAICompatible({
  name: "jetinfer",
  baseURL: "https://api.jetinfer.com/v1",
  apiKey: process.env.JETINFER_API_KEY,
});

const { text } = await generateText({
  model: jetinfer("qwen3.8-27b"),
  prompt: "Hello",
});

Coding agents

Every one of these takes a custom OpenAI-compatible endpoint. Long agent sessions are what our prefix cache is built for: the file context you resend every turn bills at the cached rate.

Cline and Roo Code. Settings → API Provider:

FieldValue
API ProviderOpenAI Compatible
Base URLhttps://api.jetinfer.com/v1
API Keysk-...
Model IDqwen3.8-27b

Continue. In ~/.continue/config.yaml:

models:
  - name: JetInfer Qwen3.8-27B
    provider: openai
    model: qwen3.8-27b
    apiBase: https://api.jetinfer.com/v1
    apiKey: sk-...

Aider. Two environment variables and a model prefix:

export OPENAI_API_BASE=https://api.jetinfer.com/v1
export OPENAI_API_KEY=sk-...
aider --model openai/qwen3.8-27b

If your agent speaks Anthropic's shape instead, we serve POST /v1/messages at the same host on the same key, so ANTHROPIC_BASE_URL=https://api.jetinfer.com reaches us with no translating proxy in between. The exact requests the Anthropic SDKs send - x-api-key auth, non-streaming and the streamed event sequence - are verified against the live endpoint. We have not run Claude Code itself against it end to end, so we are not going to tell you it works. Tell us what happens.

Frameworks, routers and proxies

All of these accept an OpenAI-compatible endpoint somewhere in their settings. Point them at https://api.jetinfer.com/v1 with model qwen3.8-27b:

ToolWhere it goes
LangChainChatOpenAI(base_url=..., api_key=..., model=...)
LlamaIndexOpenAILike(api_base=..., api_key=..., model=...)
LiteLLMmodel="openai/qwen3.8-27b" with api_base=...
Cloudflare AI Gatewayadd us as a Custom Provider; header Authorization: Bearer sk-...
Portkey, Helicone, LangDBcustom OpenAI-compatible provider, same two fields
Open WebUI, LM Studioadd an OpenAI-compatible connection

Put a caching gateway in front of us if you like: a hit inside it never reaches us and costs nothing here. Or reach us through a router you already use, once our listing there is live, and keep your existing billing. Same endpoint either way.

What to expect when things go wrong

Error codes mean specific things here, so handle them distinctly rather than retrying everything:

CodeMeaning
401Key is missing, wrong, or revoked.
402Out of credit. Retrying will not help — top up. Deliberately not a 429.
429We are at capacity. Carries Retry-After; retry is the right response. We shed load rather than queue you.
400Your request — context overflow, bad schema. The message names the real limit.
5xxOurs. You are never charged for one.

Is it us? https://api.jetinfer.com/status/public answers with no key and no account: current state, uptime over 24 hours, 7 days and 30 days, and time-to-first-token percentiles. It is served from the edge, so it still answers when the API does not.

Check your balance any time with your own key: GET https://api.jetinfer.com/v1/credits. Rates are published in /v1/models and on the specs page. We don't repeat them here, because they move with the market.