JETINFER

API docs

Three values and one request. We speak the OpenAI API, so everything below is the standard shape. This page says what we accept and what our refusals mean.

base_urlhttps://api.jetinfer.com/v1
modelqwen3.8-27b
api_keysk-... — create one in your account
Get your API key Client and framework setup sign in with Google · prepaid credit · nothing renews

Quickstart

Send the key as a bearer token. Nothing to install.

curl https://api.jetinfer.com/v1/chat/completions \
  -H "Authorization: Bearer $JETINFER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

You get an ordinary chat completion with the usage object on it. Every response carries an Inference-Id header, errors included. Quote it when you write to us and we can find the request.

The catalogue needs no key. GET https://api.jetinfer.com/v1/models returns the model id, context length, max output, quantization, supported parameters and current rates. It's generated from the table that bills you, so rates aren't repeated here.

Using an SDK you already have

There's no JetInfer SDK and there never will be. Point the OpenAI client at us: a base_url and an api_key, nothing else.

Python.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.jetinfer.com/v1",
    api_key=os.environ["JETINFER_API_KEY"],
)

r = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)

Node.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.jetinfer.com/v1",
  apiKey: process.env.JETINFER_API_KEY,
});

const r = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);

If your client speaks Anthropic's shape instead, send it to POST https://api.jetinfer.com/v1/messages with the same key - either as x-api-key, the header the Anthropic SDKs send, or as Authorization: Bearer. Same limits, same billing as everything else here. Errors come back in the Anthropic envelope, streaming emits the Anthropic event sequence, and the rate-limit headers are re-emitted under their anthropic-ratelimit-* names.

Anthropic SDK. The base URL has no /v1 - the SDK appends /v1/messages itself. Thinking arrives as a thinking block ahead of the text block.

import os
from anthropic import Anthropic
client = Anthropic(
    base_url="https://api.jetinfer.com",
    api_key=os.environ["JETINFER_API_KEY"],
)
m = client.messages.create(
    model="qwen3.8-27b",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
)
print(next(b.text for b in m.content if b.type == "text"))
RouteShape
POST /v1/chat/completionsOpenAI chat
POST /v1/completionsOpenAI legacy text
POST /v1/messagesAnthropic messages
GET /v1/modelscatalogue, no key needed
GET /v1/creditsyour own balance, with your key
GET /v1/usageyour month, day by day — the reconciliation surface

Streaming

Send "stream": true and read server-sent events until data: [DONE]. Standard SSE, so your SDK already handles it.

curl https://api.jetinfer.com/v1/chat/completions \
  -H "Authorization: Bearer $JETINFER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Usage accounting still arrives. We turn stream_options.include_usage on for every stream, so the final chunk carries the token counts whether you asked for them or not. It is the same object your bill is built from.

A 4xx on a streaming request is a real 4xx: the status is settled before any bytes go out, so a rejected request never arrives as a 200 carrying an error frame. If a stream dies after it started, you get an error frame in the OpenAI shape followed by data: [DONE], so a client waiting on the terminator is not left hanging.

What is supported

Exactly this list, which is what /v1/models declares under supported_parameters. Anything not named here we don't promise, whatever it may happen to accept.

GroupParameters
Samplingtemperature, top_p, top_k, frequency_penalty, presence_penalty
Length and stoppingmax_tokens, stop
Determinismseed
Streamingstream, stream_options
Toolstools, tool_choice
Structured outputresponse_format
Token probabilitieslogprobs, top_logprobs

max_completion_tokens, the field current OpenAI SDKs emit, is accepted too and clamped the same way max_tokens is. Omitting both is fine: generation is then bounded by what is left of the window. Reasoning output comes back in the reasoning field and is never merged into content. The model reasons before it answers and both halves count against max_tokens, so a tight budget can end in reasoning alone: finish_reason: "length" with an empty content. Give a short answer a few hundred tokens rather than a few. Full limits are on the specs page.

Errors and limits

Errors arrive in the OpenAI envelope — {"error": {"message", "type", "code", "param"}} — with the status on the response. Handle them distinctly rather than retrying everything.

CodeWhat it meansWhat to do
400The request. Context overflow, a bad schema, a non-finite number, a batched prompt, n above one. Read the message; it names the real limit.
401The key is missing, unreadable, unknown or revoked.Check the header. Retrying will not help.
402Out of credit, or a monthly ceiling reached.Top up. Deliberately not a 429.
404The model id does not exist here. Pick one from GET /v1/models.
413The body is over 8 MiB. Refused on the declared length, or cut off mid-stream if none was declared. Send less in one request.
429Your request rate, or we are at capacity. Carries Retry-After.Back off and retry. This one clears on its own.
5xxOurs.Retry. You are not charged for one.

402 and 429 are different instructions. A 429 says we are busy, so back off and retry. A 402 says the balance is empty, which retrying will never change. The message on a 402 says which it is: an empty prepaid balance, or a ceiling we set and can raise.

A request we fail to serve is never charged. Billing runs off the usage the response reported, and anything recorded as an error is skipped. We shed load with an immediate 429 rather than queueing you, so you're never billed for waiting either.

Two other ceilings: 600 requests per minute per account, raised on request, and one completion per request (n and best_of above one are refused, not quietly served). Capacity, context and concurrency are on the specs page. Whether it's us: https://api.jetinfer.com/status/public answers with no key.

Pacing yourself

Every response carries the request budget, so you can slow down before you're refused rather than after.

HeaderValue
x-ratelimit-limit-requestsrequests per minute allowed
x-ratelimit-remaining-requestshow many are left now
x-ratelimit-reset-requestsseconds until the budget is full again, as 12.3s
x-ratelimit-limit-tokenstokens per minute, when a token quota applies to your account
x-ratelimit-remaining-tokenstokens left, never below zero
x-ratelimit-reset-tokensseconds until the token budget refills

The three token headers appear only when a token quota is set on your account; the request headers are always there. Both budgets refill continuously rather than resetting on a boundary, so there's no edge to queue against — the allowance you're shown is the one you have at that instant.

Limits belong to the account, not to the key. A second key is a separate credential, not a second budget.

Billing

Prepaid credit. You buy what you want, you spend it, nothing renews and there's no invoice to chase.

Routing platforms settle by monthly invoice instead. If we issued your key directly, there's no card and no prepay: usage is metered per request and invoiced monthly at the rates in /v1/models, on the terms in your agreement. Your side of the ledger is always visible — api.jetinfer.com/partner opens with your key alone and shows month-to-date spend, requests, tokens, success rate and TTFT percentiles; GET /v1/usage?month=YYYY-MM returns the same month day-by-day in exact integer nano-USD, summing to the invoice by construction, so reconciliation is a diff rather than a conversation. Only served requests bill; shed load and refusals never do.

Credit belongs to the account, not to the key. Create a key, move your traffic, revoke the old one, and the balance is untouched. A key is shown once, at creation — we store only a hash, so a lost key is replaced rather than recovered. Check the balance any time with your own key: GET https://api.jetinfer.com/v1/credits.

Repeated prompt prefix bills at one tenth of the input rate. When a request begins with a prefix we already hold, those tokens bill at the cached rate and the rest at the ordinary input rate. Writing a prefix into the cache costs the input rate with no surcharge, output is never cached, and the cached count is in the usage object on your response. Hits are counted in 800-token blocks (the engine's cache granularity for this model): a shared prefix shorter than 800 tokens bills at the full input rate, and a 1,256-token prompt repeated bills 800 tokens as cached. Rates are in /v1/models; the ratio is on the specs page.

The account page shows where the money went. Spend and token totals over a 7, 30 or 90 day window, a per-key breakdown when you run more than one credential, your most recent requests, and a CSV export of the window for reconciliation. Set a monthly spend ceiling there too — requests are refused with a 402 above it and your credit is untouched — and arm automatic top-up so an agent never dies mid-run. When outbound email is configured on the deployment, you're warned by email when your balance runs low and if automatic top-up pauses; receipts and invoices always come from Stripe.

Leaving is self-serve too. Delete your account from the same page — account, email and keys go immediately, and any remaining credit is forfeited with them, so spend it first. Billing records stay only as long as accounting law requires, and prompts were never stored at all.

Add credit and create a key See the measurements sign up or log in with Google · one click

Questions

Anything on this page, or a request that misbehaved: hello@jetinfer.com. Send the Inference-Id and we can look it up.