Three values and one request. We speak the OpenAI API, so everything below is the standard shape. This page says what we accept and what our refusals mean.
base_url | https://api.jetinfer.com/v1 |
model | qwen3.8-27b |
api_key | sk-... — create one in your account |
Send the key as a bearer token. Nothing to install.
curl https://api.jetinfer.com/v1/chat/completions \
-H "Authorization: Bearer $JETINFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Hello"}]
}'
You get an ordinary chat completion with the usage object on it. Every response
carries an Inference-Id header, errors included. Quote it when you
write to us and we can find the request.
The catalogue needs no key.
GET https://api.jetinfer.com/v1/models returns the model id, context
length, max output, quantization, supported parameters and current rates. It's
generated from the table that bills you, so rates aren't repeated here.
There's no JetInfer SDK and there never will be. Point the OpenAI client at us:
a base_url and an api_key, nothing else.
Python.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.jetinfer.com/v1",
api_key=os.environ["JETINFER_API_KEY"],
)
r = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)
Node.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.jetinfer.com/v1",
apiKey: process.env.JETINFER_API_KEY,
});
const r = await client.chat.completions.create({
model: "qwen3.8-27b",
messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);
If your client speaks Anthropic's shape instead, send it to
POST https://api.jetinfer.com/v1/messages with the same key - either
as x-api-key, the header the Anthropic SDKs send, or as
Authorization: Bearer. Same limits, same billing as everything else
here. Errors come back in the Anthropic envelope, streaming emits the Anthropic
event sequence, and the rate-limit headers are re-emitted under their
anthropic-ratelimit-* names.
Anthropic SDK. The base URL has no /v1 - the SDK
appends /v1/messages itself. Thinking arrives as a
thinking block ahead of the text block.
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.jetinfer.com",
api_key=os.environ["JETINFER_API_KEY"],
)
m = client.messages.create(
model="qwen3.8-27b",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
print(next(b.text for b in m.content if b.type == "text"))
| Route | Shape |
|---|---|
POST /v1/chat/completions | OpenAI chat |
POST /v1/completions | OpenAI legacy text |
POST /v1/messages | Anthropic messages |
GET /v1/models | catalogue, no key needed |
GET /v1/credits | your own balance, with your key |
GET /v1/usage | your month, day by day — the reconciliation surface |
Send "stream": true and read server-sent events until
data: [DONE]. Standard SSE, so your SDK already handles it.
curl https://api.jetinfer.com/v1/chat/completions \
-H "Authorization: Bearer $JETINFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'
Usage accounting still arrives. We turn
stream_options.include_usage on for every stream, so the final chunk
carries the token counts whether you asked for them or not. It is the same object
your bill is built from.
A 4xx on a streaming request is a real 4xx: the status is settled before any bytes
go out, so a rejected request never arrives as a 200 carrying an
error frame. If a stream dies after it started, you get an error frame in the
OpenAI shape followed by data: [DONE], so a client waiting on the
terminator is not left hanging.
Exactly this list, which is what /v1/models declares under
supported_parameters. Anything not named here we don't promise,
whatever it may happen to accept.
| Group | Parameters |
|---|---|
| Sampling | temperature, top_p, top_k, frequency_penalty, presence_penalty |
| Length and stopping | max_tokens, stop |
| Determinism | seed |
| Streaming | stream, stream_options |
| Tools | tools, tool_choice |
| Structured output | response_format |
| Token probabilities | logprobs, top_logprobs |
max_completion_tokens, the field current OpenAI SDKs emit, is
accepted too and clamped the same way max_tokens is. Omitting both is
fine: generation is then bounded by what is left of the window. Reasoning output
comes back in the reasoning field and is never merged into
content. The model reasons before it answers and both halves count
against max_tokens, so a tight budget can end in reasoning alone:
finish_reason: "length" with an empty content. Give a
short answer a few hundred tokens rather than a few. Full limits are on the
specs page.
Errors arrive in the OpenAI envelope —
{"error": {"message", "type", "code", "param"}} — with the
status on the response. Handle them distinctly rather than retrying everything.
| Code | What it means | What to do |
|---|---|---|
400 | The request. Context overflow, a bad schema, a
non-finite number, a batched prompt, n above one. |
Read the message; it names the real limit. |
401 | The key is missing, unreadable, unknown or revoked. | Check the header. Retrying will not help. |
402 | Out of credit, or a monthly ceiling reached. | Top up. Deliberately not a 429. |
404 | The model id does not exist here. | Pick one from GET /v1/models. |
413 | The body is over 8 MiB. Refused on the declared length, or cut off mid-stream if none was declared. | Send less in one request. |
429 | Your request rate, or we are at capacity.
Carries Retry-After. | Back off and retry. This one clears on its own. |
5xx | Ours. | Retry. You are not charged for one. |
402 and 429 are different instructions. A 429 says we are busy, so back off and retry. A 402 says the balance is empty, which retrying will never change. The message on a 402 says which it is: an empty prepaid balance, or a ceiling we set and can raise.
A request we fail to serve is never charged. Billing runs off the usage the response reported, and anything recorded as an error is skipped. We shed load with an immediate 429 rather than queueing you, so you're never billed for waiting either.
Two other ceilings: 600 requests per minute per account, raised on request, and
one completion per request (n and best_of above one are
refused, not quietly served). Capacity, context and concurrency are on the
specs page. Whether it's us:
https://api.jetinfer.com/status/public answers with no key.
Every response carries the request budget, so you can slow down before you're refused rather than after.
| Header | Value |
|---|---|
x-ratelimit-limit-requests | requests per minute allowed |
x-ratelimit-remaining-requests | how many are left now |
x-ratelimit-reset-requests | seconds until the budget is full again, as 12.3s |
x-ratelimit-limit-tokens | tokens per minute, when a token quota applies to your account |
x-ratelimit-remaining-tokens | tokens left, never below zero |
x-ratelimit-reset-tokens | seconds until the token budget refills |
The three token headers appear only when a token quota is set on your account; the request headers are always there. Both budgets refill continuously rather than resetting on a boundary, so there's no edge to queue against — the allowance you're shown is the one you have at that instant.
Limits belong to the account, not to the key. A second key is a separate credential, not a second budget.
Prepaid credit. You buy what you want, you spend it, nothing renews and there's no invoice to chase.
Routing platforms settle by monthly invoice instead. If we
issued your key directly, there's no card and no prepay: usage is metered per
request and invoiced monthly at the rates in /v1/models, on the
terms in your agreement. Your side of the ledger is always visible —
api.jetinfer.com/partner opens
with your key alone and shows month-to-date spend, requests, tokens, success
rate and TTFT percentiles; GET /v1/usage?month=YYYY-MM returns the
same month day-by-day in exact integer nano-USD, summing to the invoice by
construction, so reconciliation is a diff rather than a conversation. Only
served requests bill; shed load and refusals never do.
Credit belongs to the account, not to the key. Create a key, move
your traffic, revoke the old one, and the balance is untouched. A key is shown
once, at creation — we store only a hash, so a lost key is replaced rather
than recovered. Check the balance any time with your own key:
GET https://api.jetinfer.com/v1/credits.
Repeated prompt prefix bills at one tenth of the input rate. When
a request begins with a prefix we already hold, those tokens bill at the cached
rate and the rest at the ordinary input rate. Writing a prefix into the cache
costs the input rate with no surcharge, output is never cached, and the cached
count is in the usage object on your response. Hits are counted in 800-token
blocks (the engine's cache granularity for this model): a shared prefix shorter
than 800 tokens bills at the full input rate, and a 1,256-token prompt repeated
bills 800 tokens as cached. Rates are in
/v1/models; the ratio is on the
specs page.
The account page shows where the money went. Spend and token totals over a 7, 30 or 90 day window, a per-key breakdown when you run more than one credential, your most recent requests, and a CSV export of the window for reconciliation. Set a monthly spend ceiling there too — requests are refused with a 402 above it and your credit is untouched — and arm automatic top-up so an agent never dies mid-run. When outbound email is configured on the deployment, you're warned by email when your balance runs low and if automatic top-up pauses; receipts and invoices always come from Stripe.
Leaving is self-serve too. Delete your account from the same page — account, email and keys go immediately, and any remaining credit is forfeited with them, so spend it first. Billing records stay only as long as accounting law requires, and prompts were never stored at all.
Anything on this page, or a request that misbehaved:
hello@jetinfer.com.
Send the Inference-Id and we can look it up.