AI inference infrastructure
Qwen3.8-27B for agent and coding workloads, on your OpenAI or Anthropic SDK. Repeated prompt prefix bills at one tenth of the input rate.
$0.21 in · $2.00 out · $0.021 cached per million tokens · live in /v1/models
Service paused. We are not taking new sign-ups or credit while inference is offline — we will not charge for requests we cannot serve. Existing credit is untouched. Email us to be told when we are back.
A cached prefix bills at one tenth of the input rate: a 6,500-token prompt with 90% cached bills as 1,235 tokens. Output is never cached, and the rates are in /v1/models. We haven't measured our own production hit rate yet. How this was measured.
OpenAI-compatible /v1/chat/completions and Anthropic-compatible
/v1/messages on the same key. Change the base URL, nothing else.
Streaming, tool calling and JSON schema output all work.
Capacity grows as demand does. When we're full you get an immediate 429 instead of a queue, so your client backs off rather than waiting.
Prompts and answers are never written to disk, logs or analytics, and never used for training. We serve from the EU.
Every throughput and latency figure is published with the prompt length, concurrency and endpoint that produced it. See them.
29 concurrent requests, 262,144 tokens of context, 65,536 max output, served from the EU. Capacity is added as demand arrives.
Exceed a limit and the request fails with an explicit error. Prompts are never silently truncated to fit. Full limits.
Directly. Creating an account is paused while inference is offline; it took one click, add credit, and create a key. There is no SDK to install: pointing an existing OpenAI client at us is a base URL and a key. The docs are one page.
base_url | https://api.jetinfer.com/v1 |
model | qwen3.8-27b |
Billing is prepaid credit: you buy what you want, you spend it, nothing renews and there's no invoice to chase. Credit belongs to your account rather than to a key, so you can rotate a key without losing your balance. A request we fail to serve is never charged.
Or through a router. If you already route through an aggregator, you'll be able to reach us that way once our listing is live and keep your existing billing. Same endpoint either way.
Questions about the measurements or the retention terms: hello@jetinfer.com.