Skip to content

One API for every leading model.

You pay the model’s rate. Nothing on top.

Keep the SDK you already use and switch models by changing one string. Seats, projects, and API keys are free.

Your request

// same SDK — only the base URL changes
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.inferray.com/v1",
  apiKey: "inferray_...",
});

const resp = await client.chat.completions.create({
  model: "gpt-6-astra",
  messages: [{ role: "user", content: "ship it" }],
});
GPT-6 Astra · per 1M tokens, input / output$10 / $50
Model the snippet calls

The snippet now calls GPT-6 Astra, at $10 / $50 per million tokens, input / output.

Zero data retention

Your prompts are never stored

Inferray follows zero data retention. Prompts, responses, and request payloads pass through and are gone — nothing is written to disk, and nothing is used for training.

  • No prompts stored
  • No responses stored
  • Never used for training

All we keep is what your bill needs: the model, token counts, latency, and cost.

Keep the tools you use

If your code already talks to OpenAI or Anthropic, this is the entire migration. Your prompts, parameters, tools, and retries are untouched — and your agents take the same two values.

client.ts2 lines changed
baseURL: "https://api.openai.com/v1",
baseURL: "https://api.inferray.com/v1",
apiKey: process.env.OPENAI_API_KEY,
apiKey: process.env.INFERRAY_API_KEY,

Two ways to pay

Pay per token with nothing added on top, or take a flat monthly plan that includes three times what it costs.

Pay as you go

0%gateway fees

  • Pay the model rate, with nothing added on top.
  • No seats, no monthly minimum, no commitment.
  • Top up any amount and spend it on any model.

Token Plans: 3× the usage you pay for

Token Plans · Solo

3×

You pay

$20

You get

$60

$40 of extra usage every month, at no extra cost.

Pro · pay $100 a month
get $300 of usage
Max · pay $200 a month
get $600 of usage

What is on your bill

The inference you use, at the rate the model is listed at. Your team, your projects, and your keys cost nothing, however many you add.

  • Streaming, with token usage tracked on every request
  • Spend, tokens, and latency by project, key, and model
  • Your prompts and responses are never stored or used for training
What an Inferray bill charges for
Line itemYou pay
InferencePer token, at the model's listed rateModel rate
Fees on topNothing added to any request$0
SeatsInvite your whole team$0
ProjectsOne per app or environment$0
API keysOne per integration$0
TotalOnly what you run

Every model, one API

Text models are priced per 1M tokens, input / output; image models show their billing unit. Nothing is added to these rates.

Every model Inferray offers, grouped by provider, with its price per million tokens (input / output).
ModelPrice
Anthropic
Claude Fable 5.1claude-fable-5-1$10 / $50
Claude Fable 5claude-fable-5$10 / $50
Claude Opus 5claude-opus-5$5 / $25
Claude Opus 4.8claude-opus-4-8$5 / $25
Claude Sonnet 5claude-sonnet-5$2 / $10
Claude Haiku 4.5claude-haiku-4-5$1 / $5
OpenAI
GPT-6 Astragpt-6-astra$10 / $50
GPT-6 Solgpt-6-sol$2 / $10
GPT-6 Lunagpt-6-luna$0.10 / $0.50
GPT-5.6 Solgpt-5.6-sol$4 / $20
GPT-5.6 Terragpt-5.6-terra$2 / $12
GPT-5.6 Lunagpt-5.6-luna$0.20 / $1.20
GPT-5.5gpt-5.5$5 / $30
GPT-5.4gpt-5.4$2.50 / $15
GPT-5.2gpt-5.2$1.75 / $14
GPT Image 2gpt-image-2$30 / 1M image tokens
Google
Gemini 3 Pro Image (Preview)gemini-3-pro-image-preview$0.13–$0.24 / image
Gemini 2.5 Flash Imagegemini-2.5-flash-image$0.30 / $30
Gemini 3.8 Flashgemini-3.8-flash$0.75 / $3.75
Gemini 2.5 Progemini-2.5-pro$1.25 / $10
DeepSeek
DeepSeek V4.1 Flashdeepseek-v4.1-flash$0.15 / $0.60
Z.ai
GLM-5.3glm-5.3$1.40 / $4.40
Moonshot AI
Kimi K3kimi-k3$3 / $15

Frequently asked questions

What to know about pricing, integration, and your data.

Do you charge gateway fees?
No. You pay the model rates shown in our catalogue, with no additional gateway fee. Seats, projects, and API keys are free.
What do I have to change in my code?
Your base URL and your key. Your SDK, your prompts, your tools, and your agents all stay exactly as they are.
How does billing work?
With pay as you go, top up your balance and pay for the inference you use, with no monthly minimum. Track spend by project, API key, and model in the console.
What happens to my prompts?
Inferray does not store your prompts or responses, or use them for training. Your console shows usage and cost without retaining that content.

Start building at the model’s rate

Create your workspace, get an API key, and top up for your first request. Your team, projects, and keys are included.