Skip to content
−30%25 models, OpenAI list price minus 30%

Every OpenAI model.
30% less.

A drop-in, OpenAI-compatible API. Keep your SDK, your prompts and your code. Change the base URL and key, and every token costs 30% less than list.

  • Same SDKs
  • Streaming passthrough
  • No subscription
POST api.blizon.tech/v1/chat/completions
streaming
{
"model": "gpt-4o-mini",
"stream": true,
"messages": [{ role: "user", content: "Summarize ticket #4821 in one sentence." }]
}
assistant · delta.content

The customer was charged twice for the March invoice after updating their card, and wants the duplicate $49 payment refunded to the original card before the next billing date.

prompt_tokens 1,284completion_tokens 38
This request$0.000151$0.000215
At 1M requests / mo$151$215

Illustrative request on gpt-4o-mini, priced from the live rate card. OpenAI list shown struck through; Blizon is 30% less.

Works with anything that speaks the OpenAI API

  • openai-python
  • openai-node
  • LangChain
  • LlamaIndex
  • Vercel AI SDK
  • LiteLLM
  • curl

Quickstart

Switch in two lines.

Blizon speaks the Chat Completions API. Point your client at a new base URL, swap the key, and everything else stays exactly where it is.

  • Same request and response shapes. Your parsing, retries and types keep working.
  • Streaming passes straight through. Server-sent events, forwarded chunk by chunk.
  • Cached input billed at the cached rate. Then 30% off that too.
base_urlhttps://api.blizon.tech/v1
import os
from openai import OpenAI
client = OpenAI(
removed: api_key=os.environ["OPENAI_API_KEY"],
added: base_url="https://api.blizon.tech/v1",
added: api_key=os.environ["BLIZON_API_KEY"],
)
stream = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)

Platform

A thin gateway with the accounting built in.

Requests go straight to the model. Everything around them, from keys to costs to failover, is handled for you.

Every request, itemized

A per-request ledger of model, tokens and exact cost, with spend by key and by model. Know what each feature costs.

Under 1 ms of overhead

Streaming responses are passed through as server-sent events, chunk by chunk. No buffering.

Automatic failover

If an upstream errors out, Blizon automatically retries the request on a healthy one.

Keys you control

One key per app or environment. Track spend per key and revoke any of them instantly.

Prompts stay yours

We log token counts and cost for billing. Prompt and completion contents are never stored.

Credits, not subscriptions

1 credit = $1. Credits are assigned to your account after approval and each request deducts its exact cost. No plans, no seats, no monthly fee.

  • 1 credit = $1
  • No subscription
  • Billed per token

Savings

See what you'd keep.

Same models, same tokens, 30% off the bill. Drag to your current monthly OpenAI spend.

Type any amount, or use the slider from $100 to $100,000.

Input / 1M tokens
$0.28$0.40
Output / 1M tokens
$1.12$1.60
At $5,000 a month you would pay $3,500 on Blizon and save $1,500 a month, $18,000 a year.

You save

$1,500/mo
$18,000 per year
OpenAI list$5,000
Blizon$3,500
1.43×
tokens for the same budget
−30%
on input, cached and output

Pricing

OpenAI list price, minus 30%.

One flat discount on input, cached input and output, for every model. No tiers, no commitments.

Models
25
Discount
−30%
GPT-5.x 7
  • gpt-5−30%
    Input
    $0.875OpenAI list $1.25
    Cached
    $0.0875OpenAI list $0.125
    Output
    $7.00OpenAI list $10.00
  • gpt-5-mini−30%
    Input
    $0.175OpenAI list $0.25
    Cached
    $0.0175OpenAI list $0.025
    Output
    $1.40OpenAI list $2.00
  • gpt-5-nano−30%
    Input
    $0.035OpenAI list $0.05
    Cached
    $0.0035OpenAI list $0.005
    Output
    $0.28OpenAI list $0.40
  • gpt-5.1−30%
    Input
    $0.875OpenAI list $1.25
    Cached
    $0.0875OpenAI list $0.125
    Output
    $7.00OpenAI list $10.00
  • gpt-5.3-codex−30%
    Input
    $1.23OpenAI list $1.75
    Cached
    $0.1225OpenAI list $0.175
    Output
    $9.80OpenAI list $14.00
  • gpt-5.4−30%
    Input
    $1.75OpenAI list $2.50
    Cached
    $0.175OpenAI list $0.25
    Output
    $10.50OpenAI list $15.00
  • gpt-5.5-instant−30%
    Input
    $3.50OpenAI list $5.00
    Cached
    $0.35OpenAI list $0.50
    Output
    $21.00OpenAI list $30.00
GPT-4.1 3
  • gpt-4.1−30%
    Input
    $1.40OpenAI list $2.00
    Cached
    $0.35OpenAI list $0.50
    Output
    $5.60OpenAI list $8.00
  • gpt-4.1-mini−30%
    Input
    $0.28OpenAI list $0.40
    Cached
    $0.07OpenAI list $0.10
    Output
    $1.12OpenAI list $1.60
  • gpt-4.1-nano−30%
    Input
    $0.07OpenAI list $0.10
    Cached
    $0.0175OpenAI list $0.025
    Output
    $0.28OpenAI list $0.40
GPT-4o 2
  • gpt-4o−30%
    Input
    $1.75OpenAI list $2.50
    Cached
    $0.875OpenAI list $1.25
    Output
    $7.00OpenAI list $10.00
  • gpt-4o-mini−30%
    Input
    $0.105OpenAI list $0.15
    Cached
    $0.0525OpenAI list $0.075
    Output
    $0.42OpenAI list $0.60
GPT-3.5 1
  • gpt-3.5-turbo−30%
    Input
    $0.35OpenAI list $0.50
    Cached
    —not available
    Output
    $1.05OpenAI list $1.50
USD per 1M tokens. Blizon = OpenAI list × 0.70. Cached input is billed at the cached rate.13 of 25 models

How it works

Live in three steps.

  1. 01

    Request access

    Tell us what you're building. Every request is reviewed by hand, and we reply by email.

    5 short fields
  2. 02

    Get approved and funded

    You get an invite to set a password, create API keys, and we load credits to your account.

    1 credit = $1
  3. 03

    Swap the base URL

    Point your OpenAI client at Blizon with your new key. That's the whole migration.

    https://api.blizon.tech/v1

Request access

Start paying less for the same tokens.

Access is request-only. Every request is read by a person, and approved accounts get an invite and credits to start sending requests.

  • A quick, human review
    We reply by email, usually within a day.
  • Keys and credits on approval
    Create as many API keys as you need. 1 credit = $1.
  • Your prompts aren't stored
    We log token counts and cost, not contents.
  • No subscription
    Credits are deducted per request at the exact cost.

Tell us about your project

Everything except company is required.

Expected monthly spend

Models, rough volume and whether you stream. It helps us size your account.

Reviewed by a person. We reply by email, usually within a day.

FAQ

Questions, answered.

Something else on your mind? Mention it in your access request and we'll answer when we reply.

Is it really compatible with the OpenAI SDK?
Yes. Blizon implements the OpenAI Chat Completions API. Use the official Python or Node SDK, or any OpenAI-compatible library, and set base_url to https://api.blizon.tech/v1 with your Blizon key.
How is the price calculated?
Every price is OpenAI's list price × 0.70, for input, cached input and output tokens. Cached input is billed at the cached rate, so prompt caching saves you money on top of the discount.
Does streaming work?
Yes. Pass stream: true and you get standard server-sent events, passed through as they arrive. The gateway adds under 1 ms of overhead, and token usage is still recorded for billing.
Which models are available?
25 models across the GPT-5.x, GPT-4.1, GPT-4o and GPT-3.5 families, including mini and nano variants and dated snapshots. The pricing table lists every one.
How do credits work?
1 credit = $1. Once your access is approved we assign credits to your account, and each request deducts its exact cost. There's no subscription. When you need more, just ask.
What do you log?
Metadata for billing and your usage dashboard: model, token counts, cost, status, latency and timestamp. We do not store prompt or completion contents.
What happens if I run out of credits?
Requests return a standard OpenAI-style 402 insufficient_quota error until credits are added. Nothing is charged beyond your balance.
Are you affiliated with OpenAI?
No. Blizon is an independent service that provides access to OpenAI models through an OpenAI-compatible API.

Same tokens. 30% less.

Request access, get approved, change two lines. Every token after that costs 30% less.