AI Model Pricing in 2026: 6 OpenAI-Compatible Models Compared

Claude Opus 5, GPT-5.6 Sol, Gemini 3.7 Flash, GLM-5.3-Flash, Qwen3.6-Plus & DeepSeek V4 Flash priced side by side, with a scene picker and 100M-token budget math.

OpenAI-compatible model prices move fast, and every vendor prices them differently — promo windows, cache tiers, peak and off-peak periods. This is the August 2026 menu: six models, each with an OpenAI-compatible endpoint you can plug straight into a browser extension like Glarity (one key, one API Host, one path). The short version: Claude Opus 5 ($5/$25 per million tokens) and GPT-5.6 Sol ($4/$20) are the premium tier; Gemini 3.7 Flash, GLM-5.3-Flash and DeepSeek V4 Flash are the flash tier; Qwen3.6-Plus ($0.50/$3.00) splits the difference with a 1M-token input window.

The six models, priced side by side

All prices are official vendor rates as of August 28, 2026, per 1M tokens (MTok). Sources: Anthropic, OpenAI, Google, Zhipu, Alibaba Cloud and DeepSeek:

Model

Input / Output per MTok

Context / Max output

Notes

Claude Opus 5

$5.00 / $25.00

1M / 128K

Same price as Opus 4.8; thinking on by default

GPT-5.6 Sol

$4.00 / $20.00 (short) · $8.00 / $30.00 (long context)

Cached input $0.40; promotional pricing at least through Nov 21, 2026

Gemini 3.7 Flash

$0.75 / $3.75 (promo)

Promo runs through Dec 31, 2026, then $1.50 / $7.50; cache $0.075; free tier available

GLM-5.3-Flash

1/10 of GLM-5.3 (promo: 1/20)

1M / 128K

Zhipu's official page: ~1/40 of Claude Opus 4.8 at the same intelligence; MIT open weights

Qwen3.6-Plus

$0.50 / $3.00 (≤256K input) · $2.00 / $6.00 (256K–1M)

1M / —

International (Singapore) rates

DeepSeek V4 Flash

$0.22 / $0.66 off-peak · $0.44 / $1.32 peak

1M / 384K

Cache-hit input from $0.007; peak UTC Mon–Fri 01:00–04:00 & 06:00–10:00

How much does 100M tokens cost a month?

A realistic summarizer workload: 100M tokens of input (pages, transcripts, documents) and 10M tokens of output (summaries, answers):

Model

100M in + 10M out

Claude Opus 5

$750

GPT-5.6 Sol

$600 (short-context standard)

Gemini 3.7 Flash

$112.50 (promo; $225 after Dec 31, 2026)

Qwen3.6-Plus

$80 (requests ≤256K input; $260 if most exceed it)

DeepSeek V4 Flash

$28.60 off-peak ($57.20 peak; cache-hit input drops this toward the $10s)

GLM-5.3-Flash

Official pricing is relative only — Zhipu states 1/10 of GLM-5.3, 1/20 in promo, ≈1/40 of Opus 4.8 at the same intelligence

Input-heavy traffic (translation, re-reading the same source page) rewards models with cheap or cached input; output-heavy traffic (rewrites, long summaries) rewards cheap output.

Which model for which job?

Job

Pick

Why

Long-document research, serious verification

Claude Opus 5

1M context, thinking on by default — spend $5/$25 where correctness matters

Big-file extraction from massive files

GPT-5.6 Sol

Strong long-context tier; $4/$20 short-context standard while the promo runs

Whole-page / whole-document translation

DeepSeek V4 Flash

Off-peak $0.22 in, cache-hit input $0.007 — re-reading the same source article is nearly free

Everyday page & YouTube summarization

Gemini 3.7 Flash or GLM-5.3-Flash

Under $1 input at promo; GLM-5.3-Flash adds MIT open weights at 1/10 of GLM-5.3

High-volume casual Q&A, light drafts

Qwen3.6-Plus

$0.50 / $3.00 for requests up to 256K input

Chronic transcript workloads

DeepSeek V4 Flash off-peak

$0.66 output per MTok — cheapest listed output among the six

Why OpenAI compatibility matters in a browser extension

Glarity's Settings → General → Connect to AI dialog takes an OpenAI API key, a model name, a context-window preset, an API Host and an API Path — so any vendor with an OpenAI-style /chat/completions endpoint works. All six above offer one: Anthropic publishes an OpenAI SDK compatibility layer for https://api.anthropic.com/v1/ (documented for testing and comparison, which is exactly what an extension's model picker does), Google's Gemini API is OpenAI-compatible, Zhipu documents an OpenAI-compatible base for GLM models, Alibaba Cloud Model Studio and DeepSeek are OpenAI-compatible by design, and OpenAI is OpenAI.

When you switch models, Glarity keeps the same task — summary, translation, Q&A — and the only things that change are the host, the key and the price per token. See how to use Qwen3.6-Plus in Glarity and how to use GLM-5.3-Flash in Glarity for step-by-step setup.

FAQ

Which of these is the cheapest in 2026? It depends on the shape of your traffic. For input-heavy jobs, DeepSeek V4 Flash has the lowest official rates among the six off-peak ($0.22 per MTok, $0.007 on cache hits). For mixed everyday summaries, Gemini 3.7 Flash's promo ($0.75/$3.75) and Qwen3.6-Plus ($0.50/$3.00, requests up to 256K) are the cheapest with no time gates.

Will these prices change? Yes — several are promotional. GPT-5.6 Sol's promo pricing runs at least through Nov 21, 2026; Gemini 3.7 Flash's rates double on Jan 1, 2027; DeepSeek V4 Flash is cheaper outside peak hours. Check the vendor's official pricing page before committing.

Which of these are open-source? GLM-5.3-Flash — Zhipu released the weights under MIT, so it can also run outside the API. The others are API-only per their official pages.

Do these models work inside Glarity? Yes. Set Settings → General → Connect to AI to OpenAI API key, add the vendor's key, pick Custom Model, set the context-window preset to the model's input limit, and fill in the API Host and API Path from the vendor's compatibility docs. Test, then Save.

Should I watch input or output price? Both, but the mix depends on the task. Summarizing a YouTube video reads a long transcript and writes a short summary — input dominates. Rewriting or translating a whole document produces many output tokens, so output price dominates.

Glarity's editorial team covers AI search, video summarization and browser productivity.

Glarity is a free browser extension for AI search and YouTube summarization, available at glarity.app.