OpenAI-compatible model prices move fast, and every vendor prices them differently — promo windows, cache tiers, peak and off-peak periods. This is the August 2026 menu: six models, each with an OpenAI-compatible endpoint you can plug straight into a browser extension like Glarity (one key, one API Host, one path). The short version: Claude Opus 5 ($5/$25 per million tokens) and GPT-5.6 Sol ($4/$20) are the premium tier; Gemini 3.7 Flash, GLM-5.3-Flash and DeepSeek V4 Flash are the flash tier; Qwen3.6-Plus ($0.50/$3.00) splits the difference with a 1M-token input window.
The six models, priced side by side
All prices are official vendor rates as of August 28, 2026, per 1M tokens (MTok). Sources: Anthropic, OpenAI, Google, Zhipu, Alibaba Cloud and DeepSeek:
Model | Input / Output per MTok | Context / Max output | Notes |
|---|---|---|---|
Claude Opus 5 | $5.00 / $25.00 | 1M / 128K | Same price as Opus 4.8; thinking on by default |
GPT-5.6 Sol | $4.00 / $20.00 (short) · $8.00 / $30.00 (long context) | — | Cached input $0.40; promotional pricing at least through Nov 21, 2026 |
Gemini 3.7 Flash | $0.75 / $3.75 (promo) | — | Promo runs through Dec 31, 2026, then $1.50 / $7.50; cache $0.075; free tier available |
GLM-5.3-Flash | 1/10 of GLM-5.3 (promo: 1/20) | 1M / 128K | Zhipu's official page: ~1/40 of Claude Opus 4.8 at the same intelligence; MIT open weights |
Qwen3.6-Plus | $0.50 / $3.00 (≤256K input) · $2.00 / $6.00 (256K–1M) | 1M / — | International (Singapore) rates |
DeepSeek V4 Flash | $0.22 / $0.66 off-peak · $0.44 / $1.32 peak | 1M / 384K | Cache-hit input from $0.007; peak UTC Mon–Fri 01:00–04:00 & 06:00–10:00 |
How much does 100M tokens cost a month?
A realistic summarizer workload: 100M tokens of input (pages, transcripts, documents) and 10M tokens of output (summaries, answers):
Model | 100M in + 10M out |
|---|---|
Claude Opus 5 | $750 |
GPT-5.6 Sol | $600 (short-context standard) |
Gemini 3.7 Flash | $112.50 (promo; $225 after Dec 31, 2026) |
Qwen3.6-Plus | $80 (requests ≤256K input; $260 if most exceed it) |
DeepSeek V4 Flash | $28.60 off-peak ($57.20 peak; cache-hit input drops this toward the $10s) |
GLM-5.3-Flash | Official pricing is relative only — Zhipu states 1/10 of GLM-5.3, 1/20 in promo, ≈1/40 of Opus 4.8 at the same intelligence |
Input-heavy traffic (translation, re-reading the same source page) rewards models with cheap or cached input; output-heavy traffic (rewrites, long summaries) rewards cheap output.
Which model for which job?
Job | Pick | Why |
|---|---|---|
Long-document research, serious verification | Claude Opus 5 | 1M context, thinking on by default — spend $5/$25 where correctness matters |
Big-file extraction from massive files | GPT-5.6 Sol | Strong long-context tier; $4/$20 short-context standard while the promo runs |
Whole-page / whole-document translation | DeepSeek V4 Flash | Off-peak $0.22 in, cache-hit input $0.007 — re-reading the same source article is nearly free |
Everyday page & YouTube summarization | Gemini 3.7 Flash or GLM-5.3-Flash | Under $1 input at promo; GLM-5.3-Flash adds MIT open weights at 1/10 of GLM-5.3 |
High-volume casual Q&A, light drafts | Qwen3.6-Plus | $0.50 / $3.00 for requests up to 256K input |
Chronic transcript workloads | DeepSeek V4 Flash off-peak | $0.66 output per MTok — cheapest listed output among the six |
Why OpenAI compatibility matters in a browser extension
Glarity's Settings → General → Connect to AI dialog takes an OpenAI API key, a model name, a context-window preset, an API Host and an API Path — so any vendor with an OpenAI-style /chat/completions endpoint works. All six above offer one: Anthropic publishes an OpenAI SDK compatibility layer for https://api.anthropic.com/v1/ (documented for testing and comparison, which is exactly what an extension's model picker does), Google's Gemini API is OpenAI-compatible, Zhipu documents an OpenAI-compatible base for GLM models, Alibaba Cloud Model Studio and DeepSeek are OpenAI-compatible by design, and OpenAI is OpenAI.
When you switch models, Glarity keeps the same task — summary, translation, Q&A — and the only things that change are the host, the key and the price per token. See how to use Qwen3.6-Plus in Glarity and how to use GLM-5.3-Flash in Glarity for step-by-step setup.
FAQ
Which of these is the cheapest in 2026? It depends on the shape of your traffic. For input-heavy jobs, DeepSeek V4 Flash has the lowest official rates among the six off-peak ($0.22 per MTok, $0.007 on cache hits). For mixed everyday summaries, Gemini 3.7 Flash's promo ($0.75/$3.75) and Qwen3.6-Plus ($0.50/$3.00, requests up to 256K) are the cheapest with no time gates.
Will these prices change? Yes — several are promotional. GPT-5.6 Sol's promo pricing runs at least through Nov 21, 2026; Gemini 3.7 Flash's rates double on Jan 1, 2027; DeepSeek V4 Flash is cheaper outside peak hours. Check the vendor's official pricing page before committing.
Which of these are open-source? GLM-5.3-Flash — Zhipu released the weights under MIT, so it can also run outside the API. The others are API-only per their official pages.
Do these models work inside Glarity? Yes. Set Settings → General → Connect to AI to OpenAI API key, add the vendor's key, pick Custom Model, set the context-window preset to the model's input limit, and fill in the API Host and API Path from the vendor's compatibility docs. Test, then Save.
Should I watch input or output price? Both, but the mix depends on the task. Summarizing a YouTube video reads a long transcript and writes a short summary — input dominates. Rewriting or translating a whole document produces many output tokens, so output price dominates.
Glarity's editorial team covers AI search, video summarization and browser productivity.
Glarity is a free browser extension for AI search and YouTube summarization, available at glarity.app.



