How to Use DeepSeek V4.1 Flash in Glarity: The Expiring 0910 Beta, Set Up in 4 Steps

Connect DeepSeek V4.1 Flash — deepseek-v4.1-flash-expires-on-0910 — to Glarity: one model ID, V4 Flash pricing, 20-request limit, a September 10 expiry.

DeepSeek's fastest model right now is also its shortest-lived. DeepSeek V4.1 Flash appeared in DeepSeek's live API on September 8, 2026 under the ID deepseek-v4.1-flash-expires-on-0910 — an internal beta whose name literally carries its own expiration date. If you already run a DeepSeek key in Glarity, connecting it is a one-field change: keep your key and your API host, and type the beta ID into Custom Model. That's the entire setup — no new endpoint, no waitlist, no new console. The price is the same as deepseek-v4-flash, and the beta is scheduled to go offline around September 10, so today is the day to test it if you want to try it at all.

What DeepSeek actually announced

DeepSeek's team posted the beta to its official community groups on September 8, and the announcement (relayed on Hacker News the same morning) is the whole official record so far:

DeepSeek V4.1 Flash is now available for internal beta testing. It uses a new model architecture, with native multimodal support, stronger capabilities, faster speed, and lower cost. Keep the base_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to use it. Pricing is currently the same as deepseek-v4-flash, with an account-level rate limit of 20 concurrent requests.

There is still no model card, no technical report, no changelog entry, and as of today DeepSeek's official Models & Pricing page lists only deepseek-v4-flash, deepseek-v4-pro and deepseek-v4-flash-vision-exp. So treat the paragraphs below the way DeepSeek does: the official claims come from that announcement, the speed numbers are first-day community measurements, and everything is fair game for your own two-day test.

The fact sheet

Field

Value

Model ID

deepseek-v4.1-flash-expires-on-0910

Status

Internal beta — an "intermediate version" (中间版本) checkpoint ahead of a final build

Announced

September 8, 2026, via DeepSeek's official community groups

Expires

September 10, 2026 — the ID encodes it; scheduled to go offline

Architecture

New model structure (per DeepSeek) — re-pretrained rather than a post-training refresh

Multimodal

Native image and text input in the base model (the Aug 21 vision model bolted a vision encoder onto the text model instead)

Context window

Not published. No figure has been released; keep Glarity's setting conservative

Pricing

Same rate card as deepseek-v4-flash (see below)

Rate limit

20 concurrent requests per account, vs 2,500 for deepseek-v4-flash

Where it runs

DeepSeek's own API — base URL and key unchanged

Connect it in four steps

Everything you need is a DeepSeek account: create or reuse an API key at platform.deepseek.com (DeepSeek's official platform — keys live there and usage is billed to that account).

  1. Open Glarity → Settings → General, scroll to Connect to AI for quick Glarity summaries and translations, and select "OpenAI API key" — the option that accepts custom keys from any AI vendor.
  2. Paste your DeepSeek API key into the API Key field.
  3. Model — choose Custom Model and type deepseek-v4.1-flash-expires-on-0910 exactly, dots included (.). This ID will not appear in the dropdown: it is a hidden beta, and early testers on Hacker News had to type it manually on purpose.
  4. Confirm the three fields that stay put — API Host https://api.deepseek.com, API Path /v1/chat/completions, Model Context Window 204800 tokens — set Temperature to 0.5 (drop to 0.2 for literal, word-for-word translation), then hit Test and Save.

Because DeepSeek says "keep the base_url unchanged", the beta deliberately reuses the exact deepseek-v4-flash wiring. If you already followed our DeepSeek V4 Flash setup guide, you only change step 3 now — and changing it back to deepseek-v4-flash later is the same one-field move.

What it costs: the V4 Flash rate card

The beta bills at the official deepseek-v4-flash rates from DeepSeek's pricing page. DeepSeek has run peak/off-peak pricing since August 16, 2026; all other hours are off-peak at half the peak rate (peak: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday).

Per 1M tokens

Off-peak

Peak

Input (cache hit)

$0.007

$0.014

Input (cache miss)

$0.22

$0.44

Output

$0.66

$1.32

The math for real Glarity sessions, off-peak:

Job

Tokens in

Tokens out

V4.1 Flash (Flash rate)

Same job on V4 Pro

1-hour YouTube video → summary + Q&A

~12,000

~1,000

~$0.003

~$0.010

30-page report → full summary

~40,000

~1,500

~$0.010

~$0.029

Book-length document → chapter notes

~200,000

~3,000

~$0.046

~$0.138

One honest warning before you load up your balance: "lower cost" is an efficiency claim, not a rate cut. The per-token price did not change; the model just generates much faster. Early testers watching a wall-clock session saw their balance drain noticeably faster than the same minutes on V4 Flash, because the beta spends tokens at a higher rate. Whether it is genuinely cheaper per finished task depends on it finishing with fewer tokens and fewer retries — which is exactly what you should measure.

What to test today, and what early numbers say

Speed is the headline, and every speed claim right now is a community measurement. The researcher whose X post circulated with the beta measured roughly 350 tokens/second average decoding; first-day developers reported ~420 output tokens/s in long generations with peaks above 500, including one widely shared test that generated more than 71,000 tokens in under three minutes. For scale, Artificial Analysis measures the public DeepSeek V4 Flash 0731 endpoint at roughly 128 t/s, so early reports point to a throughput jump of three times or more. Same-task comparisons against the August 21 vision model ran even wider — roughly 6× faster on an SVG-generation task, about 5× on a 49K-token long-context retrieval, and around 5× on a large SQL generation task. None of this is official: these are uncontrolled developer measurements, DeepSeek has published no speed benchmark, and the fastest test candidate is not always the build that ships.

Quality has no published benchmark. There is one tantalizing detail though: DeepSeek's feedback form for the beta asks testers whether V4.1 Flash "can fully replace the online DeepSeek V4 Pro" — a question DeepSeek put to its own users, not a claim about the model. Treat it as an open question, not a spec.

Run your own 48 hours. In Glarity, the beta is a low-risk test: summarize a one-hour lecture transcript, translate a long article, ask follow-up questions, then re-run the same jobs on deepseek-v4-flash and compare. Glarity's custom connection sends the text of pages and transcripts to the model — V4.1's native image input is available when you call the API directly with files, not through the extension's page flow.

Caveats before you rely on it

  • It expires. The ID says expires-on-0910; the test is scheduled to go offline around September 10, leaving roughly two days of uptime from the September 8 launch.
  • 20 concurrent requests per account. Fine for one person browsing; wrong place for a shared team workflow.
  • No specs. Parameter count, context length, and the full input-modality list are unpublished.
  • Don't wire production. No one should point a permanent workflow at a model ID that is scheduled to stop answering; treat the beta as a test, and keep deepseek-v4-flash as your daily driver.
  • A final release is expected soon — tipsters flag a launch post "very soon," and the short beta window fits that — but that expectation is unconfirmed. When the GA model lands, this setup works the same way: swap the model ID (or pick it from the dropdown if it appears).

If something fails

Symptom

Likely cause

Fix

Test returns 401

Wrong key, or a key pasted with a trailing space

Regenerate the key at platform.deepseek.com and paste over the field

"Model not found" / "model does not exist"

Typo in the ID — or the beta has expired

Type deepseek-v4.1-flash-expires-on-0910 exactly, dots included; after Sept 10, switch to deepseek-v4-flash until the final release

Same error before Sept 10 in the dropdown

The ID never appears in dropdowns — it is a hidden beta, manual entry only

Use Custom Model + typed ID, as in step 3

429 or rate limit

20 concurrent requests per account

Wait a few seconds; avoid parallel summaries and translations at once

Connection error on Test

Wrong host or path left over from another provider

API Host https://api.deepseek.com, API Path /v1/chat/completions

Balance drops much faster than usual

High token throughput — the model is 3×+ faster, not cheaper per token

Test off-peak, and compare tokens per finished task against V4 Flash

FAQ

Is DeepSeek V4.1 Flash free? No. The beta bills at deepseek-v4-flash rates: $0.22 per 1M input (cache-miss) and $0.66 per 1M output off-peak, half those rates being peak. There is no free window.

Will the beta keep working after September 10? Per the announcement and the ID, no — the test version is scheduled to expire automatically and go offline around September 10. Change the model field back to deepseek-v4-flash, and watch for the final release.

Do I need a new DeepSeek account or key for it? No. It runs on the same API, the same key, and the same platform account; only the model name changes.

Why isn't it in Glarity's model dropdown? It is a hidden internal beta and was never added to any public model list. Custom Model plus the exact typed ID is the only way in.

Should I switch my daily Glarity usage to it? Not for production-style workflows. Use the two-day window to compare quality and real per-task cost against deepseek-v4-flash; the 20-request limit and the built-in expiry make it a test, not a daily driver.

Does V4.1 Flash read images in the Glarity extension? DeepSeek describes native image-and-text input, so API calls with files work. Glarity's custom model connection summarizes and translates the text of pages and transcripts; use the API directly with image files when you want to exercise the vision side.

How do I switch back to the stable DeepSeek model? One field: set Model to deepseek-v4-flash and leave the key, host and path untouched. The endpoint, temperature and context window all stay valid.

By the Glarity Editorial Team. The Glarity Editorial Team writes about AI search, video summarization, and getting more from your browser.

Glarity is a free browser extension for AI search and YouTube summarization — try it at glarity.app.