Google's newest Flash model, Gemini 3.8 Flash, launched on September 2, 2026 — and it's a natural fit for Glarity: a 1,048,576-token context window, 65,536 tokens of output, text/image/video/audio/PDF input, a free tier, and the same introductory pricing as 3.7 Flash. Wiring it in takes about five minutes: in Settings → General, choose OpenAI API key, paste a key from Google AI Studio, and fill in the model name, context window, API host and path using Google's official OpenAI-compatible endpoint. Then every Glarity summary and translation runs on Gemini 3.8 Flash.
What's new in Gemini 3.8 Flash
Gemini 3.8 Flash is Google's "most intelligent Flash model," built for long-horizon software engineering, autonomous agents, and complex enterprise workflows — but it's also the reading-and-writing workhorse for Glarity users. It's a stable model (the newest entry in the Gemini 3 line), takes text, images, video, audio and PDFs as input, returns text, and Google's launch post highlights a 54.9% on HLE-Verified (multi-step reasoning across STEM, humanities and professional fields) and strong long-horizon software-engineering scores (DeepSWE v1.1). Google also shipped a security-focused 3.8 Flash Cyber variant alongside it, available only through the Fairwind Program for trusted defenders — the standard 3.8 Flash remains the one you connect to Glarity.
One design point matters for daily use: Google says 3.8 Flash "works harder" — it executes extra reasoning steps and calls tools iteratively, using more tokens at higher effort levels to maximize performance. Effort is limited to low / medium / high (a "minimal" effort setting is not supported). For efficiency-first workloads, Gemini 3.7 Flash remains fully supported — that's a great tradeoff to know about before you pick your default (more below).
What you need before you start
- A key from AI Studio — Google's official developer console for the Gemini API.
- The connection values below.
Configure Glarity step by step
Open Settings → General, scroll to Connect to AI for quick Glarity summaries and translations.
- Select "OpenAI API key".
- Paste your Google API key into the API Key field.
- Model — choose Custom Model and enter
gemini-3.8-flash. - Model Context Window — pick the preset nearest to 1,048,576 tokens (the official input limit). Long pages and full-hour transcripts fit comfortably.
- API Host —
https://generativelanguage.googleapis.com/v1beta/openai— Google's official OpenAI-compatible endpoint (the same one used with the 3.7 Flash settings; it serves the whole Gemini family). Google still describes this layer as beta; that's fine for reading and translation tasks. - API Path —
/chat/completions. - Temperature — 0.5. Standard parameters pass through the compatible endpoint; if your summaries ever look too loose, lower toward 0.2 for more repeatable phrasing.
- Test, then Save.
Settings at a glance
Field | Value | Meaning |
|---|---|---|
Connection | OpenAI API key | via Google's compatible endpoint |
Model | Custom Model — | Stable Gemini 3.8 Flash |
Model Context Window | 1,048,576 tokens (preset as available) | Official input limit |
API Host |
| Official OpenAI-compatible base |
API Path |
| Chat Completions route |
Temperature | 0.5 | Stable output; lower toward 0.2 if needed |
Pricing and the free tier
Per Google's official Gemini pricing page, Gemini 3.8 Flash on the Standard tier is $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026 — the same introductory rate as 3.7 Flash — after which the rates double to $1.50/$7.50. A free tier exists for Standard and Priority service (Batch and Flex are not available on it), so light use costs nothing. Context caching input is $0.075 per 1M through 2026 (then $0.15), which helps when Glarity re-reads the same document. Output for summaries and translations is short relative to input, so the numbers generally stay small; check the official page for figures before you budget.
Long-context economics
Full-hour video transcripts and very long documents both fit inside the 1M window — no splitting, no truncation. Glarity hands the whole thing to the model in one go, and the model sees the beginning and end of the document at once, which improves faithfulness on long pieces. The 65,536-token output cap is generous for summaries, subtitles and Q&A; it's worth knowing the limit exists if you ever generate near-book-length rewrites in one shot.
3.8 Flash or 3.7 Flash?
The honest recommendation: 3.8 Flash for quality-first work, 3.7 Flash for efficiency-first work. The "works harder" design means deeper reasoning and more token spend at higher effort levels. Glarity's default workload — summary, translation, subtitle generation, Q&A — is fast and inexpensive on both; reach for 3.8 when accuracy on tricky, multi-step content matters most, and stick with 3.7 (or drop effort to low) when you're processing many pages a day and token cost is the primary constraint. Switching between them is a one-field edit in Glarity's Model field — same host, same path, same key. See the Gemini 3.7 Flash setup guide for the older model's exact values.
What a Gemini-powered Glarity can do
Capability | How it behaves now |
|---|---|
Translation | Immersive page translation and side-by-side subtitles — see the bilingual subtitle tutorial |
Page summary | Whole documents summarized, including very long pages |
YouTube summary | Timestamped highlights and FAQ from full transcripts |
Video subtitles | Subtitles generated and translated into your language |
Quick email reply | Drafts in the tone and language you want |
One-click prompts | Saved tasks re-run on your own model |
Troubleshooting
- Test returns 401 / invalid key — the key is wrong or spaces crept in. Create a fresh key in AI Studio.
- Frequent 429 / quota errors — you're on the free tier quota or a busy period. Wait, or consider upgrading via the AI Studio console.
- "Model not found" — check the exact spelling
gemini-3.8-flash. - Costs jump in 2027 — yes: promotional rates end December 31, 2026; budget accordingly, or switch models in Glarity's Model field later.
- Nothing changes after saving — confirm you pressed Save after Test.
FAQ
Does Gemini 3.8 Flash do translation or just summaries? Both — one connected model covers Glarity's page translation, selection translation, subtitle translation and summaries.
Is the free tier enough? For occasional use, yes. For daily reading workflows, the paid tier's promotional rates are still very cheap.
Why 1M context instead of a smaller window? Long pages and hour-long transcripts fit whole — no splitting, no dropped context. If you never process long documents, any model works; the big window is the reason to pick Flash.
Can I switch to a Pro model later? Yes — edit the Model field and keep the same host and path; the compatible endpoint serves Google's Gemini-family models.
Is the OpenAI-compatible endpoint stable enough for regular use? Google documents it as beta, but it's the standard path for third-party clients like Glarity; testing with a real page first is the safest way to know.
Glarity's editorial team covers AI search, video summaries, and browser productivity tips.
Glarity is a free browser extension that supports AI search and YouTube video summaries, available at glarity.app.



