Headroom

Headroom · for Gemini people

Gemini usage limits, metered.

The 5-hour window, the weekly cap and the CLI's daily quota, live in the Windows taskbar — read from Google's own reporting, not reconstructed from token counts.

Windows 10 & 11 · signed auto-updates · no account needed

Gemini usage limits, explained

Two compute windows for the app, a daily request quota for the CLI — and no published compute numbers for the windows, only multipliers.

Google stopped counting prompts in late May 2026. Gemini usage limits are now compute budgets: each request draws an amount that depends on the prompt's complexity, the features it touches, and how long the conversation has been running. The budget refreshes every 5 hours until a separate weekly cap ends the refills. Absolute numbers are published nowhere — Google states only multipliers, so a paid plan changes the size of your allowance, not the mechanics: Plus gets 2x the standard allowance, Pro 4x, Ultra 5x or 20x of Pro depending on the subscription.

The command line plays by different rules. Gemini CLI meters model requests per day against your signed-in tier, and a Gemini API key has its own per-model rate limits, managed in AI Studio. Three surfaces, three unrelated quotas — which is why “how much Gemini do I have left” has no single answer, and why Google's help page warns that limits may change without notice.

5 h · refresh

The 5-hour usage limit

Every Gemini Apps request draws from this budget, and the draw varies: media generation and Deep Research consume more, Flash-Lite consumes nothing, failed requests don't count, and a single prompt's consumption is capped. It refreshes every 5 hours on every tier. When it's empty, subscribers can keep going on Flash-Lite; everyone else waits out the countdown.

7 days

The weekly cap

The 5-hour refresh only restores usage until you've spent the week's total; after that, premium models are gated until the weekly reset. Google doesn't document whether the week rolls or resets on a fixed day — its dashboard reports a timestamp, and that's all anyone gets. This is the limit that ends workdays, because there's no waiting it out over lunch.

daily

Gemini CLI model requests

A free personal account gets 1,000 model requests a day and 60 a minute. Paid plans raise the ceiling: Google's published Code Assist tiers run 1,500 and 2,000 requests a day, and AI Pro and Ultra reportedly match them. One prompt can fire several model requests, and agent mode shares the same pool. Run /model to see each bucket's used percentage and reset countdown — the same rows Headroom reads.

fixed

Context window per plan

Not a timed window but a hard cap on a single conversation: 32,000 tokens free, 128,000 on Plus, 1 million on Pro and Ultra. There's no countdown to wait for; when a chat hits the ceiling, you start a fresh one.

What Headroom shows for Gemini

Google's reported percentages and reset countdowns, visible while you work.

Both windows in the tray

Idle, the tray icon shows twin bars: tightest session on top, tightest week below. While you're actively prompting it switches to one big number — percent left on whichever Gemini window bites first. If Google starts reporting a new limit kind, it appears as its own row automatically, tagged new — no app update needed.

CLI quotas without typing /model

Headroom launches the official Gemini CLI, sends /model, and turns each model row into a meter of its own — Pro and Flash drain separately, so you can see which quits first. Your OAuth token is never read or replayed. A Gemini API key gets a simpler gauge: key validated, models counted, billing left where Google keeps it, in the AI Studio dashboard.

A toast before the week runs dry

One warning at your threshold, one at critical, and a 6-hour cooldown per quota between them. Take the weekly one seriously: unlike the 5-hour window, waiting doesn't fix it.

Fair questions

Does Gemini have a usage limit?

Yes. Since late May 2026 Gemini Apps budgets compute instead of counting prompts: your allowance refreshes every 5 hours until you reach a separate weekly cap. What each request costs depends on prompt complexity, the features involved (media generation and Deep Research draw more) and how long the chat has run. Paid plans multiply the standard allowance: Plus 2x, Pro 4x, Ultra 5x or 20x of Pro.

When does my Gemini limit reset?

The main allowance refreshes every 5 hours; the weekly cap stops those refreshes once it's spent, and Google doesn't say whether the week rolls or resets on a fixed day — it reports a timestamp and nothing more. You can check the countdown at gemini.google.com/usage, or with /model inside Gemini CLI for its separate daily quota. Headroom keeps both countdowns in the Windows taskbar.

How many prompts do I get with Google AI Pro?

There's no prompt count anymore. The “100 prompts a day” figure still quoted around the web describes the system Google retired in May 2026. AI Pro now gets 4x the standard compute-based allowance across the 5-hour and weekly windows, plus a 1-million-token context window; what the standard allowance actually is, Google doesn't publish — only the multiplier.

Why did I hit my Gemini limit so fast?

Compute-based limits mean one expensive request can eat a disproportionate share — long conversations, media generation and Deep Research all draw far more than a short text prompt. Google has since capped how much a single prompt can consume and stopped counting failed requests, which softens the worst of it. Subscribers can also drop to Flash-Lite, which draws nothing from the quota.

What are the Gemini CLI daily limits?

A free personal Google account gets 1,000 model requests a day and 60 a minute. Paid plans raise the daily cap: Google's published Code Assist tiers run 1,500 and 2,000 requests a day, and AI Pro and Ultra are reported to match them. One prompt can trigger several model requests, and agent mode draws from the same pool, so your effective prompt budget is lower than the raw number.

Are Gemini app and Gemini CLI limits the same?

No — three separate systems. The Gemini app uses compute-based 5-hour and weekly windows; Gemini CLI meters model requests per day by sign-in tier, bucketed per model; a Gemini API key has its own rate limits, managed and billed in Google AI Studio. Emptying one doesn't touch the others.

Gemini is rarely the only meter running. Headroom also watches: