Observed Signal · May 3, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Neutral
Gemini API Cheatsheet 2026: Models, Limits, Endpoints
A developer-focused cheatsheet that consolidates Google’s Gemini model names, recommended defaults, Google AI Studio free‑tier quotas, API endpoints, examples (REST, streaming, Rust reqwest), error codes, and token‑counting guidance. The article lists current Gemini models (e.g., gemini-2.5-flash-preview, gemini-1.5-pro), their context window sizes and recommended use cases, a free‑tier limits table with RPM/TPM/RPD values per model, sample cURL and Rust calls to the generativelanguage.googleapis.com endpoints, common HTTP error codes and fixes, and steps to obtain a free Google AI Studio API key. Published on 2026-05-03, the cheatsheet aims to be a single reference for developers integrating Google’s Gemini LLMs.
Consolidates Google Gemini model specifications, free‑tier quotas and API usage examples — a practical reference for developers integrating Google’s LLMs that can affect implementation and capacity planning.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Lists Gemini models and context sizes, including gemini-2.5-flash-preview (1M tokens) and gemini-1.5-pro (2M tokens).
- Reports Google AI Studio free-tier limits per model (examples: Gemini 2.5 Flash Preview — RPM 10, TPM 250,000, RPD 500; Gemini 1.5 Flash — RPM 15, TPM 1,000,000, RPD 1,500).
- Provides example REST and streaming API calls to generativelanguage.googleapis.com, including required headers and request bodies.
- Includes a Rust (reqwest + serde_json) example showing how to call the Gemini generateContent endpoint and extract response text.
- Documents common HTTP error codes (400, 403, 429, 500, 503) with suggested fixes and a token‑counting rough guide.
Connected Companies & Entities
4 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Google updates Gemini usage limits
Google changed usage limits for its Gemini apps, with the new rules taking effect on May 17, 2026. Limits are computed based on resource usage rather than a fixed request count: factors include request complexity, chat length, and use of features such as video, music, image generation, Deep Think, Deep Research and Pro models. Google says limits refresh every five hours until a user's weekly cap is reached. Free accounts have standard limits; paid tiers increase capacity (AI Plus = 2x, AI Pro = 4x, AI Ultra = 5x or 20x depending on plan). Users can view their consumption while logged into the same Google account at gemini.google.com/usage. The changes apply to users aged 18 and over.
Google Launches Gemini 3.8 Flash and Cyber AI Models
Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 3, 2026, both derived from the same foundation model. Flash is optimized for coding, reasoning, and AI agents, while Flash Cyber specializes in defensive cybersecurity and vulnerability patching. The models lead the HLE-Verified benchmark (54.9%) and perform strongly on DeepSWE v1.1, with Flash Cyber scoring 86.2% on Cybergym (beating GPT-5.5 Cyber) and 47.2% Pass@1 on CWE-Bench. Pricing remains $0.75 per million input tokens and $3.75 per million output tokens through 2026, then doubles in 2027, undercutting rivals like OpenAI and Anthropic. Cyber access is restricted via the Fairwind Program to select partners (e.g., CrowdStrike, Snowflake, government, critical infrastructure). Models are available across Google platforms, with MrBeast promoting them.
Google Restricts Free Access to Latest Gemini Models
Google will update its Gemini terms of service on October 9, 2026, limiting free users to the lightweight Flash-Lite model. Paid AI Plus subscribers ($4.99/month in the US) will retain Flash and Flash-Lite but lose access to Pro. AI Pro subscribers ($19.99/month) keep all three models and gain Deep Think mode, previously exclusive to AI Ultra. AI Ultra subscribers retain access to all models and get early access to Gemini 4 Argon. Google also introduces a reasoning effort slider across all models and may add three computational effort tiers. These changes likely free resources for Gemini 4 Argon's launch. Usage limits remain unchanged. Notably, OpenAI is moving in the opposite direction, offering GPT-6 to free users in a smaller Luna variant.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
