llm
All models
chat, vision
REQUEST
llm
Context window
Max output
Credits
9.375 credits / 1K input tokens, 56.25 credits / 1K output tokens
USD equivalent
$1.50 / $9.00 per 1M tokens (input/output)
MODEL NOTES
Gemini 3.5 Flash is Google's fast multimodal model, and its standout feature is context: a 1,000,000 token window, roughly five times larger than Claude's and eight times larger than GPT-4o's. That's enough to fit an entire book, a large codebase, or hours of transcript in a single request.
If your use case involves genuinely large inputs — analyzing a full repository, summarizing a lengthy document collection, or maintaining very long conversation history — this is the model built for that, not a workaround.
For typical short-context tasks, it's not automatically the best choice; the value here is specifically the context window, not raw output quality relative to Claude or GPT-4o.