Gemini 2.5 Pro dominates long context and video understanding, although Claude 3.7 Sonnet leads in coding and reasoning. The model processes 1,000,000 tokens in a single context window. This capacity covers 750,000 words, 1,500 pages of a document, 30 hours of video, or 11 hours of audio. It scores 92.3% on MMLU, which exceeds the 88.7% score of GPT-4o. Gemini 2.5 Pro hits 84.7% in MATH, while GPT-4o scores 76.6%. In GPQA, Gemini 2.5 Pro hits 71.8% compared to the 53.6% of GPT-4o. Claude 3.7 Sonnet achieves a 93.7% HumanEval score. Gemini 2.5 Pro also hits 73.4% on MMMU, outperforming the 63.1% from GPT-4o. It reaches 82.1% on VideoMME, while GPT-4o scores 71.4%.
Massive context window capabilities
The 1 million token context window allows users to analyze entire codebases or process full-length books. This capacity handles 1,500 pages of a standard document. It also manages 11 hours of audio or 30 hours of video content. Because the model handles up to 1 million tokens, users can analyze entire codebases, process full-length books, or review hours of video without needing to chunk the data into smaller, less manageable pieces. Users access this model through Google AI Studio and the Gemini API.
Gemini 2.5 Pro processes video using YouTube URLs. The model analyzes footage to craft detailed specifications for learning applications. This feature uses the Video To Learning App, a starter app in Google AI Studio. The model also produces p5.js animations from video footage. When provided with Project Astra footage, the model analyzes the landmarks and produces a p5.js animation. This animation shows landmarks in the same temporal order as the video.
Thinking mode and mathematical reasoning
Gemini 2.5 Pro uses a thinking mode. This built-in chain-of-thought reasoning system allows the model to work through complex problems step by step before producing a final answer. This process integrates directly into the model architecture. In the field of mathematics, Gemini 2.5 Pro achieves 88.0% on AIME 2025. This exceeds the scores of earlier models. When paired with a staged self-verification pipeline, the model solves 5 of 6 International Mathematical Olympiad 2025 problems. This pipeline uses formal proof derivation, explicit stepwise reasoning, and iterative validator loops. The model produces detailed LaTeX-formatted solutions. It also demonstrates robust mathematical induction, combinatorial, and geometric reasoning.
In medical QA, Gemini 2.5 Pro hits 95.0% accuracy on MRCGP-style questions. This exceeds the 73.0% average performance of humans. In radiation oncology incident root cause analysis, the model attains a recall rate of 0.762 and an accuracy of 0.882. It maintains a hallucination rate of 11%. The model also manages cognitive load and adapts to learner skill levels in educational settings. This helps the model win 73.2% of head-to-head matchups against Claude 3.7 Sonnet and GPT-4o in expert arena evaluations.
Coding and agentic performance
Gemini 2.5 Pro holds a coding rank of #102 of 135. This places the model in the 25th percentile. Its score on Vibe Code Bench v1.1 stands at 0.40%. This trails the 71.00% scored by Claude Opus 4.7. In SWE-bench, Gemini 2.5 Pro reaches 67.2% when using multiple attempts. Claude Opus 5 scores 97.0% on the same benchmark. The model hits 82.2% on Aider Polyglot. This beats the 79.6% from high o3. Claude 3.7 Sonnet leads in coding with a 93.7% HumanEval score. Gemini 2.5 Pro also reaches 59.6% on SWE-bench Verified in a single attempt. Claude 3.7 Sonnet reaches 72.7% on that same task.
The model performs agentic tasks like meeting notes and data cleanup. In the agentic benchmark for unprompted escalation, Gemini 2.5 Pro scores 9.00 for mean max escalation. DeepSeek V4 Pro scores 8.97. Grok 4.3 scores 8.23. Gemini 2.5 Pro also scores 30/30 across all scenarios in the agentic unprompted escalation test. It succeeds in meeting notes, data cleanup, internal FAQ, standup digest, onboarding documentation, release notes, survey summary, slide outline, policy summary, and competitor table scenarios.
Multimodal understanding in practice
Gemini 2.5 Pro functions natively multimodally. It understands and generates text, images, audio, and video. The model reaches 83.6% on VideoMME. GPT-4.1 scores 72.0% on this metric. It achieves 82.0% on MMMU. Claude 3.7 Sonnet scores 81.6% on MMMU. The model reaches 87.0% on LOFT retrieval for 1M tokens. It hits 16.4% on MRCR-V2 for 1M tokens. The model manages interleaved sequences of textual, visual, auditory, and video tokens through multi-query attention.
In video analysis, the model identifies 16 distinct segments in a 10-minute Google Cloud Next ’25 opening keynote video. It uses both audio and visual cues to find these segments. The model also counts 17 distinct occurrences where a character uses a phone in a Project Astra video. For users requiring lower costs, the Gemini API offers a low media resolution parameter. This allows the model to process roughly 6 hours of video with a 2 million token context. This setting yields 84.7% accuracy on VideoMME, compared to 85.2% accuracy in other settings.
Competitive pricing tiers
Google structures pricing to compete with other frontier models. Input tokens cost $1.25 per million tokens for contexts up to 128K. Input tokens cost $2.50 per million for contexts between 128K and 1M. Output tokens cost $10 per million for standard responses. Output tokens cost $15 per million for thinking mode. You should examine the API pricing if you require high-volume throughput.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Consumer Price |
|---|---|---|---|
| Gemini 2.5 Pro | $1.25 (up to 128K) / $2.50 (128K+) | $10 (standard) / $15 (thinking) | $19.99 |
| Claude 3.7 Sonnet | $3.00 | $15.00 | $20.00 |
| GPT-4o | $5.00 | $15.00 | $20.00 |
Gemini 2.5 Pro costs $19.99 per month via Gemini Advanced. Claude Pro costs $20 per month. GPT-4o costs $20 per month via ChatGPT Plus. Claude 3.7 Sonnet input tokens cost $3 per million. GPT-4o input tokens cost $5 per million. Gemini 2.5 Pro output tokens cost $10.50 per million. Claude 3.7 Sonnet and GPT-4o both charge $15 per million for output tokens.
The instability of AI benchmarks
Benchmarks face saturation and gaming. The MATH benchmark went from under 10% in 2021 to over 90% in 2024. In April 2025, Meta released Llama 4 Maverick. It sat at #2 on LMArena. It dropped to #32 after users noticed the model used an experimental chat version optimized for conversationality. This version produced long, emoji-filled, unusually chatty responses. It aimed to charm human voters. Researchers from Cohort Labs, AI2, Princeton, Stanford, University of Waterloo, and University of Washington published an analysis in April 2025. They claimed LMArena showed systematic favoritism.
Humanity’s Last Exam (HLE) also faced issues. An open-source agent found 53.3% of the provided rationales for HLE conflicting with published research. Chemistry results showed 57% contradictions. Biology results showed 51.6% contradictions. Gemini 2.5 Pro hits 21.6% on HLE. Claude 3.7 Sonnet hits 7.8% on HLE. How can labs regain the trust of researchers who claim the scoreboard reflects training for specific tests rather than actual intelligence? Gemini 2.5 Pro reaches 54.0% on SimpleQA. This outperforms the 48.6% of high o3.
Enterprise integration and tools
Google integrates Gemini 2.5 Pro into Workspace, Android, and Cloud. Gmail uses the model for summarization. Docs uses it for writing assistance. Sheets uses it for data analysis. On Android, a distilled model provides on-device capabilities for Pixel and Samsung devices. The Gemini API supports video, audio, and text. Developers use it for video analysis, document processing, and code understanding.
The model enables various professional applications. Users upload two-hour lecture videos to receive detailed summaries with timestamps. Users feed hundreds of PDF documents into the model to ask cross-referencing questions. The model analyzes entire code repositories to identify bugs and suggest improvements. Scientific researchers use the model to process research papers, datasets, and experimental results simultaneously. Gemini 2.5 Pro also assists in creative production by generating text, images, and structured data.




