Google Gemini 3.1 Pro

by Google DeepMind

Gemini 3.1 Pro is Google DeepMind's flagship reasoning model released February 19, 2026, as a mid-cycle update to the Gemini 3 series. It achieves 77.1% on ARC-AGI-2 (more than double its predecessor), supports a 1-million-token context window, and introduces three-tier 'Thinking' levels for balancing speed and computational depth. It uses a Transformer-based Mixture-of-Experts architecture.

Large Language Models
Proprietary

Primary Use Cases

  • Complex multi-step reasoning and problem-solving
  • Long-context document analysis and retrieval
  • Multimodal synthesis (video, audio, large documents)
  • Autonomous web research and agentic workflows
  • Software engineering and code generation
  • SVG and 3D interactive content creation

✅ Strengths

  • 77.1% on ARC-AGI-2 (2x improvement over Gemini 3 Pro)
  • 94.3% on GPQA Diamond scientific knowledge benchmark
  • 80.6% on SWE-Bench Verified software engineering benchmark
  • 85.9% on BrowseComp autonomous web research
  • Cost-effective vs GPT-5.5 (2-2.5x cheaper per token)
  • Superior native multimodal processing

⚠️ Limitations

  • Slower than GPT-5.5 on structured instruction-following tasks
  • 64K output token cap vs GPT-5.5's 128K
  • Paid-only since April 1, 2026 (no free tier for Pro models)
  • Thinking tokens billed at standard output rate, increasing costs for complex tasks
  • Tier 2 rate limits require $250 cumulative spend

🏆 Performance Benchmarks

Standardized benchmark scores for objective comparison

ARC-AGI-2(Reasoning)
77.1%

More than double the performance of Gemini 3 Pro

GPQA Diamond(Scientific Knowledge)
94.3%

Expert-level scientific knowledge evaluation

SWE-Bench Verified(Software Engineering)
80.6%

Real-world software engineering task resolution

BrowseComp(Agentic Web Research)
85.9%

Autonomous web research benchmark

Technical Specifications

Developer
Google DeepMind
Category
Large Language Models
Type
🧠 Foundation Model
License
UNKNOWN
Version
3.1 Pro

Performance Benchmarks

ARC-AGI-277.1%
GPQA Diamond94.3%
SWE-Bench Verified80.6%
BrowseComp85.9%

Source: View benchmark details

Trust & Privacy

Health Status
🟢 Active
Trains on User Data
Does not train
Certifications
None listed

Ecosystem & Integrations

API Access
✅ Available
Chrome Extension
❌ No
Mobile App
❌ No
Pricing Model
💰 Pay Per Use

Pricing Breakdown

Standard Context

Prompts up to 200K tokens

Input$2/M tokens
Output$12/M tokens
Extended Context

Prompts exceeding 200K tokens

Input$4/M tokens
Output$18/M tokens