Glossary · Term

Gemini

← all terms

Definition

Plain language

Google's family of large language models.

As stated in the literature

Google DeepMind's series of foundation models including Gemini 2.5, 3, and 3.1 Pro/Flash variants.

Also called: Gemini 2.5, Gemini 2.5 Pro, Gemini 3, Gemini 3.1, Gemini 3.1 Pro, Gemini 3.1 Flash, Gemini three Flash, Gemini three Pro, Gemini three-point-one, Gemini-3.1-flash-lite

Why it matters: It's one of the three or four model families that define the current capability frontier, and many academic comparisons include it as a baseline.

For example, Gemini 3 Pro powers Google's AI Studio chat experience and is one of the default frontier models in coding-agent benchmarks.

Heard on the show

“On Gemini 3.1 Pro, the protected number never shows up.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 58 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  3. 242
    Making a Vision Model Better by Showing It Blurry Images
  4. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  5. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  6. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  7. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  8. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
  9. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  10. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  11. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  12. 198
    The Model That Knows the Answer and Can't Say It
  13. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  14. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  15. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  16. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  17. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  18. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  19. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  20. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  21. 166
    A Router That Beats the Frontier Models It Calls
  22. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  23. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  24. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  25. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  26. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  27. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  28. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  29. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  30. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  31. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  32. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  33. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  34. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  35. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  36. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  37. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  38. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  39. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  40. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  41. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  42. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  43. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  44. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  45. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  46. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  47. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  48. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  49. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  50. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  51. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  52. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  53. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  54. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  55. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  56. 004
    The Sycophancy Circuit That Survives Alignment Training
  57. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
  58. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related terms