Glossary · Term

GPT-4

← all terms

Definition

Plain language

OpenAI's family of large language models, including the 4 and 5 series.

As stated in the literature

OpenAI's foundation model series including GPT-4, GPT-4o, GPT-4 Turbo, and the GPT-5 family, including reasoning variants.

Also called: GPT-4o, GPT-4 Turbo, GPT-4o-mini, GPT-4.1, GPT-5, GPT-5-mini, GPT-5-nano, GPT-5.1, GPT-5.2, GPT-5.2R, GPT-5.4, GPT-5.5

Why it matters: It set the expectations and pricing for the current generation of frontier chatbots, and successor models are still measured against it.

For example, GPT-4 was the model that popularized using ChatGPT for serious coding and writing tasks.

Heard on the show

“Thirty thousand images, fully black-box, six different models doing the answering, including Claude Sonnet, GPT-5.4, and Llama 4 Maverick.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 75 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  3. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  4. 242
    Making a Vision Model Better by Showing It Blurry Images
  5. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  6. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  7. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  8. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  9. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  10. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  11. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
  12. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
  13. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  14. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  15. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
  16. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  17. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  18. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  19. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
  20. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  21. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  22. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  23. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  24. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  25. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  26. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  27. 166
    A Router That Beats the Frontier Models It Calls
  28. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  29. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  30. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  31. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  32. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  33. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  34. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  35. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  36. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  37. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  38. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  39. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  40. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  41. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  42. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  43. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  44. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  45. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  46. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  47. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  48. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  49. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  50. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  51. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  52. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  53. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  54. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  55. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  56. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  57. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  58. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  59. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  60. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  61. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  62. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  63. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  64. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  65. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  66. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  67. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  68. 020
    The Compliance Gap: Why AI Says Yes and Does No
  69. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  70. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  71. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  72. 014
    Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1
  73. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  74. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  75. 004
    The Sycophancy Circuit That Survives Alignment Training