Glossary · Term

ground truth

← all terms

Definition

Plain language

The known-correct answer used as the yardstick for grading whatever the system produced.

As stated in the literature

Reference labels or values treated as authoritative for evaluation or reward computation; in verifier and benchmark design, often embedded as privileged information the system under test cannot see.

Also called: ground-truth

Why it matters: Without a trusted reference answer, there is no reliable way to score whether a system's output is right or to train it toward correctness.

For example, when grading a model's answer to '2+2', the ground truth is the known-correct '4' that its output is compared against.

Heard on the show

“Separating starvation from deception took the element registry and the ground-truth anchor the threat model denies the attacker.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 64 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 242
    Making a Vision Model Better by Showing It Blurry Images
  3. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  4. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  5. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  6. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  7. 223
    When Grok Graded Its Own Encyclopedia And Marked Itself Down
  8. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  9. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  10. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
  11. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  12. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  13. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  14. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  15. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  16. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  17. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  18. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  19. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  20. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  21. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  22. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  23. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  24. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  25. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  26. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  27. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  28. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  29. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  30. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  31. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  32. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  33. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  34. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  35. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  36. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  37. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  38. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  39. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  40. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  41. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  42. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  43. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  44. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  45. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  46. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  47. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  48. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  49. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  50. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  51. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  52. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  53. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  54. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  55. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  56. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  57. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  58. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  59. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  60. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  61. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  62. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  63. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  64. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms