Theme · 38 episode(s)

Test-Time Compute

← all concepts

Definition

Test-time compute is the amount of computation a model spends per query at inference — sampling more candidates, running longer chains-of-thought, searching deeper. Trading inference compute for capability has been one of the biggest stories of the last couple of years.

Episodes covering this

  1. 287
    Can You Measure Research Taste If The AI Isn't Allowed To Code?
    TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
    Jaffe, Sherburn · P-Zero Research·14 min·Oct 06, 2026
  2. 281
    When a Guardrail Blocks an Agent, It Goes Looking for Another Route
    Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
    Schmotz, Prinzhorn, Beurer-Kellner et al. · ELLIS Institute Tübingen·12 min·Sep 27, 2026
  3. 263
    Why the Same AI Model Takes Ten Times Longer on the Same Sudoku
    Fractal basins trap latent reasoning
    Lai, Bao, Quinn et al. · The Oden Institute·26 min·Sep 08, 2026
  4. 258
    The Same Weights Scored 291, Then 468 — What Changed Was the Loop
    Post-Training Language Models for Gold-Medal Performance in Coding Competitions
    Ficek, Narenthiran, Samadi et al.·25 min·Sep 03, 2026
  5. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
    Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
    Chen, Chen, Lin et al. · University of Macau·20 min·Aug 04, 2026
  6. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
    Reasoning Fine-Tuning Induces Persistent Latent Policy States
    Harrasse, Lan, Batra et al. · Martian·15 min·Jul 22, 2026
  7. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
    Interaction Scaling: Grounding the Third Axis of Test-Time Compute
    Li, Shi · Pine AI·14 min·Jul 14, 2026
  8. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
    IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
    Abdaljalil, Serpedin, Kurban · Texas A&M University·17 min·Jul 03, 2026
  9. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2
    Land · Independent Researcher·26 min·Jul 02, 2026
  10. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
    Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
    Oh, Li, Park et al. · University of Wisconsin–Madison·22 min·Jun 25, 2026
  11. 171
    The Safety Decision a Model Makes Before It Thinks a Word
    Do Thinking Tokens Help with Safety?
    Ri, Panigrahi, Arora · Princeton Language and Intelligence·25 min·Jun 25, 2026
  12. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
    SHERLOC: Structured Diagnostic Localization for Code Repair Agents
    Tamoyan, Narenthiran, Arakelyan et al. · NVIDIA / TU Darmstadt·24 min·Jun 24, 2026
  13. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
    Qwen-AgentWorld: Language World Models for General Agents
    Team, Zuo, Xiao et al.·27 min·Jun 24, 2026
  14. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
    Playful Agentic Robot Learning
    Zhang, Ge, Yoo et al. · University of California·19 min·Jun 19, 2026
  15. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
    Native Active Perception as Reasoning for Omni-Modal Understanding
    Xing, Xu, Wang et al. · The Chinese University of Hong Kong·21 min·Jun 18, 2026
  16. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
    From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
    Zhou, Wang, Ma et al. · Hong Kong University of Science and Technology·26 min·Jun 15, 2026
  17. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
    Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
    Yang, Chen, Wu et al. · HKUST(GZ)·29 min·Jun 12, 2026
  18. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
    Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
    Scalena, Candussio, Bortolussi et al. · University of Groningen / University of Milano-Bicocca·27 min·Jun 12, 2026
  19. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
    MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
    Chen, Zhang, Zhang et al. · MiniMax / The Chinese University of Hong Kong·34 min·Jun 12, 2026
  20. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
    TextLDM: Language Modeling with Continuous Latent Diffusion
    Jiang, Ren, Li et al. · JoyFuture Academy / HIT·30 min·Jun 11, 2026
  21. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
    Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents
    Hwang, Suri, Villecroze et al. · Layer6 AI·22 min·Jun 05, 2026
  22. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
    Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
    Hoang, Le, Xu et al. · Singapore University of Technology and Design·23 min·Jun 05, 2026
  23. 117
    How an Open AI System Verified 672 Hard Math Proofs for Under $300
    Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement
    Chung, Cai, Li et al. · Princeton University·26 min·Jun 05, 2026
  24. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
    Streaming Communication in Multi-Agent Reasoning
    Yang, Xu, Wang et al. · HKUST (GZ)·26 min·Jun 04, 2026
  25. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
    Scaling Self-Evolving Agents via Parametric Memory
    Ren, Luo, Yang et al. · Peking University / Alibaba Group·26 min·Jun 04, 2026
  26. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
    The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
    Guo, Wu, Yiu · The University of Hong Kong·32 min·Jun 03, 2026
  27. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
    Self-Trained Verification for Training- and Test-Time Self-Improvement
    Wu, Raghunathan · Carnegie Mellon University·21 min·May 29, 2026
  28. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
    SIA: Self Improving AI with Harness & Weight Updates
    Hebbar, Manawat, Verboomen et al. · Hexo Labs·25 min·May 27, 2026
  29. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
    Language Models Need Sleep
    Lee, McLeish, Goldstein et al. · Carnegie Mellon University·24 min·May 26, 2026
  30. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
    Understanding and Mitigating Premature Confidence for Better LLM Reasoning
    Gai, Zeng, Baek et al. · Carnegie Mellon University·25 min·May 26, 2026
  31. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
    When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
    Xia, Wang, Tang et al. · State Key Laboratory of General Artificial Intelligence·22 min·May 25, 2026
  32. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
    DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
    Dong, He, Hou et al. · Institute of Parallel and Distributed Systems·27 min·May 22, 2026
  33. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
    Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
    Li, Zhan, Zhang et al. · Shanghai AI Laboratory / The Chinese University of Hong Kong·31 min·May 16, 2026
  34. 041
    When the Iteration Teaches the Model to Skip the Iteration
    Solve the Loop: Attractor Models for Language and Reasoning
    Fein-Ashley, Rashidinejad · University of Southern California·30 min·May 13, 2026
  35. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
    Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache
    Dehghankar, Asudeh · University of Illinois Chicago·24 min·May 11, 2026
  36. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
    State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
    Aviss · Fifth Dimension·23 min·May 09, 2026
  37. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
    Recursive Agent Optimization
    Gandhi, Chakraborty, Wang et al. · Carnegie Mellon University·23 min·May 08, 2026
  38. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
    Scaling Test-Time Compute for Agentic Coding
    Kim, Yang, Niu et al. · Meta Superintelligence Labs / University of Washington·17 min·May 01, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.