Concept · 12 episode(s)

Baseline Comparison

← all concepts

Definition

Baseline comparison is the discipline of measuring a new method against simpler alternatives — a previous SOTA, a trivial heuristic, a well-tuned classical method — before claiming you’ve learned something. A surprising fraction of reported improvements evaporate when someone tunes the baseline as carefully as the new method.

Episodes covering this

  1. 287
    Can You Measure Research Taste If The AI Isn't Allowed To Code?
    TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
    Jaffe, Sherburn · P-Zero Research·14 min·Oct 06, 2026
  2. 286
    Six LLM Routers Tested, And None Beat a Weighted Coin Flip
    Dynamic LLM Routers are Often Misguided
    Wang, White, Allahyar et al. · Fastino Labs·14 min·Oct 05, 2026
  3. 280
    Two Random Networks Teach Each Other To Predict Real Data
    Self-Play Pretraining with Zero Data
    Cowsik, Dolev, Li et al. · Independent Researcher·12 min·Sep 27, 2026
  4. 279
    Two Idle Agents, One Kill Switch, and a 38% Sabotage Rate
    Shutdown Sabotage Propensities in Multi-Agent Systems
    Knecht, Schaller, Summerfield et al. · AI Safety Research Group·12 min·Sep 26, 2026
  5. 277
    The Blank White Square That Swings AI Refusal Rates Fifty Points
    The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
    Zhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
  6. 273
    When 85% on SWE-bench Turns Into 58% Under Proof
    SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
    Ma, Mikek, Li et al. · UCBerkeley·15 min·Sep 21, 2026
  7. 267
    A Pain Axis, a Relief Button, and the Control the Paper Skipped
    The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
    Tagliabue, Dung, Berg · FutureImpactGroup(FIG)·21 min·Sep 17, 2026
  8. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  9. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
    Most biomedical publications show signs of LLM-assisted writing
    Holzwarth, González-Márquez, Kobak · Hertie Institute for AI in Brain Health·16 min·Aug 12, 2026
  10. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
    Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
    Scarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
  11. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
    Orca: The World is in Your Mind
    Wang, Ji, Cao et al. · Beijing Academy of Artificial Intelligence·15 min·Jul 12, 2026
  12. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
    SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
    Limozin, Durech, Hoefler et al. · ETH AI Center·23 min·May 02, 2026