Definition
Plain language
A benchmark of real spreadsheet-editing tasks used to test AI agents.
As stated in the literature
An evaluation suite of Excel manipulation tasks, used to test streaming skill-library construction and refinement for agents like Excel Copilot.
Why it matters: It tests agents on realistic spreadsheet work, providing a way to check whether they can build and refine reusable skills for everyday office tasks.
For example, a task might ask an agent to clean up a messy sales sheet and add a column that totals each region.
Heard on the show
“SpreadsheetBench has a runtime that checks cell values.”Episode 078 — Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training