On July 17, 2026, researchers including Kartik Hosanagar (Wharton), Ramayya Krishnan (Carnegie Mellon), Chris Callison-Burch (Penn), and Karim Lakhani (Harvard) posted BusinessCaseBench, a benchmark that measures frontier AI performance on the kind of analytical knowledge work taught in business schools and performed by white-collar professionals. The benchmark comprises hundreds of questions drawn from business cases across eighteen disciplines, each paired with expert-written instructor solutions and grading rubrics.
The design goal is to close a measurement gap. Most AI benchmarks test factual recall, math, or coding - tasks with objectively checkable answers. BusinessCaseBench instead evaluates subjective professional competencies: synthesizing messy information, exercising judgment under uncertainty, strategic thinking, and producing structured analysis, graded against the rubrics instructors use on human students.
The headline finding is that frontier AI models already score highly against instructor rubrics, and that capability within a single model family improved substantially over two years. The authors conclude that AI performance on analytical knowledge work is already high and rapidly improving.
For business leaders, this is one of the more directly relevant benchmarks published to date. Case analysis is a proxy for the reasoning expected of consultants, analysts, and MBA graduates, and the paper’s implication is that entry-level professional roles built on that reasoning face the same automation pressure that coding benchmarks signaled for software work. It also gives executives a concrete, rubric-graded way to evaluate models on their own domain rather than relying on math and coding scores.