On July 24, 2026, Anthropic’s Frontier Red Team published Project Pilot, an evaluation built with Andon Labs that asks whether current AI models can autonomously control a commercial drone. The test task is deliberately surveillance-shaped: locate and follow a designated person in an indoor office environment. It decomposes into five subtasks - reconstruct (build a 3D model of the space), localize (track position), navigate (pathfind), detect (identify the target), and follow (keep tracking them). Andon Labs ran the evaluations independently, without Anthropic having access to the benchmark.
Fifteen models from three developers were tested, spanning several model generations. The headline finding is how close the frontier already is: by mid-2026, four of the five subtasks approached the human baseline, with 3D reconstruction the remaining bottleneck at roughly 47 percent of baseline. The best performer, Claude Fable 5, pushed past baseline on every task except reconstruction - it could detect and follow a person but failed at autonomous multi-room navigation because its reconstruction errors compounded. The report also notes a consistency gap: models’ average runs trail their best runs by about six months of progress toward baseline.
Anthropic chose aerial drones because they are cheap, ubiquitous dual-use technology - the same platforms that raise crop yields in agriculture are used to target opposing forces in warfare. The stated goal is situational awareness: measuring how close AI is to autonomously piloting physical robots, with all the benefits and risks that implies.
For business and policy readers, the result reframes embodied AI risk from hypothetical to measurable. General-purpose language models, given drone controls, are already near human baseline on most components of an aerial surveillance task, and the remaining bottleneck is a perception problem that the field is actively working on. Capability tracking of this kind is likely to become a standard input to both regulation and corporate risk assessments.