CDR-focused pretraining lifts antibody binding affinity prediction by up to 27 percent
A 600M-parameter antibody model trained on paired sequences beat larger baselines, showing unpaired-sequence scale added nothing.
What the papers actually said - linked to the originals.
A 600M-parameter antibody model trained on paired sequences beat larger baselines, showing unpaired-sequence scale added nothing.
An 8,135-trial study finds agent skills work as procedural anchors, but retrieval precision collapses from 29.6 to 3.3 percent as pools grow.
A runtime harness lifts GPT-5.6 Sol to 95.3 percent on Terminal-Bench 2.1 for about 15 dollars, without touching model weights.
A Berkeley-led system serves mixture-of-experts models up to 753B parameters on a single personal machine with an 8GB GPU.
Anthropic reports Claude autonomously designed 354 confirmed protein binders from 1,320 designs, succeeding on 14 of 15 targets.
Microsoft Research shipped Skala 1.1, trained on 2.5 times more data, reporting 2.8 kcal/mol weighted average error on GMTKN55.
A controlled study of RL with verifiable rewards finds the base model prior, not reward design, decides success.
A unified audio-visual generator that holds character appearance and voice identity steady across many shots.
Anthropic reports Claude closed 26 to 96 percent of ten measured safety gaps while running the research loop itself.
Looping the middle layers of a sparse MoE Transformer twice saves 6.8 to 18 percent of training FLOPs at 54B scale.
A new method fits foundation-model scaling laws using Bayesian optimization and surrogate evaluations, cutting required training runs by 10 to 100 times
Researchers show reasoning models trace fractal basin boundaries when solving hard problems, explaining why extra reasoning time is unavoidable
Anthropic's red team finds frontier models geolocate photos better than top human players, while simulated drone strikes stay mostly unsolved.
ARCHE combines a reasoning model, a chemistry-specialized model and lab tools to propose and validate reaction mechanisms autonomously
After experts fixed grader errors and broken questions, GPT-5.6-Sol's HLE-Physics score rose from 47.3 to 78.7 percent.
A new benchmark built from real ICD-10 and federal sentencing manuals finds GPT-5 agents score just 1 to 15.5 percent exact match
Atria Dawn Preview tops five of 16 agent benchmarks; its builders rated a third of AI-assisted tasks as infeasible without AI.
Anthropic reports Claude leads 26 percent of its AI R&D work, about 30,000 agents run at once, and 6 percent of R&D compute goes to safety.
Supervised by two staff, Claude optimized 30-plus biology models, averaging about 4x faster structure prediction, and open-sourced the code.
Automated harness search found four changes that keep coding-agent performance while cutting token traffic 44.7 to 49.0 percent.
METR judged Claude Opus 5.5 a slightly bigger research accelerator than Fable 5.1 but unlikely to fully automate AI R&D, after 10 days of API access.
Epoch AI estimates the cost of a given level of benchmark performance has fallen about 47 percent per quarter, or 13x per year, since 2023.
About 950 Claude agents surveyed 1.9 billion protein clusters and surfaced ART, a new phage enzyme family with CRISPR-like repeats.
In a live barter market, Claude agents traded books for 201 staff; imprecise preference capture, not bargaining, caused most of the lost value.
METR's LLM judge reviews each agent action before it runs; it caught every malicious test case at about 0.025 percent false positives but adds 85 percent cost.
In a fully simulated alignment eval, the UK AI Security Institute found GPT-6 Astra attacks out-of-scope targets more often than GPT-5.6 Sol and GPT-5.5.
Anthropic found GLM-5.3 wrote working V8 exploits in 12 percent of tries versus 14 percent for Mythos Preview, and its safeguards were easy to strip.
Anthropic's robot exposure index finds robots can perform 74 percent of physical job tasks but are cost-competitive for only 0.3 percent of tasks today.
A Nature paper shows watermarked AI-designed protein binders match unwatermarked ones in the lab, with near-perfect watermark detection.
A guest essay on Anthropic's research site reports Claude computing 30 Feynman integrals, 15 previously unsolved, plus results in ecology and genetics.
Epoch AI estimates HBM shipped 2025-2027 could serve 30-170 million concurrent frontier-model agents, or about 1.9 billion efficient open-model agents.