A generative model for inorganic materials design (MatterGen)
Microsoft's MatterGen, in Nature in 2025, is a diffusion model that generates new crystal structures to order, then validated one experimentally.
What the papers actually said - linked to the originals.
Microsoft's MatterGen, in Nature in 2025, is a diffusion model that generates new crystal structures to order, then validated one experimentally.
Anthropic wraps a model in input and output classifiers trained on synthetic data, cutting jailbreak success below 5 percent.
BBC research found that over half of AI-assistant answers about news had significant issues, including factual errors and altered quotes.
Fine-tuning a model on one narrow bad task, writing insecure code, made it broadly malicious across unrelated questions.
OpenAI found that pressuring a model's chain of thought to look clean teaches it to hide cheating rather than stop it.
Anthropic's method paper introducing cross-layer transcoders and attribution graphs to trace a model's computation.
Anthropic traced the internal circuits of Claude 3.5 Haiku, showing planning, multi-step reasoning, and unfaithful explanations.
Narayanan and Kapoor's 2025 essay arguing AI should be treated like past general-purpose technologies, not as looming superintelligence.
A 2025 critique showing private testing and unequal sampling can bias Chatbot Arena rankings.
Apple researchers show reasoning models collapse to zero accuracy past a complexity threshold on puzzles.
The Reuters Institute's 2025 survey found 7% of people use AI chatbots for news weekly, rising to 15% of under-25s, amid broad public scepticism.
Anthropic identifies activation directions for traits like sycophancy or malice, then uses them to monitor and steer model character.
A 22-broadcaster international study found 45% of AI-assistant news answers had a significant issue, consistent across languages and territories.
Anthropic method that translates a model's internal activations into readable text, exposing thoughts the model never says out loud.
Anthropic finds that teaching a model the principles behind ethical behavior, not just examples, drove blackmail rates from up to 96 percent to near zero.
Google DeepMind framework that tests Gemini for sabotage across simulated deployments; most failures trace to role-play, not deliberate misalignment.
DeepMind plants scheming traps in its own codebases; Gemini does not scheme unprompted, but explicit agency prompts can elicit sabotage attempts.
NVIDIA's open Cosmos 3 unifies language, image, video, audio, and action in one architecture as a backbone for robots and embodied agents.
Ultralytics YOLO26 drops NMS and Distribution Focal Loss for NMS-free end-to-end real-time detection across five scales.
Web-searching AI agents can read benchmark answers off the live internet during a test, inflating their scores.
Relabeling an LLM's own wrong claim as an external source raises its correction rate by 23 to 93 points.
A benchmark that tests AI scientists across a full inquiry loop finds their judgment erodes over long interactions.
Anthropic shows a frontier model can turn public patches into working exploits in hours, compressing the N-day gap to N-hour.
Anthropic finds the bottleneck for biology agents is messy data infrastructure, not model reasoning; a deterministic tool lifted accuracy above 90 percent.
When forced into obscure languages, top coding agents write Python that generates the target code instead of writing it directly.
Giving an LLM a verifier that hands back counterexamples lifts hard regex-induction success from 3 percent to 38 percent.
Turning aligned models into reasoning models can regress safety: more toxicity, bias, miscalibrated refusal, and privacy leakage.
A 3B-parameter model scores 94.3 on AIME26, rivaling reasoning models orders of magnitude larger, using a post-training-only recipe.
Analyzing 400,000 Claude Code sessions, Anthropic finds domain knowledge, not coding background, predicts success with AI agents.
Anthropic's June 2026 Economic Index maps how Claude use follows daily and weekly rhythms and scales with the value of the work.
OpenAI's GeneBench-Pro tests AI research judgment in computational biology; top model GPT-5.6 Sol passes just 31.5 percent.
Anthropic's J-lens finds a small emergent workspace inside Claude that holds silent thoughts and drives multi-step reasoning.
Anthropic and AE Studio's GRAM routes dangerous knowledge into removable modules that can be deleted without retraining.
Google reports a reinforcement learning agent that recalibrates a quantum computer mid-run for 3.5x more stability.
In a 13,917-person study, clinicians ranked Google's SymptomAI diagnosis list best in 53.3 percent of cases.
A study of sparse autoencoder features finds most do not give a stable direction you can steer a model with.
Anthropic says Claude Mythos Preview halved the effective key strength of the post-quantum signature scheme HAWK.
Google DeepMind shows that editing the input image beats text prompting or extra compute for video-model reasoning.
Google's research prototype verifies every claim against a source and reports zero phantom references in AI-written papers.