Research
147 posts
How Canada uses Claude
Adoption of Claude is high in Canada. Based on a sample of Claude.ai conversations in February 2026, 2.6% of global traffic is in Canada. Adjusting for popul…
How Claude's values vary by model and language
We analyzed 300,000 real conversations to measure the values Claude expresses across models and languages, compressed into four interpretable axes.
How Claude Performs on Robotics Tasks
Do language models’ strengths transfer to robotics? Can a model perceive a scene, understand a particular robot’s state, and issue actions that reliably effe…
Frontier Red Team
Anthropic's Frontier Red Team stress-tests AI systems to understand the full extent of their current capabilities and anticipate what comes next. We provide …
An off switch for dual use knowledge in AI models
New results on a method of controlling access to potentially dangerous AI capabilities
A global workspace in language models
Interpretability research on Claude's internal thoughts.
Anthropic Economic Index report: Cadences
In the latest Anthropic Economic Index report, we look at when people come to Claude, what they produce with it, and how they perceive AI’s impact on their w…
Project Fetch: Phase two
Results from our latest test of whether Claude can help Anthropic employees perform sophisticated robotics tasks. We found that Claude Opus 4.7, operating wi…
How Claude Code is used in practice
New Anthropic research looking at interactive agentic coding. We evaluate the composition of tasks, human-AI collaboration, and success rates.
Paving the way for AI agents in biology
In this Anthropic Science post, Laura Luebbert argues that we need to make biological data infrastructure more agent-friendly.
Measuring LLMs' impact on N-day exploits
In cybersecurity, a large fraction of real-world harm comes from N-days: vulnerabilities that have already been publicly disclosed, but only patched on some …
Making Claude a chemist
Anthropic is working with world-class synthetic, computational, and analytical chemists to make Claude better at chemistry. In this post, we share our first …
Mapping AI-enabled cyber threats
We’ve spent the past year investigating how threat actors are weaponizing AI to conduct cyber operations. Today, we’re sharing a new analysis that maps these…
Coding agents in the social sciences
Results from a survey of 1,260 social scientists about AI and coding agent use.
Project Glasswing: An initial update
An early update on what we've learned from Project Glasswing.
Measuring LLMs’ ability to develop exploits
We've developed two new, challenging academic benchmarks measuring AI models’ ability to develop exploits, and an updated version of the benchmark measuring …
2028: Two scenarios for global AI leadership
Our views on the AI competition between the US and China.
Teaching Claude why
New research on how we've reduced agentic misalignment
Alignment Research
Can Claude develop, test, and analyze alignment ideas of its own? We ran an experiment to find out.
Natural Language Autoencoders
AI models like Claude talk in words but think in numbers. In this study, we train Claude to translate its thoughts into human-readable text.
Interpretability Research
AI models like Claude talk in words but think in numbers. In this study, we train Claude to translate its thoughts into human-readable text.
Donating our open-source alignment tool
Updating Petri to version 3.0 and donating it to Meridian Labs
Focus areas for The Anthropic Institute
At The Anthropic Institute (TAI), we’ll be using the information we can access from within a frontier lab to investigate AI’s impact on the world, and sharin…
How people ask Claude for personal guidance
We look at what types of guidance people ask of Claude, and describe how this research shaped the training of our newest models.
Evaluating Claude’s bioinformatics research capabilities with BioMysteryBench
In this post, Brianna, a researcher on the discovery team, shares results from a recent bioinformatics benchmarking effort.Almost as soon as large language m…
Announcing the Anthropic Economic Index Survey
The Economic Research team is launching the Anthropic Economic Index Survey, a monthly survey conducted through Anthropic Interviewer.
What 81,000 people told us about the economics of AI
Anthropic's recent survey of 81,000 Claude users provides a way to connect people’s economic concerns with what we’ve quantified in Claude traffic.
Automated Alignment Researchers: Using large language models to scale scalable oversight
Large language models’ ever-accelerating rate of improvement raises two particularly important questions for alignment research.
Trustworthy agents in practice
AI “agents” represent the latest major shift in how people and organizations are using AI. A couple of years ago, AI models were only broadly available as ch…
Assessing Claude Mythos Preview’s cybersecurity capabilities
Claude Mythos Preview is a new general-purpose language model that is strikingly capable at computer security tasks. This post provides technical details for…
Emotion concepts and their function in a large language model
All modern language models sometimes act like they have emotions. What’s behind these behaviors? Our interpretability team investigates.
How Australia Uses Claude: Findings from the Anthropic Economic Index
Anthropic is expanding to Australia. We’re opening a new office in Sydney in the coming weeks, and we’ve signed a Memorandum of Understanding with the Austra…
Anthropic Economic Index report: Learning curves
The Anthropic Economic Index uses our privacy-preserving data analysis system to track how Claude is being used across the economy. It’s part of our effort t…
Economic Research
Anthropic's fifth Economic Index report studies Claude usage in February 2026, building on the economic primitives framework introduced in our previous report.
Introducing our Science Blog
We’re launching a new blog about AI and science. We’ll share work happening at Anthropic and elsewhere, our collaborations with external researchers and labs…
Long-running Claude for scientific computing
In this post, Siddharth Mishra-Sharma, a researcher on the Discovery team, explains how to apply multi-day agentic coding workflows—test oracles, persistent …
Vibe physics: The AI grad student
Can AI do theoretical physics? In this guest post, professor of physics Matthew Schwartz decided to find out by supervising Claude through a real research ca…
Societal Impacts Research
We invited Claude.ai users to share how they use AI, what they dream it could make possible, and what they fear it might do. Nearly 81,000 people participate…
A “diff” tool for AI: Finding behavioral differences in new models
Every time a new AI model is released, its developers run a suite of evaluations to measure its performance and safety. These tests are essential, but they a…
Reverse engineering Claude's CVE-2026-2796 exploit
This post dives deep into how Claude wrote an exploit for one of the vulnerabilities it found in Firefox.
Labor market impacts of AI: A new measure and early evidence
Key findingsWe introduce a new measure of AI displacement risk, observed exposure, that combines theoretical LLM capability and real-world usage data, weight…
An update on our model deprecation commitments for Claude Opus 3
As we develop increasingly capable AI models, it’s currently necessary to deprecate and retire our past models due to the cost and complexity of maintaining …
The persona selection model
A theory of why AI models act like humans.
Anthropic Education Report: The AI Fluency Index
Anthropic's AI Fluency Index measures 11 observable behaviors across thousands of Claude.ai conversations to understand how people develop AI collaboration s…
Measuring AI agent autonomy in practice
AI agents are here, and already they’re being deployed across contexts that vary widely in consequence, from email triage to cyber espionage. Understanding t…
India Country Brief: The Anthropic Economic Index
India, already the world’s largest exporter of IT services, is home to one of the world’s fastest-growing AI user bases. Understanding how AI is being used i…
LLM-discovered 0 days
AI models can now find high-severity vulnerabilities at scale. This is a moment to empower defenders. We're now using Claude to find and help fix vulnerabili…
How AI assistance impacts the formation of coding skills
Research shows AI helps people do parts of their job faster. In an observational study of Claude.ai data, we found AI can speed up some tasks by 80%. But doe…
Disempowerment patterns in real-world AI usage
AI assistants are now embedded in our daily lives—used most often for instrumental tasks like writing code, but increasingly in personal domains: navigating …
The assistant axis
Who is the Assistant? We investigate the character that most modern language models inhabit when interacting with users.
AI models on realistic cyber ranges
In a recent evaluation of AI models’ cyber capabilities, current Claude models can now succeed at multistage attacks on networks with dozens of hosts using o…
The Anthropic Economic Index report: New building blocks for understanding AI use
Is artificial intelligence really making people faster at work? What sort of tasks does AI support best? And how might it change the nature of people’s occup…
Anthropic Economic Index report: Economic primitives
This report introduces new metrics of AI usage to provide a rich portrait of interactions with Claude in November 2025, just prior to the release of Opus 4.5.
Finding bugs with Claude and property-based testing
Ensuring that programs are bug-free is one of the most challenging aspects of software engineering. We developed an agent that can efficiently identify bugs …
Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks
Last year, we described a new approach to defend against jailbreaks, which we called Constitutional Classifiers. We’ve now developed the next generation.
AI to defend critical infrastructure
AI could help defenders of critical infrastructure identify the vulnerabilities that attackers might exploit—and close them before they are exploited. Anthro…
Introducing Bloom: an open source tool for automated behavioral evaluations
We're releasing Bloom, an open source agentic framework for generating behavioral evaluations of frontier AI models. Bloom takes a researcher-specified behav…
Project Vend: Phase two
How Claude turned around its failing vending machine business
Introducing Anthropic Interviewer
What 1,250 professionals told us about working with AI
How AI Is Transforming Work at Anthropic
How AI Is Transforming Work at Anthropic
AI agents find smart contract exploits
We evaluated AI agents' ability to exploit smart contracts using a new benchmark comprising contracts that were actually exploited. On contracts exploited af…
Estimating AI productivity gains
Anthropic economic research on productivity gains
Mitigating the risk of prompt injections in browser use
Claude Opus 4.5 sets a new standard in robustness to prompt injections—adversarial instructions hidden within the content that AI models process. Our new mod…
Natural emergent misalignment from reward hacking
We show for the first time that realistic AI training processes can accidentally produce misaligned models.
Project Fetch: Can Claude train a robot dog?
A practical experiment on AI's ability to affect the physical world
Commitments on model deprecation and preservation
Claude models are increasingly capable: they're shaping the world in meaningful ways, becoming closely integrated into our users’ lives, and showing signs of…
Emergent introspective awareness in large language models
Research from Anthropic on the ability of large language models to introspect
Preparing for AI’s economic impact: exploring policy responses
We’ve asked economists and researchers to explore policy responses to the potential economic effects of powerful AI. We share some of the initial ideas and f…
A small number of samples can poison LLMs of any size
Anthropic research on data-poisoning attacks in large language models
Petri: An open-source auditing tool to accelerate AI safety research
A new automated auditing tool for AI safety research
Building AI for cyber defenders
How we've improved Claude's cyber defense skills.
Anthropic Economic Index report: Uneven geographic and enterprise AI adoption
To study such patterns of early AI adoption, we extend the Anthropic Economic Index along two important dimensions, introducing a geographic analysis of Clau…
Anthropic Economic Index: Tracking AI's role in the US and global economy
New research from Anthropic exploring geographic patterns of AI use
LLMs and biorisk
This article explains why we believe that evaluating biorisk and safeguarding against it is a critical element of responsible AI development.
Developing Nuclear Safeguards for AI
Together with the NNSA and DOE national laboratories, we have co-developed a classifier—an AI system that automatically categorizes content—that distinguishe…
Claude Opus 4 and 4.1 can now end a rare subset of conversations
An update on our exploratory research on model welfare
Claude does cyber competitions
Throughout 2025, we have been quietly entering Claude in cybersecurity competitions designed primarily for humans. In many of these competitions Claude did p…
Persona vectors: Monitoring and controlling character traits in language models
A paper from Anthropic describing persona vectors and their applications to monitoring and controlling model behavior
Cyber evaluations of Claude 4
We partnered with Pattern Labs on a range of cybersecurity evaluations of Claude Opus 4 and Claude Sonnet 4, with Opus demonstrating especially notable impro…
Project Vend: Can Claude run a small shop? (And why does that matter?)
We let Claude run a small shop in the Anthropic office. Here's what happened.
Agentic misalignment: How LLMs could be insider threats
New research on simulated blackmail, industrial espionage, and other misaligned behaviors in LLMs
Confidential Inference via Trusted Virtual Machines
Announcing a new collaborative research paper on Confidential Inference, a set of tools to improve the security of our model weights and of our users' data
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
A new set of evaluations to test the sabotage and monitoring capabilities of LLM AI models
Cyber toolkits for LLMs
Large Language Models (LLMs) that are not fine-tuned for cybersecurity can succeed in multistage attacks on networks with dozens of hosts when equipped with …
Open-sourcing circuit-tracing tools
In our recent interpretability research, we introduced a new method to trace the thoughts of a large language model. Today, we’re open-sourcing the method so…
Anthropic Economic Index: AI's impact on software development
Data on how software developers are using Claude
Exploring model welfare
Announcing a new research program at Anthropic on model welfare
Values in the wild: Discovering and analyzing values in real-world language model interactions
An Anthropic research paper testing which values AI models express in the real world
Reasoning models don't always say what they think
Research from Anthropic on the faithfulness of AI models' Chain-of-Thought
Tracing the thoughts of a large language model
Anthropic's latest interpretability research: a new microscope to understand Claude's internal mechanisms
Auditing language models for hidden objectives
A collaboration between Anthropic's Alignment Science and Interpretability teams
Forecasting rare language model behaviors
Anthropic research on predicting rare, undesirable AI behaviors
Insights on Crosscoder Model Diffing
At the link above, we report some developing work from the Anthropic Interpretability team on Crosscoder Model Diffing, which might be of interest to researc…
Constitutional Classifiers: Defending against universal jailbreaks
A paper from Anthropic describing a new way to guard LLMs against jailbreaking
Claude SWE-Bench Performance
Explore Claude's breakthrough performance on SWE-Bench, demonstrating advanced software engineering capabilities and code generation accuracy. Learn about ou…
Building Effective AI Agents
Discover how Anthropic approaches the development of reliable AI agents. Learn about our research on agent capabilities, safety considerations, and technical…
Alignment faking in large language models
A paper from Anthropic's Alignment Science team on Alignment Faking in AI large language models
Clio: Privacy-preserving insights into real-world AI use
A blog post describing Anthropic’s new system, Clio, for analyzing how people use AI while maintaining their privacy
A statistical approach to model evaluations
Suppose an AI model outperforms another model on a benchmark of interest—testing its general knowledge, for example, or its ability to solve computer-coding …
Evaluating feature steering: A case study in mitigating social biases
A new piece of Anthropic research by Durmus et al.: "Evaluating feature steering: A case study in mitigating social biases"
Sabotage evaluations for frontier models
A new paper on AI safety evaluations from Anthropic's Alignment Science team
Using dictionary learning features as classifiers
At the link above, we report some developing work from the Anthropic interpretability team on developing feature-based classifiers, which might be of interes…
Circuits Updates – September 2024
At the above link, we report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively …
Circuits Updates – August 2024
At the link above, we report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively …
Circuits Updates – July 2024
At the link above, we report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively …
Circuits Updates – June 2024
At the link above, we report a number of developing ideas on the Anthropic Interpretability team, which might be of interest to researchers working actively …
Sycophancy to subterfuge: Investigating reward tampering in language models
Empirical evidence that serious misalignment can emerge from seemingly benign reward misspecification.
The engineering challenges of scaling interpretability
In this post, and in the above roundtable video, our researchers reflect on the close relationship between scientific and engineering progress, and discuss t…
Claude’s Character
Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train mo…
Mapping the mind of a large language model
We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed …
Circuits Updates – April 2024
At the link above, we report a number of developing ideas on the Anthropic Interpretability team, which might be of interest to researchers working actively …
Simple probes can catch sleeper agents
This “Alignment Note” presents some early-stage research from the Anthropic Alignment Science team following up on our recent “Sleeper Agents: Training Decep…
Measuring the Persuasiveness of Language Models
Anthropic developed a way to test how persuasive language models (LMs) are, and analyzed how persuasiveness scales across different versions of Claude.
Many-shot jailbreaking
We investigated a “jailbreaking” technique — a method that can be used to evade the safety guardrails put in place by the developers of large language models…
Reflections on Qualitative Research
This note offers some opinionated thoughts on why interpretability research may have qualitative aspects be more central than we're used to in other fields. …
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternat…
Evaluating and Mitigating Discrimination in Language Model Decisions
As language models (LMs) advance, interest is growing in applying them to high-stakes societal decisions, such as determining financing or housing eligibilit…
Specific versus General Principles for Constitutional AI
Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a sta…
Towards Understanding Sycophancy in Language Models
Reinforcement learning from human feedback (RLHF) is a popular technique for training high-quality AI assistants. However, RLHF may also encourage model resp…
Collective Constitutional AI: Aligning a Language Model with Public Input
Anthropic and the Collective Intelligence Project recently ran a public input process involving ~1,000 Americans to draft a constitution for an AI system. We…
Decomposing Language Models Into Understandable Components
Neural networks are trained on data, not programmed to follow rules. With each step of training, millions or billions of parameters are updated to make the m…
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
In our latest paper, Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, we outline evidence that there are better units of analys…
Challenges in evaluating AI systems
Most conversations around the societal impacts of artificial intelligence (AI) come down to discussing some quality of an AI system, such as its truthfulness…
Tracing Model Outputs to the Training Data
As large language models become more powerful and their risks become clearer, there is increasing value to figuring out what makes them tick. In our previous…
Studying Large Language Model Generalization with Influence Functions
When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source o…
Measuring Faithfulness in Chain-of-Thought Reasoning
Large language models (LLMs) perform better when they produce step-by-step, “Chain-ofThought” (CoT) reasoning before answering a question, but it is unclear …
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
As large language models (LLMs) perform more difficult tasks, it becomes harder to verify the correctness and safety of their behavior. One approach to help …
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Large language models (LLMs) may not equitably represent diverse global perspectives on societal issues. In this paper, we develop a quantitative framework t…
Circuits Updates — May 2023
We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some…
Interpretability Dreams
Our present research aims to create a foundation for mechanistic interpretability research. In particular, we're focused on trying to resolve the challenge o…
Distributed Representations: Composition & Superposition
Distributed representations are a classic idea in both neuroscience and connectionist approaches to AI. We're often asked how our work on superposition relat…
Privileged Bases in the Transformer Residual Stream
Our mathematical theories of the Transformer architecture suggest that individual coordinates in the residual stream should have no special significance (tha…
The Capacity for Moral Self-Correction in Large Language Models
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- t…
Superposition, Memorization, and Double Descent
In a recent paper, we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition, where they represent more featu…
Discovering Language Model Behaviors with Model-Written Evaluations
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evalua…
Constitutional AI: Harmlessness from AI Feedback
As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant…
Measuring Progress on Scalable Oversight for Large Language Models
Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potenti…
Toy Models of Superposition
In this paper, we use toy models — small ReLU networks trained on synthetic data with sparse input features — to investigate how and when models represent mo…
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outpu…
Language Models (Mostly) Know What They Know
We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first sh…
Softmax Linear Units
In this paper, we report an architectural change which appears to substantially increase the fraction of MLP neurons which appear to be "interpretable" (i.e.…
Scaling Laws and Interpretability of Learning from Repeated Data
Recent large language models have been trained on vast datasets, but also often on repeated data, either intentionally for the purpose of upweighting higher …
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We …
In-context Learning and Induction Heads
[Read on anthropic.com](https://www.anthropic.com/research/in-context-learning-and-induction-heads)
Predictability and Surprise in Large Generative Models
Large-scale pre-training has recently emerged as a technique for creating capable, general purpose, generative models such as GPT-3, Megatron-Turing NLG, Gop…
A Mathematical Framework for Transformer Circuits
[Read on anthropic.com](https://www.anthropic.com/research/a-mathematical-framework-for-transformer-circuits)
A General Language Assistant as a Laboratory for Alignment
Given the broad capabilities of large language models, it should be possible to work towards a general-purpose, text-based assistant that is aligned with hum…