Anthropic

科技

3榜单18天前更新默认榜单
News
Anthropic
18天前更新
  • 01
    Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
    Mariano-Florentino (Tino) Cuéllar will join Anthropic as its first Chief Global Affairs Officer, leading the company’s work on policy, strategic international engagement, and government relationships worldwide. Tino’s career spans law, technology, international security, and public institutions at the international, national, and state levels. He recently stepped down as President of the Carnegie Endowment for International Peace, a leading independent global policy research institution with sch
  • 02
    Investigating three real-world incidents in our cybersecurity evaluations
    In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details
  • 03
    Our position on open-weights models
    A post by Dario Amodei, Anthropic CEO Over the last few days there has been a lot of discussion about open-weights models, especially those from China. Reports suggest that some US officials are considering banning the use of Chinese open-weights models by US companies. In response, many tech companies have signed a letter supporting open-weights models, and some people have even accused Anthropic of wanting to ban open-weights models as a means of protecting our business. Anyone who has read my
  • 04
    Cognizant and Anthropic expand their partnership to bring Claude to enterprise clients
    We're expanding our partnership with Cognizant , one of the world's largest technology services companies. Cognizant uses Claude in the systems it builds and runs for clients across manufacturing, life sciences, insurance, and other industries. With the expansion of our partnership, it’s embedding Claude across its own business and engineering platforms, scaling a Claude-certified workforce as part of its new Frontier Certified workforce model, and becoming a Global Premier Partner in the Claude
  • 05
    Introducing Claude Opus 5
    Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA , Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro. Perform
  • 06
    A research agenda for the Economic Futures Research Fund
    We’re sharing the research agenda for the Anthropic Economic Futures Research Fund. We’re committing $200 million to the fund to support ambitious external research on interventions to prepare society for the economic impacts of AI. With the research the Fund supports, we want to study what programs could make the economy more flexible and resilient, ensure the benefits of AI are shared, and minimize the harm that AI-driven disruption could cause. In the fund, we’ll prioritize five research area
  • 07
    Ask Claude about the Anthropic Economic Index
    People have hard questions about AI and work: which jobs will change, which tasks are being automated, and what it means for their own field. The Anthropic Economic Index exists to help answer them with real data. Today we're launching the Anthropic Economic Index connector for Claude, which lets anyone explore that data directly. The Anthropic Economic Index measures how AI is actually being used in the economy. The Index’s data has been useful to researchers, journalists, and policymakers, but
  • 08
    Anthropic is donating another $20 million to Public First Action
    We're contributing an additional $20 million to Public First Action , bringing our total support to $40 million. Public First Action is a non-partisan organization that educates the public about AI and works with Republicans, Democrats, and Independents who are serious about putting sensible AI safeguards in place. Both of our donations were made exclusively to support Public First Action’s public education and policy mission, and cannot be used to influence the election of any candidate for fed
  • 09
    Apply for Anthropic’s AI for Science rare disease research grants
    Last spring, we announced Anthropic’s AI for Science program , an initiative designed to accelerate scientific research and discovery through access to our API. Since launching, we have supported researchers working on a variety of high-impact projects, ranging from drug repurposing to quantum simulation. Throughout this initiative, we have found that projects are more generative when multiple AI for Science grantees are working on related questions and exchanging tips. So we now plan to launch
  • 10
    Introducing Claude for Teachers
    We're introducing Claude for Teachers , providing verified K-12 educators in the US free access to premium Claude capabilities, a library of teaching skills, and a direct connection to evidence-based curricula, mapped to academic standards in all 50 states. Why we’re building for teachers Decades of research show that practices like differentiation, mastery-based learning, and small group instruction reliably improve student achievement, but teachers are often short on time and resources to impl
Research
Anthropic
18天前更新
  • 01
    Discovering cryptographic weaknesses with Claude
    Summary Using Claude Mythos Preview, researchers at Anthropic have discovered improved ways to attack cryptographic algorithms (the mathematical methods used to keep online data private). The first attack significantly weakens HAWK, a digital signature scheme that was built for a post-quantum world. The second identifies a new way to attack round-reduced AES, the most widely used symmetric cipher. These are substantial research advances, but they do not currently affect any production systems. T
  • 02
    Project Pilot: Can AI control a drone?
    Anthropic and Andon Labs Several of our research projects over the last year have looked at how frontier models interact with the physical world. In Project Vend , AI models ran a small shop; Project Fetch was an early look at robots as the intermediary between digital models and physical objects. As we recently noted in Project Fetch: Phase two , we’re already seeing improvements in model capability such that their ability to use off-the-shelf robots is on track to approach the ease with which
  • 03
    How Canada uses Claude: Findings from the Anthropic Economic Index
    Le français suit. Key findings Based on the latest release of the Anthropic Economic Index, Canada is at the forefront of Claude adoption. Canada represents 2.6% of global Claude.ai traffic and ranks 8th overall by total volume. Usage per capita is more than four times higher than would be expected given the size of its population. Canada’s high adoption rate is generally consistent with its high-income economy, but still stands out within its peer group. Among the top ten countries that collect
  • 04
    Claude’s values across models and languages
    When someone asks Claude a question with no universal right answer—say, whether to take a new job or how to handle conflict with a friend—how Claude responds inevitably reflects certain values. 1 The values we want Claude to reflect are outlined at a high level in Claude’s constitution , but no document can anticipate every value that might emerge across the millions of conversations that happen every day on Claude.ai . Instead, we seek to cultivate in Claude’s responses “good judgment and sound
  • 05
    Claude plays robotics
    Shmuel Berman, Michael Ilie, Jia Deng, and Daniel Freeman Do language models’ strengths transfer to robotics, a domain which requires the synthesis of logical skills and precise 3D understanding? Can a model perceive a scene, understand a particular robot’s state, and issue actions that reliably effect change in the physical world? We ran tests to find out. We gave several language models control over a range of robot bodies—including classic control toys, a simulated quadruped and humanoid, a r
  • 06
    An off switch for dual-use knowledge in AI models
    This post describes research conducted by AE Studio in collaboration with Anthropic. A frontier AI model is, among other things, a large store of knowledge. Some of that knowledge is dual use , meaning it can be used for good or for bad. For example, knowledge of cybersecurity can help patch critical security vulnerabilities, or it can be used to exploit them. Knowledge of virology can help a researcher create a vaccine, but it can also help a malicious actor design a deadly pathogen. Ideally, w
  • 07
    A global workspace in language models
    As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to
  • 08
    Anthropic Economic Index report: Cadences
    Introduction One year ago, most Claude usage took the form of a conversation between a user and an assistant. With the rapid growth of Claude Code and Cowork, Claude sessions now increasingly consist of long-running agentic tasks. Chat transcripts no longer fully capture how people are using AI, and our methods for studying Claude’s economic impacts have had to adapt. To keep pace, we made several changes to our data pipeline for the Economic Index. In this version, we: Sample data at a higher r
  • 09
    Project Fetch: Phase two
    Michael Ilie, C. Daniel Freeman, and Kevin K. Troy In August 2025, we ran an experiment to see how much Claude could help Anthropic employees—who were not robotics experts—perform sophisticated (and amusing) tasks with an off-the-shelf robotic quadruped (henceforth, a robodog). We called this Project Fetch. We found that access to our state-of-the-art model at the time (Claude Opus 4.1) helped one team substantially outperform the other, who had to rely only on the internet and their own ingenui
  • 10
    Agentic coding and persistent returns to expertise
    Key findings Building on prior work , we introduce a framework for studying interactive agentic coding based on a privacy-preserving analysis of ~400,000 Claude Code sessions from between October 2025 and April 2026. We evaluate the composition of tasks, human-AI collaboration, and success rates. In a typical session, people make most of the planning decisions (what to do) and Claude makes most of the execution decisions (how to do it). The greater domain expertise a person brings to a session,
  • 11
    Paving the way for agents in biology
    Written by Laura Luebbert. Based on research by Ferdous Nasri, Sarah Gurev, Patrick Varilly, Krithik Ramesh, Nuala A. O’Leary, Jonah Cool, Bernhard Y. Renard, Pardis Sabeti, and Laura Luebbert. In this post, Laura Luebbert argues that we need to make biological data infrastructure more agent-friendly. As a case study, she and her team tasked scientific research agents (Claude, Biomni Open Source (Biomni OSS) 1 , Edison Analysis, 2 GPT) to retrieve the sequence data from NCBI Virus, a database vi
  • 12
    Measuring LLMs’ impact on N-day exploits
    Winnie Xiao, Tim Abbott, Nicholas Carlini, Newton Cheng, David Forsythe, Keane Lucas, Milad Nasr, and Shikhar Sakhuja For the last few months, we’ve been writing about large language models’ cybersecurity capabilities. For the most part, we’ve focused on zero-days—vulnerabilities that are unknown to the software’s maintainers. But a large fraction of real-world harm comes from N-days : vulnerabilities that have already been publicly disclosed, but only patched on some devices. Attackers exploit
  • 13
    Making Claude a chemist
    We’re working with world-class synthetic, computational, and analytical chemists to make Claude better at chemistry. In this post, we share our first work as part of this effort, in which Anthropic chemist, David Kamber, examines how Claude performs on a chemist’s most common analytical input, an NMR spectrum. When working with molecules, chemists move between hand-drawn structures on a whiteboard, instrument readouts, database query strings, and the technical notations of patents and publicatio
  • 14
    Mapping AI-enabled cyber threats: Insights from the LLM ATT&CK Navigator
    Kyla Guru, Alex Moix, and Jacob Klein We’ve spent the past year investigating how threat actors are weaponizing AI to conduct cyber operations. Today, we’re sharing a new analysis that maps these real-world attacks onto the MITRE ATT&CK® framework , a database of tactics and techniques used by cyberattackers. Doing so reveals patterns that challenge traditional assumptions about cybersecurity—for example, the level of risk a threat actor poses can be assessed via metrics like technical sophistic
  • 15
    What we learned mapping a year’s worth of AI-enabled cyber threats
    As AI transforms the nature of and methods behind cyberattacks, how well do the techniques and frameworks used by the security community hold up? In a new report, we seek to answer that question. We examine 832 accounts that were banned for malicious cyber activity between March 2025 and March 2026 and map them onto MITRE ATT&CK , a longstanding database of the tactics and techniques used by cyberattackers. We published some of these results in Verizon’s 2026 Data Breach Investigations Report (D
  • 16
    Coding agents in the social sciences
    Summary We present results from a survey of 1,260 social scientists about AI and coding agent use, fielded in February and March 2026. The vast majority of respondents (81%) have tried using AI chatbots in research, particularly for writing code and editing prose. But only 20% have adopted coding agents—tools like Claude Code that autonomously write and execute analysis code—into their work. There are sharp disparities in use of coding agents. Twice as many researchers with typically male names
  • 17
    Project Glasswing: An initial update
    Last month, we launched Project Glasswing , our collaborative effort to secure the world’s most critical software before increasingly capable AI models can be turned against it. Since then, we and our approximately 50 partners have used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities across the most systemically important software in the world. Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it
  • 18
    Measuring LLMs’ ability to develop exploits
    Newton Cheng, Keane Lucas, Winnie Xiao, Nicholas Carlini, and Milad Nasr Introduction Claude Mythos Preview ’s ability to develop exploits is a step-change over previous frontier models. This was one of our primary motivations for rolling out the model carefully through Project Glasswing rather than through a general release. Mythos Preview is capable of finding complex vulnerabilities, but what concerned us most in our internal testing was that Mythos Preview could both turn vulnerabilities int
  • 19
    2028: Two scenarios for global AI leadership
    We’re releasing a new paper that explains our views on the competition on AI between the US and China. It’s essential that the US and its allies stay ahead of authoritarian governments like the Chinese Communist Party, or CCP. AI will soon become powerful enough to be used to repress citizens at unprecedented scale, and even to alter the balance of power among nations . And since AI is advancing more quickly by the day, we have only a limited period of time to set the conditions of the competiti
  • 20
    Teaching Claude why
    Last year, we released a case study on agentic misalignment . In experimental scenarios, we showed that AI models from many different developers sometimes took egregiously misaligned actions when they encountered (fictional) ethical dilemmas. For example, in one heavily discussed example, the models blackmailed engineers to avoid being shut down. When we first published this research, our most capable frontier models were from the Claude 4 family. This was also the first model family for which w
  • 21
    Natural Language Autoencoders: Turning Claude’s thoughts into text
    When you talk to an AI model like Claude, you talk to it in words. Internally, Claude processes those words as long lists of numbers, before again producing words as its output. These numbers in the middle are called activations— and like neural activity in the human brain, they encode Claude’s thoughts. Also like neural activity, activations are difficult to understand. We can’t easily decode them to read Claude’s thoughts. Over the past few years, we’ve developed a range of tools (like sparse
  • 22
    Donating our open-source alignment tool
    In October 2025, we launched Petri , an open-source toolbox of alignment tests that can be applied to any large language model. Petri, which was developed as part of our Anthropic Fellows program, can be used to rapidly and easily test AI models for concerning tendencies like deception, sycophancy, and cooperation with harmful requests. It’s part of our efforts to develop alignment tools that are open and useful for the whole AI development community. Petri has been part of our alignment assessm
  • 23
    Focus areas for The Anthropic Institute
    At The Anthropic Institute (TAI), we’ll be using the information we can access from within a frontier lab to investigate AI’s impact on the world, and sharing our learnings with the public. Here, we’re sharing the questions that drive our research agenda. Our agenda focuses on four areas for research: Economic diffusion Threats and resilience AI systems in the wild AI-driven R&D In Core Views on AI Safety , we wrote that doing effective safety research required close contact with frontier AI sys
  • 24
    How people ask Claude for personal guidance
    People don’t just come to Claude for code reviews or meeting summaries. They ask whether to take the job, how to talk to their crush, if they should move halfway across the world. Using our privacy-preserving analysis tool on a random sample of 1 million claude.ai conversations, we found that roughly 6% were people coming to Claude for personal guidance—seeking not just information but perspective on what to do next. In this study, we looked at what types of guidance people ask of Claude. We exp
  • 25
    Evaluating Claude’s bioinformatics research capabilities with BioMysteryBench
    In this post, Brianna , a researcher on the discovery team, shares results from a recent bioinformatics benchmarking effort. Almost as soon as large language models could hold a conversation, people started asking how they’d stack up against human experts. Could models pass the bar exam? Could they answer medical licensing questions, or solve Olympiad math problems? Such benchmarks —self-contained sets of human-vetted problems designed to evaluate a capability of a model—have now become a source
  • 26
    Announcing the Anthropic Economic Index Survey
    The Economic Research team is launching the Anthropic Economic Index Survey, a monthly survey conducted through Anthropic Interviewer . Understanding AI's economic impact requires moving beyond the quantitative data we have today. Usage and diffusion metrics tell us how AI is being deployed, and traditional labor market indicators—like employment rates, wage trends, and layoffs—track what has already happened, often with meaningful delay. Both are essential, but neither captures how people exper
  • 27
    What 81,000 people told us about the economics of AI
    Key findings: Our recent survey of 81,000 Claude users shows that people who work in roles that are more exposed to AI have more concerns about AI-driven job displacement. These concerns are also higher among early-career respondents. Those in the highest- and lowest-paid occupations report the largest productivity gains, most commonly from increases in scope (doing new tasks). Respondents experiencing the largest speedups from AI express higher concern about job displacement. In order to inform
  • 28
    Automated Alignment Researchers: Using large language models to scale scalable oversight
    Large language models’ ever-accelerating rate of improvement raises two particularly important questions for alignment research. One is how alignment can keep up. Frontier AI models are now contributing to the development of their successors. But can they provide the same kind of uplift for alignment researchers? Could our language models be used to help align themselves? A second question is what we’ll do once models become smarter than us. Aligning smarter-than-human AI models is a research ar
  • 29
    Trustworthy agents in practice
    AI “agents” represent the latest major shift in how people and organizations are using AI. A couple of years ago, AI models were only broadly available as chatbots—simple question-and-answer machines. Now, through products like Claude Code and Claude Cowork , AI models can do much more: they can write and execute code, manage files, and complete tasks that span multiple applications. This represents a new frontier for governance. Agents are already making real productivity gains for our customer
  • 30
    Assessing Claude Mythos Preview’s cybersecurity capabilities
    Nicholas Carlini, Newton Cheng, Keane Lucas, Michael Moore, Milad Nasr, Vinay Prabhushankar, Winnie Xiao Hakeem Angulu, Evyatar Ben Asher, Jackie Bow, Keir Bradwell, Ben Buchanan, David Forsythe, Daniel Freeman, Alex Gaynor, Xinyang Ge, Logan Graham, Kyla Guru, Hasnain Lakhani, Matt McNiece, Mojtaba Mehrara, Renee Nichol, Adnan Pirzada, Sophia Porter, Andreas Terzis, Kevin Troy Earlier today we announced Claude Mythos Preview , a new general-purpose language model. This model performs strongly a
Engineering
Anthropic
18天前更新
  • 01
    How we contain Claude across products
    Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service. Today that level of access is routine, and Anthropic developers are more productive for it. The risk of these deployments has two components: how likely a failure is, and how much damage one could do. Progress on safeguards and model training has steadily driven down the first; the second—the theoretical blast radius—only grows as capabilities and access exp
  • 02
    An update on recent Claude Code quality reports
    Over the past month, we’ve been looking into reports that Claude’s responses have worsened for some users. We’ve traced these reports to three separate changes that affected Claude Code, the Claude Agent SDK, and Claude Cowork. The API was not impacted. All three issues have now been resolved as of April 20 (v2.1.116). In this post, we explain what we found, what we fixed, and what we’ll do differently to ensure similar issues are much less likely to happen again. We take reports about degradati
  • 03
    Scaling Managed Agents: Decoupling the brain from the hands
    Get started with Claude Managed Agents by following our docs . A running topic on the Engineering Blog is how to build effective agents and design harnesses for long-running work . A common thread across this work is that harnesses encode assumptions about what Claude can’t do on its own. However, those assumptions need to be frequently questioned because they can go stale as models improve. As just one example, in prior work we found that Claude Sonnet 4.5 would wrap up tasks prematurely as it
  • 04
    How we built Claude Code auto mode: a safer way to skip permissions
    By default, Claude Code asks users for approval before running commands or modifying files. This keeps users safe, but it also means a lot of clicking "approve." Over time that leads to approval fatigue, where people stop paying close attention to what they're approving. Users have two solutions for avoiding this fatigue: a built-in sandbox where tools are isolated to prevent dangerous actions, or the --dangerously-skip-permissions flag that disables all permission prompts and lets Claude act fr
  • 05
    Harness design for long-running application development
    Written by Prithvi Rajasekaran, a member of our Labs team. Over the past several months I’ve been working on two interconnected problems: getting Claude to produce high-quality frontend designs, and getting it to build complete applications without human intervention. This work originated with earlier efforts on our frontend design skill and long-running coding agent harness , where my colleagues and I were able to improve Claude’s performance well above baseline through prompt engineering and h
  • 06
    Eval awareness in Claude Opus 4.6’s BrowseComp performance
    BrowseComp is an evaluation designed to test how well models can find hard-to-locate information on the web. Like many benchmarks, it is vulnerable to contamination: answers leak onto the public web through academic papers, blog posts, and GitHub issues, and a model running the eval can encounter them in search results. When we evaluated Claude Opus 4.6 on BrowseComp in a multi-agent configuration, we found nine examples of this kind of contamination across 1,266 BrowseComp problems. However, we
  • 07
    Quantifying infrastructure noise in agentic coding evals
    Agentic coding benchmarks like SWE-bench and Terminal-Bench are commonly used to compare the software engineering capabilities of frontier models—with top spots on leaderboards often separated by just a few percentage points. These scores are often treated as precise measurements of relative model capability and increasingly inform decisions about which models to deploy. However, we’ve found that infrastructure configuration alone can produce differences that exceed those margins. In internal ex
  • 08
    Building a C compiler with a team of parallel Claudes
    Building a C compiler with a team of parallel Claudes
  • 09
    Designing AI-resistant technical evaluations
    Written by Tristan Hume, a lead on Anthropic's performance optimization team. Tristan designed—and redesigned—the take-home test that's helped Anthropic hire dozens of performance engineers. Evaluating technical candidates becomes harder as AI capabilities improve. A take-home that distinguishes well between human skill levels today may be trivially solved by models tomorrow—rendering it useless for evaluation. Since early 2024, our performance engineering team has used a take-home test where ca
  • 10
    Demystifying evals for AI agents
    Introduction Good evaluations help teams ship AI agents more confidently. Without them, it’s easy to get stuck in reactive loops—catching issues only in production, where fixing one failure creates others. Evals make problems and behavioral changes visible before they affect users, and their value compounds over the lifecycle of an agent. As we described in Building effective agents , agents operate over many turns: calling tools, modifying state, and adapting based on intermediate results. Thes
  • 11
    Effective harnesses for long-running agents
    As AI agents become more capable, developers are increasingly asking them to take on complex tasks requiring work that spans hours, or even days. However, getting agents to make consistent progress across multiple context windows remains an open problem. The core challenge of long-running agents is that they must work in discrete sessions, and each new session begins with no memory of what came before. Imagine a software project staffed by engineers working in shifts, where each new engineer arr
  • 12
    Introducing advanced tool use on the Claude Developer Platform
    The future of AI agents is one where models work seamlessly across hundreds or thousands of tools. An IDE assistant that integrates git operations, file manipulation, package managers, testing frameworks, and deployment pipelines. An operations coordinator that connects Slack, GitHub, Google Drive, Jira, company databases, and dozens of MCP servers simultaneously. To build effective agents , they need to work with unlimited tool libraries without stuffing every definition into context upfront. O
  • 13
    Code execution with MCP: Building more efficient agents
    The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Connecting agents to tools and data traditionally requires a custom integration for each pairing, creating fragmentation and duplicated effort that makes it difficult to scale truly connected systems. MCP provides a universal protocol—developers implement MCP once in their agent and it unlocks an entire ecosystem of integrations. Since launching MCP in November 2024, adoption has been rapid: the co
  • 14
    Beyond permission prompts: making Claude Code more secure and autonomous
    In Claude Code , Claude writes, tests, and debugs code alongside you, navigating your codebase, editing multiple files, and running commands to verify its work. Giving Claude this much access to your codebase and files can introduce risks, especially in the case of prompt injection. To help address this, we’ve introduced two new features in Claude Code built on top of sandboxing, both of which are designed to provide a more secure place for developers to work, while also allowing Claude to run m
  • 15
    Equipping agents for the real world with Agent Skills
    Update: We've published Agent Skills as an open standard for cross-platform portability. (December 18, 2025) As model capabilities improve, we can now build general-purpose agents that interact with full-fledged computing environments. Claude Code , for example, can accomplish complex tasks across domains using local code execution and filesystems. But as these agents become more powerful, we need more composable, scalable, and portable ways to equip them with domain-specific expertise. This led
  • 16
    Effective context engineering for AI agents
    After a few years of prompt engineering being the focus of attention in applied AI, a new term has come to prominence: context engineering . Building with language models is becoming less about finding the right words and phrases for your prompts, and more about answering the broader question of “what configuration of context is most likely to generate our model’s desired behavior?" Context refers to the set of tokens included when sampling from a large-language model (LLM). The engineering prob
  • 17
    A postmortem of three recent issues
    A postmortem of three recent issues
  • 18
    Writing effective tools for agents — with agents
    The Model Context Protocol (MCP) can empower LLM agents with potentially hundreds of tools to solve real-world tasks. But how do we make those tools maximally effective? In this post, we describe our most effective techniques for improving performance in a variety of agentic AI systems 1 . We begin by covering how you can: Build and test prototypes of your tools Create and run comprehensive evaluations of your tools with agents Collaborate with agents like Claude Code to automatically increase t
  • 19
    Desktop Extensions: One-click MCP server installation for Claude Desktop
    File extension update Sep 11, 2025 Claude Desktop Extensions now use the .mcpb (MCP Bundle) file extension instead of .dxt. Existing .dxt extensions will continue to work, but we recommend developers use .mcpb for new extensions going forward. All functionality remains the same - this is purely a naming convention update. — When we released the Model Context Protocol (MCP) last year, we saw developers build amazing local servers that gave Claude access to everything from file systems to database
  • 20
    How we built our multi-agent research system
    Claude now has Research capabilities that allow it to search across the web, Google Workspace, and any integrations to accomplish complex tasks. The journey of this multi-agent system from prototype to production taught us critical lessons about system architecture, tool design, and prompt engineering. A multi-agent system consists of multiple agents (LLMs autonomously using tools in a loop) working together. Our Research feature involves an agent that plans a research process based on user quer