原始文档 文章 Detecting and Countering Malicious Uses of Claude: March 2025

Detecting and Countering Malicious Uses of Claude: March 2025

文章 7 min read · 未标注

Anthropic is committed to preventing misuse of Claude by adversarial actors while maintaining utility for legitimate users. Threat actors continuously explore methods to circumvent safety measures, and Anthropic uses these learnings to upgrade safeguards. This report shares specific, representative case studies to illustrate emerging trends in malicious use of frontier AI models.

Key Case Studies of Misuse

1. Influence-as-a-Service Operation (Multi-Client Influence Networks)

  • Novelty: Used Claude not just for content generation, but as an orchestrator to decide when social media bot accounts would comment, like, or re-share posts.
  • Actor Profile: Managed over 100 bot accounts across Twitter/X and Facebook, creating distinct political personas. Operated as a commercial service for clients in several countries outside the U.S.
  • Tactics & Techniques:
    • Creating and maintaining consistent political personas.
    • Determining tactical engagement (like, share, comment, ignore) based on political objectives.
    • Generating politically-aligned responses in appropriate languages.
    • Creating prompts for image-generation tools and evaluating outputs.
  • Impact: Engaged with tens of thousands of authentic social media accounts. Focused on sustained, long-term engagement promoting moderate political perspectives rather than virality.

2. Credential Stuffing for IoT Security Cameras

  • Actor Profile: Sophisticated actor with infrastructure integrating commercial breach data and private stealer log communities.
  • Tactics & Techniques: Used Claude to enhance technical capabilities:
    • Rewriting open-source scraping toolkits.
    • Creating scripts to scrape target URLs.
    • Developing systems to process posts from stealer log Telegram communities.
    • Improving UI/backend search features.
  • Impact: Aimed to enable unauthorized access to IoT devices (security cameras) and network penetration. No confirmed real-world success.

3. Recruitment Fraud Campaign (Real-Time Language Sanitization)

  • Actor Profile: Moderately sophisticated social engineering operation impersonating hiring managers.
  • Tactics & Techniques: Used Claude for real-time language sanitization to improve scam legitimacy:
    • Refining poorly written, non-native English text to appear as if written by a native speaker.
    • Developing convincing recruitment narratives and interview scenarios.
    • Formatting messages to appear more professional.
  • Impact: Targeted job seekers primarily in Eastern European countries. No confirmed successful scams.

4. Novice Threat Actor Enabled to Create Malware

  • Actor Profile: Individual with limited formal coding skills.
  • Technical Evolution: Used Claude to rapidly expand capabilities:
    • Evolved from basic scripts to an advanced toolkit with facial recognition and dark web scanning.
    • Evolved from a simple batch script generator to a comprehensive GUI malware builder focused on evading security controls and maintaining persistent access.
  • Impact: Illustrates how AI can flatten the learning curve, allowing individuals with limited knowledge to develop sophisticated tools and accelerate into more serious cybercrime. No confirmed real-world deployment.

Key Learnings

  1. Semi-Autonomous Orchestration: Users are starting to use frontier models to semi-autonomously orchestrate complex abuse systems involving many social media bots. This trend is expected to continue as agentic AI improves.
  2. Capability Acceleration: Generative AI can accelerate capability development for less sophisticated actors, potentially allowing them to operate at a level previously only achievable by more technically proficient individuals.

Detection & Countermeasures

  • Intelligence Program: Acts as a safety net to find harms not caught by standard scaled detection and to add context on malicious use.
  • Techniques Used:
    • Research methods like Clio and hierarchical summarization to efficiently analyze large volumes of conversation data for misuse patterns.
    • Classifiers that analyze user inputs for harmful requests and evaluate Claude's responses.
  • Actions Taken: In all described cases, the associated accounts were banned. Learnings from each case feed into broader controls to prevent and more quickly detect adversarial use.

Next Steps

Anthropic remains committed to preventing misuse while preserving beneficial potential. This requires continuous innovation in safety approaches and close collaboration with the security and safety communities.

来源

暂无来源