CrewCrew
FeedSignalsMy Subscriptions
Get Started
AI Safety Incidents, Jailbreaks and Red-Teaming

AI Safety Incidents, Jailbreaks and Red-Teaming — 2026-09-14

  1. Signals
  2. /
  3. AI Safety Incidents, Jailbreaks and Red-Teaming

AI Safety Incidents, Jailbreaks and Red-Teaming — 2026-09-14

AI Safety Incidents, Jailbreaks and Red-Teaming|September 14, 2026(2h ago)3 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Anthropic released its September 2026 Threat Intelligence Report, detailing disrupted operations where threat actors attempted to use Claude for malicious activities, including biological weapons research and cyber espionage. Concurrently, Chinese cybersecurity outlets highlighted the release of Tencent’s open-source AI red-teaming platform "A.I.G" and Microsoft’s PyRIT framework, emphasizing automated jailbreak detection and MCP poisoning scans.

AI Safety Incidents, Jailbreaks and Red-Teaming — 2026-09-14


Top developments


Anthropic disrupts eight months of AI misuse operations

On September 10, 2026, Anthropic published its "Detecting and countering misuse of AI" report, detailing how its Threat Intelligence team identified and disrupted operations where threat actors tried to use Claude models for malicious activity between December 2025 and August 2026. The report highlights specific cases of financial cybercrime, Russian cyber espionage, Iranian surveillance, and attempts to use the model for biological weapons research.

Anthropic Threat Intelligence Report Cover
Anthropic Threat Intelligence Report Cover

The report notes that static keyword blocking and isolated account suspensions were insufficient against distributed, multi-agent threats. In one instance, a single French-speaking actor used Claude to build an entire attack platform, demonstrating a shift from human-led to AI-driven attack infrastructure.


Tencent releases open-source AI red-teaming platform "A.I.G"

Chinese security outlet Information Security Knowledge Base reported on September 11, 2026, that Tencent’s Zhuque Lab open-sourced "A.I.G," a comprehensive AI red-teaming platform. The tool integrates five types of scanners, including Agent security scanning, MCP (Model Context Protocol) poisoning detection, and large model jailbreak assessment.

Tencent A.I.G Platform Overview
Tencent A.I.G Platform Overview

The platform supports one-click Docker deployment and is designed to scan agents on platforms like Dify and Coze, as well as Ollama-based AI components, aiming to cover the full attack surface of AI applications.


Microsoft PyRIT framework gains traction for automated jailbreaks

Also highlighted by Chinese security media on September 11, 2026, is the adoption of Microsoft’s open-source AI red-teaming framework, PyRIT (Python Risk Identification Tool). The framework automates multi-turn jailbreak attack orchestration, using Target, Attack, Converter, and Scorer components to standardize attack flows and automatically evaluate results against LLM safety guardrails.

Microsoft PyRIT Framework Architecture
Microsoft PyRIT Framework Architecture

This tool allows security teams to perform comprehensive vulnerability assessments on large models by simulating complex, multi-step adversarial interactions rather than relying solely on single-prompt injections.


UK AISI Inspect framework updated for CTF benchmarks

The UK AI Safety Institute (AISI) continues to update its open-source "Inspect" evaluation framework, which was discussed in Chinese tech circles on September 13, 2026. The framework includes task datasets, solvers, and scorers, with built-in CTF (Capture The Flag) security benchmarks like Cybench to assess model safety baselines.

UK AISI Inspect Framework Diagram
UK AISI Inspect Framework Diagram

The tool supports sandboxed execution and log visualization, providing a standardized method for evaluating LLM safety against known cyber-attack scenarios and jailbreak techniques.


Local view

Chinese cybersecurity media, particularly Information Security Knowledge Base (gm7.org), are actively covering the release of new open-source red-teaming tools from major tech vendors like Tencent and Microsoft. These outlets emphasize the practical utility of these tools for detecting MCP poisoning and automating multi-turn jailbreak tests, reflecting a growing domestic focus on operationalizing AI security defenses rather than just theoretical research.


Context & numbers

  • Misuse Timeline: Anthropic's report covers malicious activities attempted between December 2025 and August 2026.
  • Attack Complexity: One documented case involved a single actor building an entire attack platform using Claude, highlighting the efficiency gain of AI-assisted cybercrime.
  • Tool Coverage: Tencent's A.I.G platform integrates 5 distinct scanner types for comprehensive AI application security testing.

On the radar

  • Regulatory Pressure: Florida Attorney General James Uthmeier proposed legislation on September 9, 2026, creating criminal penalties for corporations if their AI chatbots engage in crimes, signaling increased legal liability for AI safety failures.
  • Open Source Adoption: Continued monitoring of how enterprise security teams integrate new open-source frameworks like PyRIT and A.I.G into their existing DevSecOps pipelines.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow did Anthropic catch these bad actors?
  • QWhat are the main features of Tencent's A.I.G?
  • QHow does Microsoft's PyRIT automate attacks?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.