OpenAI disrupts coordinated AI model distillation campaign
OpenAI has identified and disrupted a coordinated adversarial distillation campaign targeting its models, with first activity observed in July. Operators manipulated model interactions to extract protected reasoning, violating terms of service, without breaking encryption. The campaign peaked on July 24-25 with 16,000 requests from over 4,000 users, and a related cluster of 15,000 users was fully disrupted by July 28. OpenAI attributes the core cluster to individuals associated with Moonshot AI. The company strengthened protections, banned accounts, and shared findings with the Frontier Model Forum and government channels. Adversarial distillation poses safety and national security risks, potentially enabling training models without safeguards. OpenAI expects these attempts to become more sophisticated and continues to enhance defenses across first-party and partner deployments.
- •OpenAI disrupted a coordinated adversarial distillation campaign targeting its models, with first activity observed in July.
- •The campaign peaked on July 24-25 with 16,000 requests from over 4,000 users.
- •Related prompt-pattern activity involved a cluster of more than 15,000 users, fully disrupted by July 28.
