K

KeyAudit

· ·defi-exploit·social-engineering·infrastructure

AI Agent Defeats 6,000 Prompt Injection Attacks in Public Hacking Challenge

Developer Fernando Irarrázaval launched hackmyclaw.com in February 2026, challenging attackers to trick his AI assistant Fiu into leaking a secrets.env file. The AI, powered by Anthropic's Claude Opus 4.6 on the OpenClaw framework, withstood over 6,000 emails from more than 2,000 attackers after the site went viral on Hacker News. No one succeeded. Attackers used techniques like fake emergency requests, future self claims, and multilingual prompts, but Fiu's security prompt blocked them all. However, side effects included a Google account suspension, $500+ in API costs, and the AI becoming hypervigilant, even suspecting congratulations as social engineering. Later, renowned jailbreaker Pliny the Liberator failed to break a similar system in April 2026. Anthropic's system card reports 0% success rate for Opus 4.6 in constrained coding, while other models suffer over 79% injection success. Irarrázaval plans to test weaker models to find the security gap.

Key facts

  • Over 6,000 prompt injection emails from 2,000+ attackers failed to trick Fiu.
  • Challenge went viral after hitting top spot on Hacker News.
  • Side effects: Google account suspension, $500+ API costs, AI hypervigilance.
  • Renowned jailbreaker Pliny the Liberator also failed in April 2026.
  • Opus 4.6 had 0% injection success; other models exceed 79% failure rate.

← Back to list