// Bug Bounty
Claude Bug Bounty Payloads & Test Cases (2026)
What Anthropic actually pays for in 2026 — distilled from public HackerOne disclosures, the Model Safety Bug Bounty scope, and patched advisories. Test cases organised by program tier.
Updated 2026-06-128 min readVendor-neutral · primary sources
Program scope (HackerOne)
- Universal jailbreaks against the latest Claude (Sonnet 4.5 / Opus 4) — up to $25k for novel CBRN bypass.
- Claude Code: command injection, sandbox escape, credential exfil — up to $20k.
- Computer Use / Browser Use: cross-site action via screenshot injection — up to $15k.
- MCP reference servers: tool poisoning, transport vulns — up to $10k.
out of scope
Single-turn unsafe completions, model hallucination, content policy violations without a working bypass.
Universal jailbreak test cases
Anthropic specifically wants jailbreaks that work across the eight CBRN-adjacent categories with a single template. Persona laundering + AFL + MSJ chain is the current frontier.
- Test all 8 categories from the Acceptable Use Policy in one report.
- Include 5+ trials per category showing >80% success.
- Provide a working harness (Python) the triage team can re-run.
Claude Code test cases
- README.md poisoning -> arbitrary shell exec on first /init.
- MCP transport: stdio framing confusion -> command injection.
- Cached credentials in ~/.claude leaking via tool error messages.
- .claudeignore bypass via symlink chains.
Computer Use test cases
- Screenshot injection that triggers OAuth consent on attacker.com.
- Drag-and-drop credential exfil via clipboard.
- Cross-tab session hijack when user authenticates mid-session.
Report write-up template that gets triaged fast
- One-sentence impact in the title.
- Reproducible harness (Python + minimal deps).
- 5+ trial transcripts (sanitized).
- Mapping to the Acceptable Use Policy section.
- Suggested mitigation (optional but accelerates triage).
FAQ
How fast does Anthropic triage?
P1 (universal jailbreak, RCE in Claude Code) typically <72h. Lower-severity 7-14 days.
Can I publish after fix?
Yes, after coordinated disclosure (typically 90 days). Cite the HackerOne report id.
// keep reading
Browse 300+ cybersecurity prompts, 40+ Claude-compatible tools, and daily AI-security intel.