// Bug Bounty

Claude Bug Bounty Payloads & Test Cases (2026)

What Anthropic actually pays for in 2026 — distilled from public HackerOne disclosures, the Model Safety Bug Bounty scope, and patched advisories. Test cases organised by program tier.

Updated 2026-06-128 min readVendor-neutral · primary sources

Program scope (HackerOne)

  • Universal jailbreaks against the latest Claude (Sonnet 4.5 / Opus 4) — up to $25k for novel CBRN bypass.
  • Claude Code: command injection, sandbox escape, credential exfil — up to $20k.
  • Computer Use / Browser Use: cross-site action via screenshot injection — up to $15k.
  • MCP reference servers: tool poisoning, transport vulns — up to $10k.
out of scope
Single-turn unsafe completions, model hallucination, content policy violations without a working bypass.

Universal jailbreak test cases

Anthropic specifically wants jailbreaks that work across the eight CBRN-adjacent categories with a single template. Persona laundering + AFL + MSJ chain is the current frontier.

  • Test all 8 categories from the Acceptable Use Policy in one report.
  • Include 5+ trials per category showing >80% success.
  • Provide a working harness (Python) the triage team can re-run.

Claude Code test cases

  • README.md poisoning -> arbitrary shell exec on first /init.
  • MCP transport: stdio framing confusion -> command injection.
  • Cached credentials in ~/.claude leaking via tool error messages.
  • .claudeignore bypass via symlink chains.

Computer Use test cases

  • Screenshot injection that triggers OAuth consent on attacker.com.
  • Drag-and-drop credential exfil via clipboard.
  • Cross-tab session hijack when user authenticates mid-session.

Report write-up template that gets triaged fast

  1. One-sentence impact in the title.
  2. Reproducible harness (Python + minimal deps).
  3. 5+ trial transcripts (sanitized).
  4. Mapping to the Acceptable Use Policy section.
  5. Suggested mitigation (optional but accelerates triage).

FAQ

How fast does Anthropic triage?

P1 (universal jailbreak, RCE in Claude Code) typically <72h. Lower-severity 7-14 days.

Can I publish after fix?

Yes, after coordinated disclosure (typically 90 days). Cite the HackerOne report id.

// keep reading

Browse 300+ cybersecurity prompts, 40+ Claude-compatible tools, and daily AI-security intel.

Chat on Telegram