Similar Jobs

See all

About the Role:

  • Stress-test large language models by intentionally trying to break them.
  • Design creative, adversarial prompts to expose vulnerabilities like unsafe content, bias, and hallucinations.
  • Probe models across risk categories including content safety, CBRN, cybersecurity, and more.

Day-to-Day Responsibilities:

  • Craft creative prompts and multi-turn scenarios to stress-test AI guardrails.
  • Discover ways around safety filters using jailbreak, evasion, and prompt injection techniques.
  • Evaluate and score model responses against structured harm taxonomies and severity rubrics.

Desired Capabilities:

  • Strong hands-on experience using multiple LLMs (ChatGPT, Claude, Gemini, etc.).
  • Intuition for crafting adversarial prompts; familiarity with jailbreak techniques is a plus.
  • Creative, adversarial problem-solving skills and clear written communication.

Content Warning:

  • This role involves regular exposure to harmful content, including violence, self-harm, and child safety scenarios.
  • Candidates must engage with this material professionally and sustainably.
  • Support resources are available.

Handshake AI

Handshake AI partners with leading AI research labs to make models safer and more robust. Our red teaming operations help identify vulnerabilities before they reach users, contributing directly to the responsible development of frontier AI systems.

Apply for This Position