AI Safety Expert

1 week ago

Toronto ON, Toronto Census Division, ON; Ontario, Canada Mercor Full-time €16 - €22

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: AI Safety Experts - English & Odia
Type: Contract
Compensation: $16-$22/hour
Location: Remote

Role Responsibilities

  • Red team conversational AI models and agents. Conduct jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure. Follow taxonomies, benchmarks, and playbooks to ensure consistent testing.
  • Document reproducibly. Produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously . Thrive in flexible hours while improving AI model performance .

Qualifications

Must-Have

  • Native fluency in English and Odia .
  • Strong judgment about language and content.
  • Rigorous attention to detail and consistency.
  • Structured approach to guidelines and quality standards.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, task types, and customers.

Preferred

  • Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Background in Cybersecurity : penetration testing, exploit development, reverse engineering.
  • Knowledge of socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.

Resources & Support

  • For details about the interview process and platform information, please check:
  • For any help or support, reach out to: support@mercor.com
#J-18808-Ljbffr