AI Safety Experts — English & Norwegian in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Expert — English & Norwegian based in Canada.
This is a fully remote opportunity to contribute directly to the development of safer, more reliable AI systems.
You will join a red-team-focused project designed to identify vulnerabilities that automated evaluations may overlook.
The role combines adversarial thinking, structured testing, and high-quality human data generation.
You will challenge conversational AI models through jailbreaks, prompt injection, misuse scenarios, and multi-turn manipulation.
Your findings will help strengthen AI safety benchmarks and improve the robustness of systems used in real-world environments.
The work is text-based, flexible, and conducted across evolving projects and customer needs.
Sensitive topics may arise, but participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources.
- Red-team conversational AI models and agents by testing jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation scenarios.
- Generate high-quality human evaluation data by annotating model failures, classifying vulnerabilities, and identifying systemic safety risks.
- Apply established taxonomies, benchmarks, testing frameworks, and playbooks to ensure evaluations are consistent and reproducible.
- Create structured attack cases, datasets, and reports that clearly document vulnerabilities and provide actionable insights for improving AI systems.
- Identify weaknesses that automated testing may miss and develop creative adversarial approaches to uncover unexpected failure modes.
- Communicate technical and socio-technical risks clearly to both technical and non-technical stakeholders.
- Adapt quickly across different projects, AI systems, evaluation methodologies, and customer requirements.
- Contribute to expanding evaluation coverage and reducing the likelihood of unexpected AI safety issues reaching production environments.
- Native-level fluency in both English and Norwegian, with excellent written communication skills in both languages.
- Prior experience in red teaming, AI adversarial testing, cybersecurity, socio-technical probing, conversational AI evaluation, or a closely related field.
- Strong curiosity and an adversarial mindset, with the ability to systematically push AI systems toward their limits.
- Experience using structured frameworks, benchmarks, taxonomies, or testing methodologies rather than relying solely on ad hoc experimentation.
- Strong analytical and documentation skills, with the ability to explain vulnerabilities and risks clearly and precisely.
- Ability to work independently, adapt to changing project requirements, and move effectively between different AI safety challenges.
- Knowledge of adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction, is a plus.
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering is an advantage.
- Experience with harassment or misinformation testing, abuse analysis, conversational AI testing, psychology, acting, or creative writing for adversarial scenarios is also valued.
- Fully remote work with the flexibility to complete assignments on your own schedule.
- Independent contractor engagement with project-based flexibility.
- Weekly payments through Stripe or Wise based on services rendered.
- Opportunity to gain hands-on experience in human-data-driven AI safety and red teaming.
- Direct contribution to making advanced AI systems more robust, safe, and trustworthy.
- Exposure to cutting-edge AI safety projects and evolving evaluation methodologies.
- Projects may be extended, shortened, or concluded early depending on project needs and performance.
- No requirement to access confidential or proprietary information from previous employers, clients, or institutions.
- Referral opportunity: earn up to $250 for each successful referral, subject to applicable restrictions.
- Participation in higher-sensitivity projects is optional, with topics communicated in advance and supported by clear guidelines and wellness resources.
- Please note that H-1B and STEM OPT candidates cannot currently be supported for this engagement.