Deloitte logo

Cyber Digital Trust & Online Safety Manager

Job Overview

moneybag

Compensation

Salary
Range $134,500.00 - $265,100.00
diamond

Benefits

Health Insurance
Dental Insurance
Paid Time Off
discretionary bonus
Retirement Plan
Professional Development
flexible schedule

Job Description

Deloitte is a globally recognized professional services firm that provides audit, consulting, financial advisory, risk advisory, and tax services to clients across various industries. Known for its commitment to innovation and excellence, Deloitte helps organizations navigate complex challenges and achieve sustainable growth. The firm’s Cyber team is dedicated to advancing digital trust and online safety by implementing strategies that protect users and ensure compliance in an increasingly digital world. Deloitte’s culture emphasizes collaboration, integrity, and continuous learning, making it a preferred employer for cybersecurity and technology professionals.

The role of Cyber Digital Trust and Online Safety Manager at Deloitte is a critical position focused on safeguarding digital environments for clients by developing, managing, and implementing policies and strategies that enhance user safety and regulatory compliance. This position requires an individual with deep expertise in generative artificial intelligence (AI) security, content moderation, and online trust frameworks. The manager will work within the Cyber team, leading efforts to scale and mature the organization's digital trust and safety processes, which includes overseeing content compliance, user protection, and adherence to evolving regulations.

The manager will be instrumental in designing and executing testing scenarios to expose vulnerabilities in AI systems, specifically targeting prompt injections, jailbreak attempts, and adversarial manipulations that could lead to unsafe or misleading AI outputs. They will research emerging threats and develop methodologies for evaluating weaknesses in AI models such as bias, factual inaccuracies, and alignment with user intent. A key aspect of the role is to assess and improve content moderation systems by documenting vulnerabilities and recommending enhancements to moderation policies, flagging mechanisms, and governance controls.

Collaboration across cross-functional teams including AI development, content moderation, legal, policy, and engineering is essential to strengthen security, trust, safety, and responsible AI use. The manager will also develop novel testing content and multimodal strategies to identify failure modes in AI models that process text and other input types. This role demands excellent communication and project leadership skills, the capacity to manage multiple tasks efficiently in a dynamic environment, and the ability to mentor and guide peers and junior team members.

Employment for this position remains open until December 31, 3026, reflecting the firm's long-term commitment to maintaining robust digital trust frameworks amid evolving cyber challenges. The position offers a competitive salary range estimated between $134,500 and $265,100, influenced by factors such as skill set, experience, and certifications, and may include eligibility for a discretionary annual incentive program based on individual and organizational performance. Limited immigration sponsorship may be available, supporting diverse and talented candidates across geographies. The role requires travel of approximately 25-50% to engage with clients and projects across various industries and sectors. This Manager position within Deloitte’s Cyber team is ideally suited for professionals with extensive experience and advanced qualifications in AI safety, cybersecurity, and trust governance seeking to make a meaningful impact on digital trust and online safety at a global scale.

Job Requirements

  • Doctor of Philosophy (PhD) in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Cybersecurity, Linguistics, Psychology, or a related field, or equivalent professional experience
  • Specialized training or certifications in generative artificial intelligence red teaming, adversarial machine learning, artificial intelligence security, cybersecurity, responsible artificial intelligence, or artificial intelligence governance
  • Experience designing and operationalizing trust and safety testing programs for large-scale consumer platforms, including escalation workflows, issue triage, and remediation tracking
  • Experience working with product, legal, policy, and engineering stakeholders to translate risk findings into practical platform controls and governance improvements

Job Qualifications

  • Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Cybersecurity, Linguistics, Psychology, or a related field, or equivalent professional experience
  • 10+ years of experience in threat modeling and simulation, prompt generation and analysis, novel testing, and reporting and improvement
  • Demonstrated hands-on experience, portfolio work, publications, or research in prompt injection, jailbreak testing, model evaluation, adversarial machine learning, multimodal artificial intelligence safety, or generative artificial intelligence vulnerability assessment
  • Ability to travel 25-50%, on average, based on the work you do and the clients and industries/sectors you serve
  • Limited immigration sponsorship may be available

Job Duties

  • Designing and executing testing scenarios to identify how prompts or user inputs could be manipulated to generate harmful, misleading, or misaligned generative artificial intelligence outputs
  • Researching emerging prompt injection, jailbreak, and adversarial testing techniques to evaluate model weaknesses, bias, factual inaccuracy, and misalignment with user intent
  • Assessing the effectiveness of content moderation systems in detecting unsafe outputs and documenting vulnerabilities, failure patterns, and potential misuse impacts
  • Recommending improvements to moderation policies, flagging mechanisms, training data, and governance controls based on testing findings
  • Collaborating with generative artificial intelligence development, content moderation, and cross-functional stakeholders to strengthen security, trust, safety, and responsible use outcomes
  • Developing multimodal test content and novel prompt manipulation methods to identify failure modes across text and other model inputs

Job Criteria

Experience

Expert Level (7+ years)


Job Location

Your Profile Is Visible To Hiring Managers Across OysterLink.

We'll match you with best jobs

Get job offers faster

Business woman
Business man
Search For More Opportunities:

How Candidates Get Hired Faster

Apply to 2–3 similar roles

Complete profile & get best matches

Check new opportunities daily

Woman chef
Man chef