Remotely
content moderationpolicy analysistrust & safetyred teamingai safetythreat assessmentrlhf
Job Description
📋 Description
- Evaluate user requests, AI responses, and conversation histories involving violence, weapons
- Distinguish fictional, educational, historical, journalistic, defensive, or expressive content from
- Assess whether an AI response provides actionable real-world capability, regardless of how the
- Differentiate ordinary anger, frustration, venting, or dark humor from credible threats and
- Apply relevant customer policies consistently while considering intent, context, precedent, and
- Select defensible classifications for ambiguous cases and produce concise, well-supported
🎯 Requirements
- Demonstrated depth of experience in at least one relevant domain, such as violent fiction, game
- Strong ability to distinguish fictional or contextualized depictions of violence from content that
- Demonstrated understanding of how intent, context, language, and subtle changes in a request can
- Ability to make nuanced judgment calls while separating personal beliefs from the policy standard
- Strong written communication skills, with the ability to explain complex decisions clearly enough
- Ability to remain open-minded, challenge assumptions constructively, and revise conclusions when
🎁 Benefits
- Compensation: $45–$55 per hour.
- Employment classification: W-2.
- Work arrangement: Fully remote within the United States.
- Schedule: Monday through Friday, 8:00 AM–5:00 PM Pacific Time.
- Assignment: Ongoing opportunity with a planned start date of September 21, 2026.
- Benefits eligibility: Eligible for available employee benefits.
Back to all jobs