Jobiglo

No results.

RLHF Specialist – Remote

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow LoRA QLoRA Llama 2 Mistral Gemma LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

Job description

About the role

We are looking for an RLHF Specialist to design, implement, and optimise reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factual accuracy, and alignment of large language models. This fully remote position works closely with machine‑learning engineers to turn human preferences into high‑quality training signals.

Key responsibilities

  • Generate high‑quality preference data by ranking model responses on helpfulness, honesty, and harmlessness.
  • Design multi‑turn prompts that stress‑test model reasoning and safety.
  • Write chain‑of‑thought explanations to train reward models.
  • Collaborate with ML engineers to analyse failure modes and identify data gaps.
  • Develop and iterate annotation strategies, ensuring consistency across a global team.
  • Probe models for biases, hallucinations, and vulnerabilities, documenting findings.
  • Analyse edge cases where reward models behave unexpectedly and suggest data interventions.
  • Translate complex RL concepts into repeatable tasks for junior reviewers.
  • Maintain a personal test set of prompts to monitor model performance over time.

Required profile

  • Minimum 2 years of experience in data annotation, model evaluation, computational linguistics, or AI trust & safety.
  • Strong proficiency in Python and deep‑learning frameworks (PyTorch, JAX, or TensorFlow).
  • Deep understanding of reinforcement‑learning algorithms (PPO, trust‑region methods, reward hacking).
  • Hands‑on experience fine‑tuning open‑source models (Llama 2/3, Mistral, Gemma) with LoRA/QLoRA.
  • Experience with annotation platforms such as LabelBox, Scale AI, or Snorkel.
  • Familiarity with cloud ML services (AWS SageMaker, GCP Vertex AI).

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • Reinforcement Learning (PPO, Trust Regions, Reward Hacking)
  • LoRA / QLoRA fine‑tuning
  • Open‑source LLMs (Llama 2/3, Mistral, Gemma)
  • Annotation tools (LabelBox, Scale AI, Snorkel)
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 6 days from now

50 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Odixcity Consulting