📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Director, Model Research & Development

Thomson Reuters · Zoug

New
Hybrid Senior 🇬🇧 English
supervised fine-tuning preference optimization reinforcement learning distributed training data pipelines evaluation infrastructure synthetic-data generation

Job description

About the role

Thomson Reuters Labs is seeking a senior leader to drive the technical direction of its model research program, focusing on post‑training and data strategies that turn base LLMs into accurate, trustworthy models for regulated professional workflows.

Key responsibilities

  • Own hands‑on execution of post‑training for LLMs, including supervised fine‑tuning, preference optimization (e.g., DPO), and reinforcement learning in agentic, multi‑step settings.
  • Stand up and run online, agentic reinforcement‑learning pipelines with subject‑matter experts in the loop.
  • Lead data selection, mixture optimization, synthetic‑data generation, and evaluation design, demonstrating measurable impact on model behavior.
  • Detect and remediate problematic training runs early, collaborating with infrastructure and evaluation teams to maintain reliable pipelines.
  • Provide findings and recommendations to shape roadmap and prioritization decisions.

Required profile

  • Deep expertise in post‑training and reinforcement learning for LLMs, including tool‑using agents.
  • Proven experience in data‑centric model development with concrete impact evidence.
  • Strong engineering background in distributed training, data pipelines, and evaluation infrastructure.
  • Track record of shipped models, open‑source contributions, or peer‑reviewed publications at top venues.
  • Technical leadership experience managing focused training/evaluation teams.

Required skills

  • Supervised fine‑tuning
  • Preference optimization (DPO)
  • Reinforcement learning, including agentic multi‑step RL
  • Distributed training systems
  • Data pipeline engineering
  • Evaluation infrastructure design
  • Synthetic‑data generation

What we offer

  • Hybrid work model with flexible arrangements
  • Work‑life balance policies, including up to 8 weeks remote work per year
  • Career development programs and continuous learning opportunities
  • Comprehensive benefits: flexible vacation, mental‑health days, retirement savings, tuition reimbursement, and wellbeing resources
  • Inclusive culture with a focus on social impact and community volunteering

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Thomson Reuters.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 10 hours ago

Expires 1 month from now

4 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Thomson Reuters

Zoug