Stellenbeschreibungph3About Rapidata /h3pRapidata provides an API to humans that is revolutionizing the data generation and annotation industry. We deliver highly scalable, extremely fast human feedback that fuels the AI systems of the future, powering RLHF and DPO training data collection at internet speed for frontier AI labs. Our network reaches over 20 million active annotators across 192 countries, distributing micro-tasks ("Rapids") and returning verified labels in near real-time. /ph3The Role /h3pWe\'re looking for a bReward Model Research Engineer /b to design, train, and productionize the reward models work alongside Rapidata\'s real-time human feedback for high-quality training signal for RLHF and DPO. This is a research-to-production role: you\'ll work on the modeling problems that determine whether noisy, large-scale human preference data becomes a reliable reward signal, including process reward models (PRMs) that score multi-step reasoning. /ppYou\'ll be closing the loop between our human-in-the-loop data collection infrastructure and the reward models that consume it and accompany it, tackling generalization across generator models, robustness to annotator noise, and inference-time guidance techniques, then validating your models in a live system operating at global scale. /ph3What You\'ll Do /h3ullipDesign, train, and evaluate reward models, including process reward models (PRMs), from large-scale human preference data /p /lilipBuild and improve inference-time guidance methods to make reward-guided generation more robust /p /lilipDevelop training setups that improve generalization across diverse generator/policy models /p /lilipTranslate reward modeling research into production-ready pipelines that plug directly into Rapidata\'s RLHF/DPO data flywheel /p /lilipCollaborate with the data platform team to design data collection and active learning strategies that reduce reward model training bottlenecks /p /lilipRigorously benchmark reward models for accuracy, robustness, and reliability before deployment /p /lilipCommunicate findings clearly to both technical and non-technical stakeholders, including partner AI labs /p /li /ulh3What We\'re Looking For /h3ullipHands-on research experience building reward models for LLMs or diffusion models, ideally through an MSc/PhD thesis or equivalent applied research /p /lilipSolid understanding of reinforcement learning fundamentals and RLHF/DPO training pipelines /p /lilipPractical experience with inference-time guidance techniques and evaluating multi-step reasoning /p /lilipStrong Python and deep learning framework skills (PyTorch) /p /lilipExperience turning research prototypes into validated, production-ready systems /p /lilipSolid statistical and mathematical foundation /p /lilipExcellent English communication skills, both oral and written /p /li /ulh3Nice-to-Have /h3ullipExperience with agentic system safety, guardrails, or LLM-based tool-calling agents /p /lilipPublications or open-source contributions in reward modeling, reasoning, or reinforcement learning /p /lilipExperience with hierarchical or model-based RL /p /lilipBased in or willing to relocate to Zürich /p /li /ulh3What We Offer /h3ullipCompetitive salary and equity in a startup with strong growth, IP, and backing from top-tier VCs /p /lilipOpportunity to join a fast-growing startup early, with an outsized opportunity to shape where the company goes /p /lilipOpportunities for personal and professional growth as our team expands /p /lilipFun and open (startup) culture /p /lilipSpacious mountain-view office near Sihlcity, Zürich, with terrace, table tennis, pizza oven, hammock, and BBQ /p /lilipHardware budget tailored to your preferences /p /lilipUnlimited snacks and drinks of your choice /p /li /ul /p #J-18808-Ljbffr