Centific→
Research Intern — Applied Reinforcement Learning at Centific in Remote
InternshipRemoteFull-timeRemote$73k–$94k/yr
Skills
phd in csmlreinforcement learningpythonpytorchgpu-based trainingllms experiencerl environmentssimulation techniquesexperimentation practicesdistributed trainingexperiment tracking
Job Description
Summary: Centific is a frontier AI data foundry that empowers clients with scalable AI deployment. The PhD Research Intern will design and evaluate reinforcement learning systems for agentic AI workflows, developing RL environments and translating research into practical solutions.
Responsibilities:
- End-to-end RL pipelines for agentic systems (simulation → training → evaluation)
- Alignment of LLM-based agents using RLHF, DPO, PPO, and emerging methods
- Design of reward functions, verifiers, and evaluation frameworks
- Simulation environments (digital twins) for enterprise workflows
- Scalable training and inference for RL-based systems
- Build a custom RL environment simulating a real-world enterprise workflow and train an agent using PPO or GRPO
- Develop a reward modeling pipeline from human feedback and evaluate alignment improvements
- Create an evaluation harness measuring reasoning, task success, and policy safety
- Prototype an agentic system with tool use and multi-step reasoning, integrated with RL training
- Document experiments, ablations, and findings for research and productionization
Required Qualifications:
- PhD candidate in CS, ML, or related field with research in reinforcement learning or agentic AI
- Strong Python and PyTorch skills with GPU-based training experience
- Solid understanding of RL fundamentals (MDPs, policy gradients, value methods)
- Experience with LLMs and post-training techniques (RLHF, DPO, PPO, etc.)
- Strong experimentation practices (ablation, reproducibility, clear reporting)
Preferred Qualifications:
- Experience with RL environments (Gymnasium, RLlib, Stable Baselines)
- Research in offline RL, model-based RL, or hierarchical RL
- Publications at top ML conferences (NeurIPS, ICML, ICLR, ACL)
- Experience with simulation, synthetic data, or multi-agent systems
- Distributed training and large-scale experimentation
Required Skills: PhD in CS, ML, Reinforcement Learning, Python, PyTorch
Important Skills: GPU-based training, LLMs experience, RL environments, Simulation techniques
Nice-to-Have Skills: Experimentation practices, Distributed training, Experiment tracking
Benefits: Competitive stipend and real-world impactful projects, Mentorship from researchers and engineers, Access to modern GPU infrastructure, Opportunities to publish and present research
Benefits
Competitive stipend and real-world impactful projects
Mentorship from researchers and engineers
Access to modern GPU infrastructure
Opportunities to publish and present research