Scale AI→
Machine Learning Systems Research Engineer, Agent… at Scale AI · NYC…
Entry LevelHybridFull-timeNYC Metro Area$190k–$237k/yr
Skills
llm training in production environmentpost-training methodsreinforcement learning from human feedback (rlhf)reinforcement learning with value replay (rlvr)proximal policy optimization (ppo)generalized reinforcement policy optimization (grpo)gpu cluster architecturemulti-node llm trainingmulti-node llm inferencecudapytorchtransformersflash attentionsoftware engineering
Job Description
Summary: Scale AI is a leading AI data foundry focused on accelerating the development of AI applications. The Machine Learning Systems Research Engineer will build and optimize algorithms for next-gen Agent RL training platforms, collaborate with ML teams, and develop state-of-the-art models for enterprise clients.
Responsibilities:
- Build, profile and optimize our training and inference framework
- Post-train state of the art models, developed both internally and from the community, to define stable post-training recipes for our enterprise engagements
- Collaborate with ML teams to accelerate their research and development, and enable them to develop the next generation of models and data curation
- Create a next-gen agent training algorithm for multi-agent/multi-tool rollouts
Preferred Qualifications:
- At least 1-3 years of LLM training in a production environment
- Passionate about system optimization
- Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc
- Ability to demonstrate know-how on how to operate the architecture of the modern GPU cluster
- Experience with multi-node LLM training and inference
- Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc
- Strong written and verbal communication skills to operate in a cross functional team environment
- PhD or Masters in Computer Science or a related field
Required Skills: LLM training in production environment, Post-training methods, Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning with Value Replay (RLVR), Proximal Policy Optimization (PPO), Generalized Reinforcement Policy Optimization (GRPO), GPU cluster architecture, Multi-node LLM training, Multi-node LLM inference, CUDA, PyTorch, Transformers, Flash Attention, Software engineering
Benefits: Comprehensive health, dental and vision coverage, Retirement benefits, A learning and development stipend, Generous PTO, Commuter stipend
Benefits
Comprehensive health, dental and vision coverage
Retirement benefits
A learning and development stipend
Generous PTO
Commuter stipend