ByteDance→
Student Researcher (AI Foundation Models… at ByteDance · San Jose
InternshipOn-siteSan Jose, CA$125k–$125k/yr
Skills
pythonc++distributed training frameworkspytorch fsdpmegatron-style parallelismreinforcement learning training systemsgpu programmingcudatritoncompiler technologieslarge-scale inference optimizationperformance profilingsystems debuggingmachine learning systemsdistributed computingperformance optimization
Job Description
Summary: ByteDance is a leading technology company focused on pioneering new paths toward artificial general intelligence. They are seeking a PhD Intern for their Seed Infrastructures team to work on distributed training systems, reinforcement learning frameworks, and performance optimization for AI foundation models.
Responsibilities:
- Design and optimize large-scale distributed training systems (e.g., data/model/pipeline parallelism, memory efficiency, fault tolerance)
- Contribute to reinforcement learning training frameworks and large-scale post-training systems
- Improve inference performance, latency, and throughput for foundation models
- Develop compiler or runtime optimizations for heterogeneous hardware (GPU/accelerator)
- Work on system-level performance analysis, profiling, and bottleneck diagnosis
- Build tooling and automation to improve developer productivity and system reliability
Required Qualifications:
- Currently pursuing a PhD degree in Computer Science, Electrical Engineering, or related technical fields
- Strong programming skills in Python and/or C++
- Solid understanding of systems, distributed computing, machine learning systems, or performance optimization
- Experience with one or more of the following: Distributed training frameworks (e.g., PyTorch FSDP, Megatron-style parallelism); Reinforcement learning training systems; GPU programming (CUDA, Triton) or compiler technologies; Large-scale inference optimization; Performance profiling and systems debugging
- Strong problem-solving skills and the ability to work in fast-paced research-driven environments
Preferred Qualifications:
- Experience working on large-scale ML systems or infrastructure projects
- Contributions to open-source ML systems or performance tooling
- Publications in ML systems, distributed systems, or related areas (a plus but not required)
Required Skills: Python, C++, Distributed training frameworks, PyTorch FSDP, Megatron-style parallelism, Reinforcement learning training systems, GPU programming, CUDA, Triton, Compiler technologies, Large-scale inference optimization, Performance profiling, Systems debugging, Machine learning systems, Distributed computing, Performance optimization
Internship Start Date: Start in 2026
Benefits: Interns have day one access to health insurance, life insurance, wellbeing benefits and more., Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year)., Interns who are not working 100% remote may also be eligible for housing allowance.
Benefits
Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year).
Interns who are not working 100% remote may also be eligible for housing allowance.