Red Hat→
Machine Learning Systems Research Intern, PhD,… at Red Hat · Boston
InternshipHybridBoston, MA
Skills
c++cudapythonpytorchai model optimizationquantizationpruningknowledge distillationgpu performance optimizationlarge language model architecturesefficient inference techniquesopen-source ml frameworksanalytical skills
Job Description
Summary: Red Hat is the world’s leading provider of enterprise open source software solutions, and they are seeking a highly motivated summer intern to join their Machine Learning Research Team. The intern will work on cutting-edge AI inference and model optimization techniques, contributing to research and engineering efforts that enhance LLMs' performance.
Responsibilities:
- Research and implement techniques for LLM inference and LLM optimizations
- Conduct experiments to evaluate the impact of optimization methods on model accuracy, latency, and throughput
- Collaborate with researchers and engineers to integrate optimizations into real-world machine learning workflows
- Document findings and contribute to technical reports, blog posts, or research publications
Required Qualifications:
- Currently pursuing a Ph.D. degree in Computer Science, Electrical Engineering, Machine Learning, or a related field
- Strong programming skills in C++, CUDA, and Python
- Experience with tensor math libraries such as PyTorch
- Familiarity with AI model optimization techniques such as quantization (e.g., INT4, FP8), pruning, and knowledge distillation
- Deep understanding and experience in GPU performance optimizations
- Excellent knowledge of large language model architectures
- Strong analytical and problem-solving skills
- Excellent communication skills and ability to work in a team-oriented research environment
- Background in efficient inference techniques for large-scale language models or computer vision models
- Prior experience contributing to open-source ML frameworks or research publications
Preferred Qualifications:
- 1 or more co-authored papers at a top tier conference like NeurIPS, ICLR, ACL, CVPR, MLSys is a big plus
Required Skills: C++, CUDA, Python, PyTorch, AI model optimization, Quantization, Pruning, Knowledge distillation, GPU performance optimization, Large language model architectures, Efficient inference techniques, Open-source ML frameworks, Analytical skills
Internship Start Date: Start in 2026 Summer
Benefits: Hands-on experience with state-of-the-art AI inference optimization research., Mentorship from leading experts in machine learning and model efficiency., Opportunity to contribute to research papers, patents, or open-source projects., Competitive stipend and potential for full-time opportunities.
Benefits
Hands-on experience with state-of-the-art AI inference optimization research.
Mentorship from leading experts in machine learning and model efficiency.
Opportunity to contribute to research papers, patents, or open-source projects.
Competitive stipend and potential for full-time opportunities.