National Laboratory of the Rockies→
Graduate Summer Intern – HPC Operational… at National… · Remote
InternshipRemoteRemote$44k–$71k/yr
Skills
pythondata sciencedata analyticslinux command-linesoftware engineering fundamentalsgitunit testinghigh-performance computing (hpc)workload managerstime-series datasystem logslarge language model (llm) apissupercomputing architecture
Job Description
Summary: The National Laboratory of the Rockies (NLR) is seeking a Graduate Summer Intern for its Advanced Computing program in High-Performance Computing (HPC) Operational Analytics. The role involves assisting in building data foundations and AI-driven intelligence infrastructure, focusing on optimizing the operations of the Kestrel supercomputer.
Responsibilities:
- Develop scalable Python pipelines to ingest and analyze diverse HPC telemetry
- Build and evaluate machine learning models to characterize complex HPC workloads
- Assist in developing automated, agentic AI workflows
- Process and package operational data
- Write robust, test-driven code and document data schemas within a collaborative Git repository
Required Qualifications:
- Minimum of a 3.0 cumulative grade point average
- Undergraduate: Must be enrolled as a full-time student in a bachelor's degree program from an accredited institution
- Post Undergraduate: Earned a bachelor's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate: Must be enrolled as a full-time student in a master's degree program from an accredited institution
- Post Graduate: Earned a master's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate + PhD: Completed master's degree and enrolled as PhD student from an accredited institution
- Proficient in Python, specifically the data science and analytics stack
- Comfortable operating in a Linux command-line environment
- Software engineering fundamentals, including Git and unit testing
Preferred Qualifications:
- Experience running code on HPC systems or using workload managers (e.g., Slurm)
- Experience handling complex, large-scale data formats (time-series telemetry, system logs, or text-based data)
- Familiarity with integrating Large Language Model (LLM) APIs into software workflows or building agentic systems
- Basic understanding of supercomputing architecture or facility operations
Required Skills: Python, Data Science, Data Analytics, Linux Command-Line, Software Engineering Fundamentals, Git, Unit Testing, High-Performance Computing (HPC), Workload Managers, Time-Series Data, System Logs, Large Language Model (LLM) APIs, Supercomputing Architecture
Benefits: Medical, dental, and vision insurance, 403(b) Employee Savings Plan with employer match, Sick leave (where required by law), Performance-, merit-, and achievement- based awards that include a monetary component, Relocation expense reimbursement
Benefits
Medical, dental, and vision insurance
403(b) Employee Savings Plan with employer match
Sick leave (where required by law)
Performance-, merit-, and achievement- based awards that include a monetary component
Relocation expense reimbursement