SpaceX→
High Performance Computing (HPC) Systems… at SpaceX · Hawthorne
Entry LevelOn-siteFull-timeHawthorne, CA$125k–$145k/yr
Skills
hpc clusterslinuxbashpythonslurmpbslsfprometheusgrafananagioscfdfeapytorchtensorflowgpu computingcudadockerpodmansingularitypuppetansibleenterprise networkingvirtualizationsecurity technologiesself-motivator
Job Description
Summary: SpaceX is a company dedicated to enabling human life on Mars through advanced technologies. They are seeking an HPC Systems Engineer to manage HPC clusters, provide application support, and integrate Linux-based compute clusters within a fast-paced engineering environment.
Responsibilities:
- Administer and manage HPC clusters, storage systems, and high-speed networks
- Provide application support to SpaceX employees across engineering disciplines
- Install and integrate Linux-based compute clusters
- Write instructional documentation and convey highly technical ideas in non-technical terms
Required Qualifications:
- 1+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
- Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree
- Experience with Linux
- Must be willing to work extended hours and weekends as needed
- To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. 1157, or (iv) Asylee under 8 U.S.C. 1158, or be eligible to obtain the required authorizations from the U.S. Department of State
Preferred Qualifications:
- 3+ years of professional experience building, deploying and troubleshooting Linux systems
- Experience with a scripting language (Bash, Python) to automate and solve reoccurring tasks
- Experience building, deploying and troubleshooting HPC clusters
- Familiarity with cluster resource managers (Slurm, PBS, LSF)
- Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
- Familiarity with scientific and engineering computing (CFD, FEA)
- Familiarity with ML frameworks (PyTorch, Tensorflow)
- Familiarity with GPU usage in a compute cluster and Cuda
- Experience with containers (Docker, Podman, Singularity)
- Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
- Comfortable working with mission critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
- Eligibility for access to classified material up to TS/SCI with Polygraph
Required Skills: HPC clusters, Linux, Bash, Python, Slurm, PBS, LSF, Prometheus, Grafana, Nagios, CFD, FEA, PyTorch, Tensorflow, GPU computing, Cuda, Docker, Podman, Singularity, Puppet, Ansible, Enterprise networking, Virtualization, Security technologies, Self-motivator
Benefits: Long-term incentives, in the form of company stock, stock options, or long-term cash awards, Potential discretionary bonuses, Ability to purchase additional stock at a discount through an Employee Stock Purchase Plan, Access to comprehensive medical, vision, and dental coverage, Access to a 401(k) retirement plan, Short and long-term disability insurance, Life insurance, Paid parental leave, Various other discounts and perks, 3 weeks of paid vacation, Eligible for 10 or more paid holidays per year, Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law
Benefits
Long-term incentives, in the form of company stock, stock options, or long-term cash awards
Potential discretionary bonuses
Ability to purchase additional stock at a discount through an Employee Stock Purchase Plan
Access to comprehensive medical, vision, and dental coverage
Access to a 401(k) retirement plan
Short and long-term disability insurance
Life insurance
Paid parental leave
Various other discounts and perks
3 weeks of paid vacation
Eligible for 10 or more paid holidays per year
Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law