NVIDIA→
Deep Learning Software Engineer, FlashInfer - New… at NVIDIA · US
Entry LevelOn-siteFull-timeUS, CA, Santa Clara$124k–$196k/yr
Skills
deep learning frameworksinference enginespythonc++domain specific compilersgpu kernel developmentcuda c++machine learning compilersperformance optimization
Job Description
Summary: NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. They are looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack, focusing on innovative AI systems software to accelerate for AI inference.
Responsibilities:
- Innovating and developing new AI systems technologies for efficient inference
- Designing, implementing, and optimizing kernels for high impact AI workloads
- Designing and implementing extensible abstractions for LLM serving engines
- Building efficient just-in-time domain specific compilers and runtimes
- Collaborating closely with other engineers at NVIDIA across deep learning frameworks, libraries, kernels, and GPU arch teams
- Contributing to open source communities like FlashInfer, vLLM, and SGLang
Required Qualifications:
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field (or equivalent experience); PhD are preferred
- Strong experience in developing or using deep learning frameworks (e.g. PyTorch, JAX, TensorFlow, ONNX, etc) and ideally inference engines and runtimes such as vLLM, SGLang, and MLC
- Strong Python and C/C++ programming skills
Preferred Qualifications:
- Background in domain specific compiler and library solutions for LLM inference and training (e.g. FlashInfer, Flash Attention)
- Expertise in inference engines like vLLM and SGLang
- Expertise in machine learning compilers (e.g. Apache TVM, MLIR)
- Strong experience in GPU kernel development and performance optimizations (especially using CUDA C/C++, cuTile, Triton, or similar)
- Open source project ownership or contributions
Required Skills: Deep Learning Frameworks, Inference Engines, Python, C++, Domain Specific Compilers, GPU Kernel Development, CUDA C++, Machine Learning Compilers, Performance Optimization
Benefits: Equity, Benefits
Benefits
Equity
Benefits