NVIDIA→
Software Engineer, AI Inference Systems - New College… at NVIDIA · US
Entry LevelOn-siteFull-timeUS, CA, Santa Clara$108k–$178k/yr
Skills
pythonc++gorustalgorithmsdata structuresoperating systemscomputer architectureparallel programmingdistributed systemsdeep learning theoriesperformance engineering in ml frameworksml frameworks - pytorchinference engines - vllminference engines - sglanggpu programmingcudagpu memory hierarchygpu streamsncclprofiling tools - nsight systemsprofiling tools - nsight computecontainersdockerkubernetesslurmlinux namespaceslinux cgroupsdebugging
Job Description
Summary: NVIDIA is a leader in AI research and development, seeking highly skilled software engineers to build AI inference systems. The role involves architecting high-performance inference stacks, optimizing GPU kernels, and collaborating across teams to advance accelerated computing for AI.
Responsibilities:
- Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation
- Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization; build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization
- Define and build inference benchmarking methodologies and tools; contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite
- Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds
- Conduct and publish original research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products
Required Qualifications:
- Bachelor's, Master's, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE) (or equivalent experience)
- Strong programming skills in Python and C/C++; experience with Go or Rust is a plus; solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, distributed systems, deep learning theories
- Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang)
- Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute)
- Experience with containers and orchestration (Docker, Kubernetes, Slurm); familiarity with Linux namespaces and cgroups
- Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting
Preferred Qualifications:
- Experience building and optimizing LLM inference engines (e.g., vLLM, SGLang)
- Hands-on work with ML compilers and DSLs (e.g., Triton, TorchDynamo/Inductor, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores)
- Experience contributing to containerization/virtualization technologies such as containerd/CRI-O/CRIU
- Experience with cloud platforms (AWS/GCP/Azure), infrastructure as code, CI/CD, and production observability
- Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts
Required Skills: Python, C++, Go, Rust, Algorithms, Data structures, Operating systems, Computer architecture, Parallel programming, Distributed systems, Deep learning theories, Performance engineering in ML frameworks, ML frameworks - PyTorch, Inference engines - vLLM, Inference engines - SGLang, GPU programming, CUDA, GPU memory hierarchy, GPU streams, NCCL, Profiling tools - Nsight Systems, Profiling tools - Nsight Compute, Containers, Docker, Kubernetes, Slurm, Linux namespaces, Linux cgroups, Debugging
Benefits: Equity, Benefits
Benefits
Equity
Benefits