Apple→
Machine Learning Engineer at Apple in Sunnyvale, CA
Job Description
Imagine what you could do here. At Apple, great ideas have a way of becoming great products, services, and customer experiences very quickly. Bring passion and dedication to your job and there’s no telling what you could accomplish.
Do you want to make Siri and Apple products more intelligent for our users? The Machine Learning Platform Technology team is building groundbreaking technology for search, natural language processing, artificial intelligence and machine learning. Our infrastructure and research for data curation form the backbone of Apple Intelligence. It powers the largest Apple foundation models on servers and a wide gamut of services at Apple including Apple Search, Apple Music, AppleTV, AppStore, iMessages, Photos & Camera, Spotlight, Safari, Siri and upcoming ever exciting Apple products serving millions of queries every day with incredible low latencies, drawing every ounce of compute from our hardware.
As part of this group, you will work with one of the most exciting high performance computing environments, with petabytes of data, millions of queries per second, and have an opportunity to imagine and build products that delight our customers every single day. You will have a chance to work on optimizing billions of parameter language and vision and speech models using state of the art technologies and make it run at the scale of Apple.
We design, build, and maintain large-scale ML systems that make petabytes of data easy to process and query while upholding Apple's rigorous privacy standards, for training the next generation of Apple Foundation Models.
Contribute to pre-training data curation for LLM development across all domains and modalities (eg: web, STEM, code, multilingual, image, audio, video)
Design, implement, and deploy scalable ML models for information extraction, data selection, and synthetic data generation over trillions of unstructured records
Design, execute, and analyze scientific experiments to advance our understanding of large language models
Independently lead small-scale research projects while contributing to cross-functional, larger-scale research initiatives
Accelerate workflows by building high-quality data tooling that improves data velocity, reliability, and scalability
Apply expertise in one or more of the following areas: multimodal generation and perception (text, image, video, or audio), OCR, data scaling laws, or data mixing
4+ years of experience with large language models (LLMs), large multimodal models (LMMs), computer vision, or related AI/ML technologies
Experience developing and evaluating machine learning models, with a strong understanding of data and model quality
Strong programming skills and hands-on experience using one or more deep learning frameworks, such as PyTorch, TensorFlow, or JAX
Experience building large-scale machine learning systems and working with distributed data processing frameworks
Strong problem-solving skills with a results-oriented mindset
Experience leading rapid prototyping and proof-of-concept development for AI/ML applications
Excellent communication skills, with the ability to work independently and collaborate effectively with cross-functional teams
Master's degree or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field, or equivalent industry experience
Experience developing or advancing state-of-the-art LLMs and/or LMMs
Experience with modern deep learning architectures, including Transformers, mixture-of-experts (MoE), and multimodal models
Experience deploying and scaling machine learning and deep learning models, including VLMs, for production-scale inference
Publications at leading AI conferences (e.g., NeurIPS, ICML, ICLR, ACL, CVPR, ICCV, AAAI, KDD) and/or demonstrated significant industry impact in AI