Workday→
Machine Learning Engineer III at Workday in Pleasanton, CA
Entry LevelOn-siteFull-timePleasanton, CA
Job Description
Machine Learning Engineer III
Location: USA, CA, Pleasanton
Your work days are brighter here.
We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.
About the Team
This is a very exciting opening in the AI Platform team in our Information Retrieval and Agent Evaluation team. We believe if you do what you love, you’ll love what you do. There’s a lot to love at Workday. We are part of a global, high-growth technology company and our team has the opportunity to develop the next generation of Workday’s groundbreaking collaborative products supporting a customer base of more than 31 million strong. Over 65% of the Fortune 500 are Workday customers.
The Agent Evaluation Platform project is the "Ground Truth" engine for Workday’s AI transformation and we have an ambitious roadmap. As Workday infuses AI Agents into every facet of our enterprise suite, our team provides the critical infrastructure and algorithms needed to prove they work—and make them better. We build the platform that enables agent engineering teams to be empowered with rigorous, data-driven optimization, evaluation and validation of their agents.
The AI Platform Information Retrieval products are at the heart of Workday’s intelligence layer. We bridge the gap between human language, search, and enterprise data, including reasoning over knowledge. Our products utilize advanced semantic search to navigate Workday’s massive data model, as well as turning natural language questions into precise SQL and Python executions.
Workday’s AI Platform organization is bringing “AI first” products to life at every step of the Workday product offering. We’re looking for highly creative, results-focused, and deeply skilled Machine Learning Engineers/scientists to work with us on a range of these challenges.
Why Workday?
1. The Data: Work with exclusive, high-integrity enterprise datasets that most researchers never see. You'll be working at the absolute frontier of Agentic AI - "how do we validate, scale, and optimize an agent" and "how do we extract the correct data for agents".
2. The Scale: Your code will empower the world’s largest companies to make data-driven decisions. You are the gatekeeper of quality for products reaching 31 million users.
3. The Culture: A "people-first" environment that balances high-intensity innovation with sustainable work-life integration.
About the Role
We are looking for a highly skilled and pragmatic Machine Learning Engineer to work with us on the applied research, development, deployment, and optimization of advanced ML systems, Agent Evaluation, Search, Information Retrieval (IR) and Semantic Parsing products. You will be a crucial driver in embedding cutting-edge AI agent technology, directly into Workday products, moving quickly from deep applied research to robust production features. You will use Workday’s vast computing resources on rich, exclusive datasets to deliver value that transforms the way our customers make decisions and run their businesses. We will challenge you to apply your best creative thinking, analysis, problem-solving, and technical abilities to make an impact on thousands of enterprises and millions of users. Sound like your kind of challenge?
In this role, you would:
Agent Optimization (Meta-ML): Develop algorithms for automated node-level optimization within agent graphs - determining the best LLM model and prompt configuration for every step of a workflow. Work on enabling life-long agent learning.
Architect Evaluation Pipelines: Design and scale offline evaluation systems using Kubeflow on the cloud to handle massive, distributed test suites for LLM agents. Build recommender systems for engineering teams to drive optimal evaluation for their agents. Work on failure attribution as an outcome of evaluation.
Enable Online Evaluation and Observability: Implement real-time evaluation and A/B testing frameworks for live agents, enabling teams to experiment with new architectures in production safely.
Advanced Information Retrieval: Build and refine hybrid search systems (Vector, Keyword, Agentic, and Multi-step Reasoning) that navigate Workday’s proprietary data objects with high precision and recall. Apply deep learning techniques to enhance recommender systems, ranking models, and information retrieval systems, driving relevance and personalization across core Workday applications.
Semantic Parsing & Code Gen: Design and optimize LLM, model and agent based products for Text-to-SQL and Text-to-Python, focusing on schema grounding, few-shot prompting, and fine-tuning for structured output.
Evaluation Platform Development: Establish rigorous and scalable methods for evaluating AI Agents, and Information Retrieval products, LLMs, and other ML model accuracy and performance, including developing metrics for quality, safety, latency, and user experience.
AI Agent Engineering: Design, build, and deploy sophisticated AI agents (e.g. orchestration agents, reasoning agents, planning and execution node agents, tool selection agents, autonomous workflow agents, conversational interface agents, swarm agents) that interact seamlessly with enterprise data. Work on continuous learning for the agents.
Prompt Engineering & Optimization: Develop, test, and maintain advanced prompt engineering and prompt optimization strategies and guardrails to ensure LLM-powered features are accurate, safe, and reliable at scale.
Production and MLOps: Own the entire ML lifecycle, ensuring high-quality, scalable deployment, monitoring, and continuous improvement of all models and agents in production environments.
Own exploration, design and execution of advanced ML models, algorithms and frameworks that deliver value to our users.
Collaborate with other ML engineers, software engineers, product managers, and across teams to deliver your products through Workday end user applications.
Be given autonomy and ownership over your work, but with the support of the entire organization.
Keep abreast of the latest advancements in ML/AI, Information Retrieval, Agentic AI, Generative AI, NLP research, techniques and tools.
Have extraordinary opportunities for career growth and learning in a fast-growing, forward-looking company
About You
Basic Qualifications
3+ years of professional experience as a Machine Learning Engineer, focusing on researching, developing, building, training, and deploying deep learning, Information Retrieval, NLP solutions, and generative/agentic AI systems into production.
Proven, hands-on experience building and launching Generative AI products powered by long context LLMs, specifically applied to structured data tasks e.g. Text to SQL.
Demonstrated experience in building and evaluating AI agents, including familiarity with agent frameworks, RAG architectures, and agent evaluation frameworks (e.g., DeepEval, RAGAS, or custom internal systems).
Programming Mastery: Expert-level Python skills, specifically for building modular libraries and frameworks that other engineers will use, and including experience with AI coding tools. Experience with topics relating to multi-threading, api design, matrix processing, runtime memory design, and asynchronous call patterns.
Solid understanding and experience with MLOps, scalability, and cloud services (e.g., AWS, GCP, Azure), containerization technologies (e.g. Docker), Kubernetes, and large-scale ML systems.
Other Qualifications
A relevant advanced degree (Master’s or Ph.D.) in Computer Science, Machine Learning, AI, Computer/Software engineering or a related quantitative programming field.
Proficiency in modern ML and deep learning frameworks (e.g. PyTorch, TensorFlow, Huggingface), and agentic frameworks (e.g. LangChain/LangGraph/LangSmith).
Proven theoretical and practical understanding of statistical analysis and machine learning algorithms, natural language processing, optimization, recommender systems, especially for supervised, unsupervised and self-supervised methods. Solid understanding of A/B testing, statistical significance, and how to design experiments that differentiate signal from noise in non-deterministic LLM outputs.
Expertise in prompt engineering, prompt optimization, and developing robust strategies for integrating LLMs into user-facing products.
Systems Design: Strong understanding of how to build scalable "Agent-in-the-loop" systems, including error handling and state management.
Experience researching, developing and deploying production-grade recommender systems, information retrieval systems and/or ranking models, e.g. vector databases, embedding models, and RAG (Retrieval-Augmented Generation) architectures.
Expertise in language model fine-tuning techniques (e.g., parameter-efficient fine-tuning, domain adaptation) and building models with mid and large model architectures (e.g., BERT family, as well as LLMs).
Take ownership for finding creative algorithmic and system design solutions that move projects, work-streams and products forward. Have determination to turn ideas into reality and improve user experience.
Optimization-Focused: Interest in optimization techniques, prompt optimization techniques (e.g. DSPy), and cost-benefit analysis of different LLM architectures.
Data Mastery: Proficiency in PySpark, Pandas, and SQL. Experience with large-scale data processing, and data ingestion pipelines.
Experience with advanced techniques such as reinforcement learning, imitation learning, graph neural networks, and multi-modal models.
Schema Expertise: Familiarity with Knowledge Graphs or complex relational data modeling is a huge plus.
Evaluation Obsession: You don't trust a model until you've built a statistically sound way to break it. Experience with A/B testing, online evaluation experimentation, and "Golden Dataset" curation.
Standout leader, strong communication skills, with experience working across functions and teams, and working on ambiguous problems.
Bonus points for machine learning related research publications.
Resilience to obstacles, and the ability to lead the solving of problems independently.
The annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please click here.
Primary Location: USA.CA.Pleasanton
Primary Location Base Pay Range: $160,000 USD - $240,000 USD
Additional US Location(s) Base Pay Range: $136,200 USD - $240,000 USD
With Flex Work, we’re combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.
Pursuant to applicable Fair Chance law, Workday will consider for employment qualified applicants with arrest and conviction records.
Workday is an Equal Opportunity Employer including individuals with disabilities and protected veterans.
accommodations@workday.com.
Are you being referred to one of our roles? If so, ask your connection at Workday about our Employee Referral process!
At Workday, we value our candidates’ privacy and data security. Workday will never ask candidates to apply to jobs through websites that are not Workday Careers.
Please be aware of sites that may ask for you to input your data in connection with a job posting that appears to be from Workday but is not.
In addition, Workday will never ask candidates to pay a recruiting fee, or pay for consulting or coaching services, in order to apply for a job at Workday.
Additional Locations: USA, CA, Santa Clara, USA, CA, San Francisco