Pathway→
AI Benchmark & Datasets Engineer/ Researcher… at Pathway · Palo Alto
InternshipHybridFull-timePalo Alto, CA
Skills
ml/llm evaluationdata sciencebenchmarkingtechnical product roleshigh-quality datareproducible experimentsenglish fluency
Job Description
Summary: Pathway is an innovative company focused on building advanced AI models that surpass traditional transformer architectures. They are seeking an AI Benchmark & Datasets Engineer/Researcher Intern to support benchmarking processes for model evaluation, which will directly influence the development and market presentation of their AI solutions.
Responsibilities:
- Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets
- Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap
- Run benchmarks with baseline models to validate setup, uncover edge cases, and de-risk R&D runs
- Hand off "benchmark-ready" packages to R&D (specs, data, evaluation scripts, expected metrics, constraints)
- Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats that GTM and R&D can both use
- Track and organize benchmark results, model leaderboards, and "what good looks like" for different customers and scenarios
- Contribute to demos and public-facing proof points based on benchmark outcomes
Required Qualifications:
- You are expected to meet at least one of the following criteria:
- You were an ICPC World Finalist, or an IOI, IMO, IOAI or IPhO medalist in High School
- You have published a research paper at an A-rated o A*-rated venue (according to ICORE)
- You have completed coding projects - ideally with a GitHub repository showcasing previous work
- You were an intern at a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA)
- You can get a warm recommendation from your university faculty member
- Have experience with ML/LLM evaluation, data science, or technical product roles, ideally around benchmarks or experimentation
- Are comfortable reading papers, leaderboards, and Github repos, and turning them into clear, repeatable benchmark specs
- Can talk comfortably with both engineers and customers, and translate between technical detail and business value
- Care about high‑quality data, reproducible experiments, and crisp documentation
- Are respectful of others
- Are fluent in English
Required Skills: ML/LLM evaluation, data science, benchmarking
Important Skills: technical product roles
Nice-to-Have Skills: high-quality data, reproducible experiments, English fluency
Benefits: Join an intellectually stimulating work environment., Be a pioneer: you get to work with a new type of 'Live AI' challenges., Be part of one of an early-stage AI startup that believes in impactful research and foundational changes.
Benefits
Join an intellectually stimulating work environment.
Be a pioneer: you get to work with a new type of 'Live AI' challenges.
Be part of one of an early-stage AI startup that believes in impactful research and foundational changes.