Regeneron→
Data Engineer, Analytical & Biological Mass… at Regeneron · TARRYTOWN
Entry LevelOn-siteFull-timeTARRYTOWN$94k–$153k/yr
Skills
data engineeringpythonsqlapi developmentetl/elt pipelinesrelational database designcloud platforms - awsdata validationloggingerror handlingversion controlsoftware testing
Job Description
Summary: Regeneron Pharmaceuticals is seeking a Data Engineer to join their Data Enablement and Analytics team within the Analytical & Biological Mass Spectrometry group. The role involves designing, developing, and maintaining data platforms to support automated data pipelines for LC-MS analytics, enhancing the efficiency of biotherapeutic drug development programs.
Responsibilities:
- Design, develop, and maintain scalable, automated data pipelines that integrate LC-MS analytical workflows from raw data acquisition through cloud storage, processing, structured archival, and visualization
- Design, build, and optimize ETL/ELT workflows, data integrations, and APIs to enable seamless interoperability across heterogeneous systems (e.g., LC-MS instruments, LIMS, SDMS, data processing software, and enterprise data stores)
- Partner with IT and DEA teams to design, deploy, and manage data infrastructure (e.g., data lakes, data warehouses) for robust data ingestion, processing, and storage
- Drive platform reliability and performance through proactive monitoring, observability practices, and continuous improvement
- Contribute to ABMS data strategy, architecture, and governance by establishing standards for data quality, code quality, compliance, accessibility, and platform reliability; support adoption through documentation, training, and best practice guidance
- Stay current with emerging technologies and evaluate innovative approaches in data engineering, scientific informatics, and operational analytics
Required Qualifications:
- Bachelor's or Master's degree in Computer Science, Data Engineering, Software Engineering, Data Science, Bioinformatics, Computational Biology, Computer Engineering, Information Systems, or a related quantitative discipline
- 0–5 years of hands-on experience in data engineering, scientific data infrastructure, or a closely related technical role
- Proficiency in Python and SQL; demonstrated ability to write production-quality code for data pipelines, data/file parsing, and automated workflows, apply software engineering best-practices such as version control, testing, documentation, and code review
- Experience developing APIs, ETL/ELT pipelines, or data access layers to connect backend systems with analytical and visualization tools
- Experience with relational database design and building structured data stores from semi-structured or unstructured scientific data
- Experience with cloud platforms (AWS preferred), including data storage and compute services
- Understanding of data validation, logging, and error‑handling concepts in production data pipelines
- Strong communication skills, with the ability to translate scientific data requirements into well-documented, maintainable technical solutions and collaborate effectively across scientific, analytical, and IT teams
Preferred Qualifications:
- Prior experience in biopharmaceutical, biotech, or life sciences industry, particularly within an analytical laboratory environment
- Familiarity with mass spectrometry raw data and processed data formats
- Familiarity with LC-MS software ecosystems (e.g., Skyline, LabKey Panorama, Protein Metrics Byosphere, Genedata Expressionist, or Waters UNIFI/Empower)
- Familiarity with Laboratory Information Management System (LIMS), Scientific Data Management System (SDMS), and Electronic Lab Notebooks (ELN) platforms (e.g., Benchling, NuGenesis, and IDBS), including API integration or workflow configuration
- Familiarity with workflow orchestration tools (e.g., Nextflow or similar)
- Familiarity with shell scripting (e.g., Bash) and configuration formats such as JSON
- Experience with containerization technologies such as Docker for packaging and deploying data pipelines
- Experience building data connectors or API integrations for visualization platforms (e.g., Power BI, Spotfire, and Tableau)
Required Skills: Data Engineering, Python, SQL, API Development, ETL/ELT Pipelines, Relational Database Design, Cloud Platforms - AWS, Data Validation, Logging, Error Handling, Version Control, Software Testing
Benefits: Health and wellness programs (including medical, dental, vision, life, and disability insurance), Fitness centers, 401(k) company match, Family support benefits, Equity awards, Annual bonuses, Paid time off, Paid leaves (e.g., military and parental leave)
Benefits
Health and wellness programs (including medical, dental, vision, life, and disability insurance)
Fitness centers
401(k) company match
Family support benefits
Equity awards
Annual bonuses
Paid time off
Paid leaves (e.g., military and parental leave)