Roche→
Data Delivery Specialist (Clinical & Multimodal… at Roche · South…
Entry LevelHybridFull-timeSouth San Francisco$110k–$205k/yr
Skills
clinical data structurespythonpandassqlbashcsvjsonparquetfastqvcfdicomaws s3google cloud storagejupyter notebooksairflowcdisc sdtmcdisc adamomopfhirmetadata standardsontologiescontrolled vocabulariesai/ml data workflowsmultimodal datasets
Job Description
Summary: Roche is a company dedicated to advancing science and healthcare access. They are seeking an Associate Data Delivery Specialist to support the integration and delivery of clinically anchored, multimodal datasets across sequencing, imaging, and proteomics domains, ensuring high-quality datasets for analytics and research.
Responsibilities:
- Assist in the ingestion, validation, and harmonization of clinical datasets, including patient demographics, longitudinal records, and outcomes data. Apply data standards and controlled vocabularies to improve consistency and usability
- Support the integration of clinical data with sequencing, imaging, and proteomics datasets to create coherent, analysis-ready multimodal datasets
- Prepare, validate, and document datasets for delivery to internal stakeholders across research, bioinformatics, and data science teams. Perform quality control checks and support issue resolution in data workflows
- Assist in structuring datasets and metadata to support downstream analytics and machine learning use cases. Contribute to early-stage AI-enabled data curation and harmonization efforts
- Work closely with data engineers, data scientists, and domain experts to support data integration efforts and ensure alignment with scientific and technical requirements
Required Qualifications:
- Bachelor's or Master's degree in Bioinformatics, Data Science, Biomedical Engineering, Computer Science, Clinical Sciences, or a related field and 0–2 years of experience working with clinical, biomedical, or scientific data
- Foundational knowledge of clinical data structures, including patient-level and longitudinal datasets
- Detail-oriented with a strong focus on data quality and consistency and are motivated to learn and grow in a data-intensive, scientific environment
- Technical skills for: Programming: Python (Pandas), SQL; familiarity with Bash is a plus
- Data Formats: Experience with structured data (CSV, JSON, Parquet); exposure to scientific formats (e.g., FASTQ, VCF, DICOM) is a plus
- Data Platforms: Exposure to cloud storage environments such as AWS S3 or Google Cloud Storage
- Tools: Familiarity with Jupyter notebooks and basic workflow tools (e.g., Airflow) is beneficial
Preferred Qualifications:
- Exposure to clinical data standards such as CDISC (SDTM/ADaM), OMOP, or FHIR
- Experience or coursework involving multimodal datasets (e.g., clinical + omics or imaging)
- Familiarity with metadata standards, ontologies, or controlled vocabularies
- Basic understanding of AI/ML data requirements and workflows
- Interest in applying data to translational research and drug discovery
Required Skills: Clinical data structures, Python, Pandas, SQL, Bash, CSV, JSON, Parquet, FASTQ, VCF, DICOM, AWS S3, Google Cloud Storage, Jupyter notebooks, Airflow, CDISC SDTM, CDISC ADaM, OMOP, FHIR, Metadata standards, Ontologies, Controlled vocabularies, AI/ML data workflows, Multimodal datasets
Benefits: A discretionary annual bonus may be available based on individual and Company performance., This position also qualifies for the benefits detailed at the link provided below.
Benefits
A discretionary annual bonus may be available based on individual and Company performance.
This position also qualifies for the benefits detailed at the link provided below.