RefinedScience→
Bioinformatics Machine Learning Intern at RefinedScience in Remote
InternshipRemoteFull-timeRemote$71k–$79k/yr
Skills
single-cell omicsmachine learning expertiseprogramming proficiencystatistical foundationdeep learning frameworksbioinformatics experiencelinux/unix cliversion control (git)containerization (docker)cloud computing
Job Description
Summary: RefinedScience is dedicated to advancing care through innovative science and data integration. They are seeking a highly motivated Bioinformatics Machine Learning Intern to contribute to projects in single-cell biology and precision medicine by applying machine learning techniques to biological datasets.
Responsibilities:
- Analyze single-cell and multiomics datasets to extract biological insights supporting precision medicine and drug development programs
- Apply and evaluate machine learning and deep learning approaches to single-cell data for tasks such as cell type classification, biomarker discovery, and patient stratification
- Explore and prototype generative AI and LLM-based approaches to accelerate biological data interpretation and scientific workflows
- Collaborate with scientists, clinicians, and data scientists to design and execute data-driven research projects
- Document and optimize computational workflows following reproducible research best practices
- Present findings through technical reports, visualizations, and presentations to cross-functional teams
Required Qualifications:
- Current Ph.D. candidate in Bioinformatics, Computational Biology, Computer Science, Biostatistics, or a related quantitative field
- Single-cell omics experience: Demonstrated ability to process, analyze, and interpret single-cell data (scRNA-seq, scATAC-seq, CITE-seq, or spatial transcriptomics) using frameworks such as Scanpy/scverse, Seurat, or Bioconductor
- Machine learning expertise: Applied experience developing and evaluating ML/deep learning models on biological data, including neural network architectures (GNNs, transformers, autoencoders), model selection and benchmarking, and integration of ML approaches into analytical workflows
- Programming proficiency: Python and/or R for data analysis, statistical modeling, and visualization
- Statistical foundation: Understanding of statistical methods for biological data (hypothesis testing, differential expression, multiple testing correction, clustering)
- Strong problem-solving skills and ability to communicate complex insights effectively
Preferred Qualifications:
- Experience with deep learning frameworks (PyTorch, TensorFlow, JAX)
- Familiarity with graph neural networks, attention mechanisms, or transformer architectures applied to biological data
- Experience with ML experiment tracking and reproducibility (MLflow, Weights & Biases)
- Exposure to representation learning, variational autoencoders, or contrastive learning methods
- Familiarity with scikit-learn, XGBoost, or similar ML libraries
- Interest in or experience with LLMs, RAG systems, or agentic AI tooling
- Experience with multimodal single-cell integration (Seurat WNN, scvi-tools/MultiVI/totalVI, Muon)
- Familiarity with spatial transcriptomics analysis (Squidpy, cell2location, nf-core/spatialvi)
- Experience with cell-cell communication inference (CellChat, NicheNet, LIANA)
- Knowledge of drug-gene interaction resources (CMap/LINCS, OpenTargets, ChEMBL)
- Familiarity with Linux/Unix CLI and version control (Git/GitHub)
- Experience with containerization (Docker, Singularity) and environment management (conda, venv)
- Exposure to cloud computing platforms (GCP preferred)
- Familiarity with workflow managers (Nextflow, Snakemake)
- Adherence to best-practices for conduct reproducible computational research
Required Skills: Single-cell omics, Machine learning expertise, Programming proficiency
Important Skills: Statistical foundation, Deep learning frameworks, Bioinformatics experience
Nice-to-Have Skills: Linux/Unix CLI, Version control (Git), Containerization (Docker), Cloud computing