Cambia Health Solutions→
AI Scientist Intern at Cambia Health Solutions in Burlington, WA
InternshipHybridFull-timeBurlington, WA$82k–$82k/yr
Skills
data sciencemachine learningdeep learningstatistical modelinghealthcare payer domainalgorithmic designpython programmingmachine learning pipelinesclassic machine learning algorithmssupervised learningsemi-supervised learningunsupervised learningreinforcement learningregressionclassificationtime series modelingtransfer learningobjective functionsregularizationoverfittinghyperparameter tuningmodel evaluation metricsfeature importancescikit-learnxgboosttensorflowpytorchnatural language processingtokenizationpart-of-speech taggingnamed entity recognitionsentiment analysismachine translationpre-trained nlp modelsbertgptspacytext preprocessingfeature extractionembedding methodsnltkhugging face transformerstext classificationlanguage generationinformation extractionlarge language modelsprompt engineeringfine-tuning techniquesretrieval-augmented generationgenerative model evaluationalignment techniquesmultimodal generative modelsresponsible aidata analysisresearch designdata visualizationsqlalgorithmsdata structuresinterest in healthcare industry
Job Description
Summary: Cambia Health Solutions is a purpose-driven company in the healthcare sector, seeking an AI Scientist Intern for a 12-week full-time position starting in May or June 2026. The intern will work under the mentorship of an AI Scientist, engaging in the design, development, and implementation of data-driven solutions using machine learning and statistical modeling techniques to address business problems in healthcare.
Responsibilities:
- Researches, designs, develops, and implements data-driven models and algorithms using machine learning, deep learning, statistical, and other mathematical modeling techniques
- Trains and tests models and develops algorithms to solve business problems
- Adheres to standard best-practices and establishes principled experimental frameworks for developing data-driven models
- Develops models and performs experiments and analyses that are replicable by others
- Uses open-source packages when appropriate to facilitate model development
- Identifies, measures, analyzes, and visualizes drivers to explain model performance (e.g., feature importance, interpretability, bias and error analysis), both offline (in the development phase) and online (in production)
- Uses appropriate metrics and quantified outcomes to drive model and algorithm improvements
- Analyzes, diagnoses, and resolves bugs in production machine learning models and systems
- Evaluates model/use case feasibility by quickly generating prototypes
- Takes models from prototype stage and improves performance as needed
- Writes clean, well-commented, tested, version-controlled, and maintainable python code
- Collaborates with team members and Cambia business partners
- Actively participates in group meetings and discussions
- Communicates effectively both orally and in writing with both technical and non-technical audiences
- Keeps current with the state of the art in machine learning and AI and its application to healthcare
- Keeps current with evolving commercial and open-source tools, techniques, and brings these practices to projects
- Over time develops familiarity and insight with various subdomains of healthcare data
Required Qualifications:
- Demonstrated knowledge of data science, machine learning, and modeling
- Ability to use well-understood techniques and existing patterns to build, analyze, deploy, and maintain models
- Effective in time and task management
- Able to develop productive working relationships with colleagues and business partners
- Strong interest in the healthcare industry
- Ability to read, understand, and apply the latest research to enhance our products where possible
- Strong mathematical foundation and theoretical grasp of the concepts underlying machine learning, optimization, etc
- Demonstrated understanding of how to structure simple machine learning pipelines
- Classic ML algorithms (e.g., linear and logistic regression, decision and boosted trees, SVM, collaborative filtering, ranking)
- Approaches (e.g., supervised, semi-supervised, unsupervised, reinforcement learning, regression, classification, time series modeling, transfer learning)
- Foundational ML concepts such as objective functions, regularization and overfitting
- Data partitions (train/dev/test) and model development
- Hyperparameter tuning and grid search
- Evaluation concepts (metrics, feature importance, etc.)
- Familiarity with standard python packages (scikit-learn, XGBoost, TensorFlow, PyTorch, etc.)
- Familiarity with structure of machine learning pipelines
- Experience with natural language processing (NLP) techniques such as tokenization, part-of-speech tagging, named entity recognition, sentiment analysis, and machine translation
- Familiarity with pre-trained models and frameworks like BERT, GPT, and spaCy for NLP tasks
- Understanding of text preprocessing, feature extraction, and embedding methods
- Experience with NLP libraries and tools (NLTK, Hugging Face Transformers, etc.)
- Capable of building and refining NLP models for tasks such as text classification, language generation, and information extraction
- Awareness of the ethical considerations and challenges in deploying NLP models
- Large Language Models (LLMs) and their capabilities (e.g, in-context learning, few-shot learning, zero-shot learning)
- Prompt engineering techniques and best practices
- Fine-tuning approaches (e.g, full fine-tuning, parameter-efficient methods like LoRA, QLoRA)
- Retrieval-Augmented Generation (RAG) and knowledge integration
- Evaluation methods for generative models (e.g, perplexity, BLEU, ROUGE, human evaluation, LLM as a Judge)
- Alignment techniques (e.g, RLHF, constitutional AI, red-teaming)
- Multimodal generative models (text-to-image, text-to-video, multimodal understanding)
- Responsible AI considerations specific to generative models (e.g, bias, hallucinations, safety)
- Familiarity with Gen AI frameworks and tools (e.g, Hugging Face and LangChain)
- Strong foundation in data analysis
- Research and experiment design
- Visualization with data
- Answering questions with data
- Strong python programming skills. Familiarity with standard data science packages
- Familiarity with standard software development best practices. Strong SQL skills a plus
- Understanding of standard algorithms and data structures (ex. search & sort) and their analysis
Required Skills: Data Science, Machine Learning, Deep Learning, Statistical Modeling, Healthcare Payer Domain, Algorithmic Design, Python Programming, Machine Learning Pipelines, Classic Machine Learning Algorithms, Supervised Learning, Semi-supervised Learning, Unsupervised Learning, Reinforcement Learning, Regression, Classification, Time Series Modeling, Transfer Learning, Objective Functions, Regularization, Overfitting, Hyperparameter Tuning, Model Evaluation Metrics, Feature Importance, Scikit-learn, XGBoost, TensorFlow, PyTorch, Natural Language Processing, Tokenization, Part-of-Speech Tagging, Named Entity Recognition, Sentiment Analysis, Machine Translation, Pre-trained NLP Models, BERT, GPT, spaCy, Text Preprocessing, Feature Extraction, Embedding Methods, NLTK, Hugging Face Transformers, Text Classification, Language Generation, Information Extraction, Large Language Models, Prompt Engineering, Fine-tuning Techniques, Retrieval-Augmented Generation, Generative Model Evaluation, Alignment Techniques, Multimodal Generative Models, Responsible AI, Data Analysis, Research Design, Data Visualization, SQL, Algorithms, Data Structures, Interest in Healthcare Industry
Benefits: Generous benefits, Work from home options, Participating in Cambia-supported outreach programs
Benefits
Generous benefits
Work from home options
Participating in Cambia-supported outreach programs