Northeastern University→
AI/ML Engineer for Assessment Systems at Northeastern… · Boston
Entry LevelOn-siteFull-timeBoston, MA (Main Campus)$52k–$52k/yr
Skills
python programmingnatural language processinglarge language model apisnlp pipeline designevaluation metricsversion control (git)prompt engineeringeducational assessmentbias detection metricsretrieval-augmented generation
Job Description
Summary: Northeastern University is seeking an AI/ML Engineer for Assessment Systems to develop and validate AI scoring systems for evaluating student transcripts. The role involves working with natural language processing and educational assessment methodologies to create scalable evaluation tools in collaboration with the research team.
Responsibilities:
- Design AI scoring architecture combining rule-based and LLM components
- Develop prompt engineering strategies for rubric-aligned evaluation
- Implement quantitative indicator scoring (frequency counts, behavioral thresholds)
- Create LLM-based qualitative judgment systems for complex indicators
- Build data processing pipelines for transcript preparation and analysis
- Compare AI ratings against human expert ground truth (n=200+ transcripts)
- Calculate agreement statistics (correlation, Cohen's kappa, confusion matrices)
- Identify systematic rating discrepancies and refine prompts/algorithms
- Implement confidence scoring to flag uncertain AI judgments
- Conduct bias audits across demographic subgroups (gender, race/ethnicity)
- Deploy validated scoring system for 300+ student transcripts
- Create technical documentation for scoring methodology
- Develop user-facing explanations of AI ratings for educators
Required Qualifications:
- Graduate student or professional in Computer Science, Data Science, Computational Linguistics, or related field
- Strong Python programming skills (pandas, scikit-learn, numpy)
- Experience with large language model APIs (OpenAI GPT-4, Claude, or similar)
- Demonstrated ability to design and implement NLP pipelines
- Understanding of evaluation metrics (precision, recall, F1, correlation, agreement statistics)
- Experience with version control (Git) and collaborative coding practices
Preferred Qualifications:
- Experience with prompt engineering and LLM optimization
- Background in natural language processing or computational linguistics
- Familiarity with educational assessment or psychometrics
- Knowledge of inter-annotator agreement statistics (Cohen's kappa, Krippendorff's alpha)
- Experience with retrieval-augmented generation (RAG) systems
- Understanding of bias detection and fairness metrics in AI systems
Required Skills: Python programming, Natural language processing, Large language model APIs
Important Skills: NLP pipeline design, Evaluation metrics, Version control (Git)
Nice-to-Have Skills: Prompt engineering, Educational assessment, Bias detection metrics, Retrieval-augmented generation
Benefits: Medical, Vision, Dental, Paid time off, Tuition assistance, Wellness & life, Retirement-, Commuting & transportation
Benefits
Medical
Vision
Dental
Paid time off
Tuition assistance
Wellness & life
Retirement-
Commuting & transportation