iManage→
NLP Data Engineering Intern at iManage in Chicago, IL
InternshipOn-siteChicago, IL
Skills
pythonnatural language processingtokenizationembeddingssemantic searchspacyhuggingface datasetsnltkdata structuresalgorithmsstatisticsgitcuriositycollaborative mindset
Job Description
Summary: iManage is dedicated to Making Knowledge WorkTM, providing a dynamic internship program for students. As an NLP Data Engineering Intern, you will gain hands-on experience in designing and optimizing text data pipelines to support AI and machine learning solutions, collaborating with various teams to enhance document data for intelligent features across the platform.
Responsibilities:
- Performing exploratory analyses on large text corpora and developing preprocessing pipelines for training and evaluation data
- Supporting the design of automated workflows for text normalization, deduplication, language identification, PII redaction, and metadata enrichment
- Assisting with building automated data validation processes to ensure accuracy and consistency of NLP datasets
- Contributing to dataset curation, prompt dataset preparation, labeling coordination, and text quality validation to support model fine-tuning, semantic search, and Gen AI evaluations
- Partnering with the Applied AI team to understand data requirements and help build data interfaces for machine learning systems
- Learning and applying data lineage best practices and data privacy, security, and governance principles
- Maintaining highest quality standards through processes that identify and correct mistakes and inconsistencies
Required Qualifications:
- Current enrollment in a Master's, or PhD program in Computer Science, Data Engineering, Data Science, Applied Mathematics, Computational Linguistics, or a related quantitative field
- Proficiency in Python and experience using it to extract, structure, classify, and analyze text data
- Foundational understanding of NLP concepts such as tokenization, embeddings, and semantic search
- Familiarity with standard NLP libraries such as SpaCy, HuggingFace Datasets, or NLTK
- Solid knowledge of data structures, algorithms, and statistics
- Proficiency with Git and collaborative development workflows
- A passion to learn and improve, and an eagerness to share knowledge with colleagues
- Problem-solving, creativity, curiosity, and a collaborative mindset
Preferred Qualifications:
- Exposure to Microsoft Azure services such as Fabric, ADLS, AI Foundry, or Azure ML
- Experience with data pipeline orchestration or workflow automation tools like Databricks
- Familiarity with knowledge graphs or semantic data modeling
Required Skills: Python, Natural Language Processing, Tokenization, Embeddings, Semantic search, SpaCy, HuggingFace Datasets, NLTK, Data structures, Algorithms, Statistics, Git, Curiosity, Collaborative mindset