MKThink→
Summer Internship - Data Science / Data Engineering at MKThink · San…
InternshipOn-siteSan Francisco, CA
Skills
pythondata pipelinesmachine learningnatural language processingunstructured data processingdocument parsingfeature engineeringrule-based logicinformation extractionschema mappingconfidence scoringvalidationerror-handling
Job Description
Summary: MKThink is a future-forward design firm dedicated to creating intelligent spaces that improve quality of life. They are seeking a Data Science or Data Engineering intern to support the development of an unstructured data extraction pipeline, building systems that ingest documents and improve output accuracy through user feedback.
Responsibilities:
- Build data pipelines for ingesting and processing unstructured documents, including PDFs with inconsistent structure and content
- Develop extraction workflows that combine document parsing, feature engineering, ML/NLP methods, and rule-based logic to identify relevant fields
- Design methods to evaluate what content is useful, discard irrelevant content, and align extracted information to a predefined schema
- Implement confidence scoring, validation, and error-handling logic to improve extraction accuracy and reliability
- Build a human-in-the-loop feedback workflow where users can confirm, reject, or correct extracted fields and trigger reruns toward improved output
Required Qualifications:
- Graduate student preferred, advanced undergraduate considered
- Strong Python
- Experience with data pipelines
- Experience with machine learning
- Experience with unstructured data processing
- Build data pipelines for ingesting and processing unstructured documents
- Develop extraction workflows that combine document parsing, feature engineering, ML/NLP methods, and rule-based logic
- Design methods to evaluate what content is useful, discard irrelevant content, and align extracted information to a predefined schema
- Implement confidence scoring, validation, and error-handling logic to improve extraction accuracy and reliability
- Build a human-in-the-loop feedback workflow where users can confirm, reject, or correct extracted fields and trigger reruns toward improved output
Preferred Qualifications:
- Graduate student preferred in Data Science, Computer Science, Engineering, or related field
- Experience working with unstructured data, document intelligence, information extraction, or schema mapping
- Comfortable working on applied modeling and data engineering problems with ambiguous inputs and variable document formats
- Proactive and highly self-motivated, able to operate independently with minimal guidance and supervision
Required Skills: Python, Data pipelines, Machine learning, Natural language processing, Unstructured data processing, Document parsing, Feature engineering, Rule-based logic, Information extraction, Schema mapping, Confidence scoring, Validation, error-handling