Abaka AI→
Data Engineer at Abaka AI in Palo Alto, CA
Entry LevelOn-siteFull-timePalo Alto, CA$150k–$225k/yr
Skills
data engineeringartificial intelligencelarge-scale data systemsmultimodal data workflowstext data processingimage data processingaudio data processingvideo data processingdata pipeline architecturedistributed processingautomation frameworksdata privacydata securityaccess controldata anonymizationexecutionresilience
Job Description
Summary: Abaka AI is built on a mission to be the world’s most trusted data partner for AI companies, serving over 1,000 industry leaders. The Data Engineer will shape data engineering standards and systems, taking ownership of multimodal data processes to deliver high-quality datasets for advanced AI teams.
Responsibilities:
- Work closely with foundation model clients to understand their data requirements, and coordinate internal teams to create tailored delivery plans that ensure on-time, high-quality data delivery, including meeting expectations for format, precision, and volume
- Lead the development of mid- to long-term plans for the data engineering function. Build scalable, end-to-end pipelines for multimodal data (text, image, audio, video, 3D point cloud, etc.) across data sourcing, cleaning, annotation, QA, storage, and iterative optimization for training, fine-tuning, and evaluation
- Develop solutions to core technical challenges in multimodal data processing, including cross-modal alignment (for example, image-text semantic matching), large-scale data cleaning (deduplication, denoising, format normalization), annotation efficiency, and data encryption and security
- Partner with algorithm, product, and business teams by providing feedback on data bottlenecks, refining internal tooling and services, and supporting client-facing teams with technical documentation and pre-sales materials
- Evaluate and optimize the cost structure of data processing operations, including headcount, infrastructure, and tooling, to balance quality, efficiency, and scalability
Required Qualifications:
- Strong background in computer science, data engineering, artificial intelligence, or related fields, with hands-on experience building or operating large-scale data systems
- 1+ years of experience in data engineering or data operations. Leadership experience is highly valued, and experience with LLM or multimodal dataset preparation is a strong plus
- Deep understanding of end-to-end multimodal data workflows, with hands-on experience in at least two modalities (text, images, audio, or video)
- Proficiency in designing technical architectures for large-scale data pipelines, including distributed processing and automation frameworks, along with familiarity with data privacy and security best practices such as access control and data anonymization
- Strong execution and team management capabilities, with the ability to translate high-level objectives into actionable plans and drive team results
- Excellent communication and cross-functional collaboration skills, with the ability to clearly articulate technical and operational requirements, resolve conflicts, and manage stakeholder expectations
- High sense of ownership and resilience, with comfort working in a fast-paced, rapidly evolving AI environment and the ability to manage urgent delivery timelines
Required Skills: Data engineering, Artificial intelligence, Large-scale data systems, Multimodal data workflows, Text data processing, Image data processing, Audio data processing, Video data processing, Data pipeline architecture, Distributed processing, Automation frameworks, Data privacy, Data security, Access control, Data anonymization, Execution, Resilience
Benefits: Equity, Comprehensive benefits package, Health, Dental, Vision, PTO, Flexible work schedule
Benefits
Equity
Comprehensive benefits package
Health
Dental
Vision
PTO
Flexible work schedule