Abaka AI→
Data Engineer at Abaka AI in Palo Alto, CA
Entry LevelOn-siteFull-timePalo Alto, CA$150k–$225k/yr
Skills
data engineeringmultimodal data workflowslarge-scale data systemstechnical architecture designdata privacysecurityteam managementcross-functional collaborationcommunicationhigh sense of ownership
Job Description
Summary: Abaka AI is built on a mission to be the world’s most trusted data partner for AI companies, serving over 1,000 industry leaders. The Data Engineer will shape data engineering standards and systems, taking ownership of multimodal data processes to deliver high-quality datasets for advanced AI teams.
Responsibilities:
- Work closely with foundation model clients to understand their data requirements, and coordinate internal teams to create tailored delivery plans that ensure on-time, high-quality data delivery, including meeting expectations for format, precision, and volume
- Lead the development of mid- to long-term plans for the data engineering function. Build scalable, end-to-end pipelines for multimodal data (text, image, audio, video, 3D point cloud, etc.) across data sourcing, cleaning, annotation, QA, storage, and iterative optimization for training, fine-tuning, and evaluation
- Develop solutions to core technical challenges in multimodal data processing, including cross-modal alignment (for example, image-text semantic matching), large-scale data cleaning (deduplication, denoising, format normalization), annotation efficiency, and data encryption and security
- Partner with algorithm, product, and business teams by providing feedback on data bottlenecks, refining internal tooling and services, and supporting client-facing teams with technical documentation and pre-sales materials
- Evaluate and optimize the cost structure of data processing operations, including headcount, infrastructure, and tooling, to balance quality, efficiency, and scalability
Required Qualifications:
- Strong background in computer science, data engineering, artificial intelligence, or related fields, with hands-on experience building or operating large-scale data systems
- 1+ years of experience in data engineering or data operations. Leadership experience is highly valued, and experience with LLM or multimodal dataset preparation is a strong plus
- Deep understanding of end-to-end multimodal data workflows, with hands-on experience in at least two modalities (text, images, audio, or video)
- Proficiency in designing technical architectures for large-scale data pipelines, including distributed processing and automation frameworks, along with familiarity with data privacy and security best practices such as access control and data anonymization
- Strong execution and team management capabilities, with the ability to translate high-level objectives into actionable plans and drive team results
- Excellent communication and cross-functional collaboration skills, with the ability to clearly articulate technical and operational requirements, resolve conflicts, and manage stakeholder expectations
- High sense of ownership and resilience, with comfort working in a fast-paced, rapidly evolving AI environment and the ability to manage urgent delivery timelines
Required Skills: Data engineering, Multimodal data workflows, Large-scale data systems
Important Skills: Technical architecture design, Data privacy, security
Nice-to-Have Skills: Team management, Cross-functional collaboration, communication, High sense of ownership
Benefits: Equity, Comprehensive benefits package, Health, Dental, Vision, PTO, Flexible work schedule
Benefits
Equity
Comprehensive benefits package
Health
Dental
Vision
PTO
Flexible work schedule