Goldman Sachs→
Data Engineering - Data, Lakehouse and… at Goldman Sachs · New York
Entry LevelOn-siteFull-timeNew York, NY$110k–$130k/yr
Skills
pythonjavasqlsoftware engineering fundamentalsversion controltestingci/cd practicestemporal data modellingschema designschema evolutiondata compatibilitydata partitioningdata clusteringdata qualitydata reconciliationroot-cause analysisproduction data pipelinesdistributed data processing frameworksapache sparkjsonavroparquetansi sqlkafkasnowflakeapache icebergdatabrickshadoop ecosystemsybase iqcontainerized deploymentkubernetesjudgement in technical trade-offswillingness to work with stakeholders
Job Description
Summary: Goldman Sachs is seeking a Data Engineer to join their Lakehouse and AI Data Platform team. The role involves designing, building, and supporting data pipelines and curated datasets, ensuring data products are reliable and scalable while collaborating with various stakeholders to meet business needs.
Responsibilities:
- Build, enhance and support batch and streaming data pipelines on the Lakehouse and AI data platform
- Refactor or modernise existing data flows where needed to improve reliability, performance and maintainability
- Where needed, build reusable tooling to improve delivery, consistency and operational support
- Ensure data pipelines are production-ready, well tested and operationally supportable
- Develop raw, refined and curated datasets that support analytics, reporting and AI use cases
- Apply sound data modelling principles to represent business entities, relationships and historical change accurately
- Work with consumers to shape data products that are usable, well documented and aligned to business needs
- Implement controls to validate completeness, accuracy and consistency of data across pipelines and datasets
- Use reconciliation approaches to build confidence in production outputs and investigate breaks where they arise
- Contribute to clear standards for testing, monitoring and issue resolution
- Contribute to practical improvements in testing, monitoring or reconciliation tooling where these strengthen platform reliability and day-to-day delivery
- Work closely with engineers, platform teams and data consumers to deliver agreed outcomes to time and quality expectations
- Communicate clearly on progress, risks, dependencies and design choices, including where delivery would benefit from improvements to shared platform tooling
Required Qualifications:
- Bachelor's or master's degree in a relevant discipline, or equivalent practical experience, with evidence of strong quantitative skills or data engineering expertise
- Strong hands-on programming experience in Python or Java
- Good working knowledge of SQL, including troubleshooting, optimization and data analysis
- Ability to learn new tools, internal platforms and delivery workflows quickly
- Familiarity with software engineering fundamentals, including version control, testing, release discipline and CI/CD practices
- Understanding of temporal data modelling, including the handling of historical state and change over time
- Knowledge of schema design, schema evolution and data compatibility considerations
- Understanding of partitioning, clustering and other techniques used to improve data performance at scale
- Ability to make sensible design choices across normalized and deformalized models, and between natural and surrogate keys
- Practical approach to data quality, reconciliation and root-cause analysis
- Experience building or supporting production data pipelines in a collaborative engineering environment
- Experience working with distributed data processing frameworks such as Apache Spark
- Working knowledge of common data formats such as JSON, Avro and Parquet
Preferred Qualifications:
- Stronger ownership of technical design across multiple datasets or pipeline domains
- Experience guiding implementation standards, code quality and engineering practices within a team
- Ability to lead delivery for a workstream, manage dependencies and support less experienced engineers
Required Skills: Python, Java, SQL, Software engineering fundamentals, Version control, Testing, CI/CD practices, Temporal data modelling, Schema design, Schema evolution, Data compatibility, Data partitioning, Data clustering, Data quality, Data reconciliation, Root-cause analysis, Production data pipelines, Distributed data processing frameworks, Apache Spark, JSON, Avro, Parquet, ANSI SQL, Kafka, Snowflake, Apache Iceberg, Databricks, Hadoop ecosystem, Sybase IQ, Containerized deployment, Kubernetes, Judgement in technical trade-offs, Willingness to work with stakeholders
Benefits: Discretionary bonus, Valuable and competitive benefits and wellness offerings
Benefits
Discretionary bonus
Valuable and competitive benefits and wellness offerings