Infinite Computer Solutions→
Senior Data Engineer at Infinite… · Location…
ExperiencedOn-siteNot specified
Job Description
Senior Data Engineer (GCP & Legacy Migration Specialist)
Role Overview
We are seeking a highly skilled Senior Data Engineer with extensive experience in cloud modernization, large-scale data architecture, and legacy database migration. In this role, you will lead the architecture, design, and execution of data pipelines on Google Cloud Platform (GCP), modernizing traditional RDBMS and Hadoop/Big Data infrastructure into cloud-native BigQuery environments.
Key Responsibilities
Role Overview
We are seeking a highly skilled Senior Data Engineer with extensive experience in cloud modernization, large-scale data architecture, and legacy database migration. In this role, you will lead the architecture, design, and execution of data pipelines on Google Cloud Platform (GCP), modernizing traditional RDBMS and Hadoop/Big Data infrastructure into cloud-native BigQuery environments.
Key Responsibilities
- Cloud Architecture & GCP Data Pipeline Engineering (30%)
- Architect, build, and optimize enterprise-scale cloud data warehouses on GCP using BigQuery, Cloud Storage, Pub/Sub, and Cloud Dataproc.
- Design and manage robust cloud workflows using Cloud Composer (Apache Airflow), troubleshooting complex pipeline DAG dependencies and orchestration issues.
- Implement cross-project service accounts and secure integration patterns to ingest data safely from heterogeneous cloud and on-premises systems.
- Legacy Migration & Database Modernization (25%)
- Drive the end-to-end migration of legacy on-premises ETL workloads (Informatica PowerCenter, IDQ) and traditional databases (Teradata, Oracle) to GCP.
- Translate, refactor, and optimize complex Teradata SQL and stored procedures into highly performant BigQuery SQL and native cloud constructs.
- Leverage deep expertise in legacy utilities (BTEQ, FastLoad, MultiLoad, TPump, TPT) to extract and transfer massive historical datasets smoothly.
- Big Data Processing & Distributed Computing (20%)
- Process massive volumes of distributed data utilizing Apache Spark (Spark Core, Spark SQL, Spark Streaming) using Scala and PySpark.
- Refactor legacy Hive/SQL queries into high-performance Spark DataFrames and distributed data transformations.
- Implement advanced data optimization techniques, including partitioning and bucketing strategies across Hive, managed/external tables, and BigQuery tables.
- Data Governance, Quality & Automation (15%)
- Establish enterprise Data Quality framework rules and validation pipelines (CDQ, IDQ) to guarantee accurate reporting and downstream analytics.
- Write complex UNIX shell scripts and build automation utilities to execute routine data transport, loading, and staging workloads.
- Enforce robust automated workflow scheduling across time-driven and data-driven triggers using Oozie and Cloud Composer.
- DevOps, CI/CD & Production Operations (10%)
- Deploy data platform pipelines and infrastructure code using Continuous Integration and Continuous Deployment (CI/CD) practices.
- Perform advanced query optimization, schema refactoring, and database tuning across RDBMS platforms (MySQL, SQL Server, Oracle, Teradata) and BigQuery.
- Collaborate with cross-functional analytics and business intelligence teams to deliver clean, staged, and curated data models.