CGI→
Databricks Engineer at CGI in Hybrid - Hyderabad
ExperiencedHybridHybrid - Hyderabad
Skills
databrickspysparkData LakeSQLData BricksData
Job Description
Role & responsibilities
We are seeking a skilled Databricks Engineer with strong expertise in PySpark, Databricks, and Azure/AWS Cloud platforms. The ideal candidate will be responsible for building scalable data pipelines, processing large datasets, optimizing ETL workflows, and supporting data engineering initiatives within a modern cloud-based data ecosystem.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark and Databricks.
- Build and optimize ETL/ELT processes for processing large-scale structured and unstructured data.
- Develop data ingestion frameworks from multiple data sources including databases, APIs, and cloud storage.
- Work with Delta Lake, Spark SQL, and Databricks Workflows for efficient data processing.
- Implement data quality checks, monitoring, and troubleshooting of data pipelines.
- Collaborate with data architects, analysts, and business stakeholders to understand data requirements.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Integrate Databricks solutions with cloud platforms such as Azure, AWS, or GCP.
- Support CI/CD implementation and deployment of data engineering solutions.
- Ensure data governance, security, and compliance standards are followed.
Required Skills
- 6+ years of experience in Data Engineering.
- Strong hands-on experience with Databricks and PySpark.
- Expertise in Apache Spark, Spark SQL, and DataFrame APIs.
- Strong knowledge of ETL/ELT development and data warehousing concepts.
- Experience with Delta Lake, Databricks Notebooks, and Databricks Workflows.
- Proficiency in Python programming.
- Strong SQL skills and database concepts.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Knowledge of Git, CI/CD pipelines, and DevOps practices.
- Experience working with large-scale distributed data processing systems.
Preferred Skills
- Experience with Azure Data Factory (ADF), Azure Data Lake Storage (ADLS), or AWS S3.
- Knowledge of Kafka, Event Hubs, or streaming technologies.
- Exposure to Data Modeling and Data Governance concepts.
- Experience with Airflow or other workflow orchestration tools.
- Databricks Certification is a plus.