GreenGas USA→
Data Engineer - Intern at GreenGas USA in Houston, TX
InternshipHybridFull-timeHouston, TX
Skills
pythonsqlazureetl toolsdata warehousingbig data technologiesjavadatabase systemsdata modelingscriptingautomation
Job Description
Summary: GreenGas USA is a company focused on leading the global energy transition through renewable energy solutions. The Data Engineer Intern will support data-driven initiatives by assisting in building and maintaining data infrastructure, contributing to data pipelines, and ensuring data quality and accessibility for analysis and reporting.
Responsibilities:
- Build robust Extract, Transform, Load (ETL) or Extract, Load, Transform (ELT) pipelines to move data from various sources (databases, APIs, streaming sources, external data providers) into data warehouses, data lakes, or other data repositories
- Create efficient and scalable data pipelines to handle large volumes of structured and unstructured data
- Data engineers will write code and scripts to automate repetitive tasks in data collection, processing, and delivery. Ie Databases to PBIs
- Continuously monitor pipeline performance, identify and resolve data-related issues, errors, and performance bottlenecks
- Selecting appropriate technologies for data storage such as relational databases, NoSQL databases, data lakes, data warehouses, cloud storage services) and designing efficient data models, schemas
- Work to ensure data is stored in a way that allows for high-performance queries and efficient analytical and operational use cases
- Research and implement new data technologies, tools, and frameworks to improve data infrastructure and processes
- Proficiency in Azure to deploy and manage data solutions in the cloud, including automation, databases and PowerBI’s
- Implement validation rules, cleansing procedures, and monitoring systems to detect and rectify anomalies, ensuring data accuracy
- Define standards and policies for data usage, ensuring consistency, reliability, and compliance (e.g., GDPR, HIPAA)
- Implement security controls and access management policies to protect sensitive information from unauthorized access or theft
- Set up dashboards, analytics tools, and API endpoints to make processed data accessible to end-users and applications
- Create comprehensive documentation for data pipelines, architecture, and processes to facilitate system transparency
- By providing clean, reliable, and accessible data, data engineers enable organizations to make informed business decisions
Required Qualifications:
- Programming Languages: Python, Java, SQL and related technologies
- Database Systems: Strong knowledge of relational databases (e.g., PostgreSQL, MySQL, SQL Server) and NoSQL databases (e.g., MongoDB, Cassandra)
- Data Warehousing: Experience with data warehousing and platforms (e.g., Snowflake)
- ETL Tools: Proficiency with various ETL tools such as snowflake and Databricks
- Cloud Platforms: Experience with Azure Suite, Jira DevOps and Ticketing system
- Big Data Technologies: (Hadoop, Spark, Kafka, Hive)
- Data Modeling and Schema Design: Ability to design efficient data models
- Scripting and Automation: For automating data processes
Required Skills: Python, SQL, Azure
Important Skills: ETL Tools, Data Warehousing, Big Data Technologies
Nice-to-Have Skills: Java, Database Systems, Data Modeling, Scripting, Automation