- Posted 2 days ago
- Be among the first 20 applicants
Job Description
Experience: 3–5 Years (Maximum 5 Years)
Relevant Experience: 3+ Years
Locations: Bangalore – 6 positions | Hyderabad – 2 positions
Work Mode: WFO – 5 Days/Week
CTC: 20–22 LPA
Role Overview
We are looking for Data Engineers to join a Datastore Migration Factory team responsible for end-to-end migration from an on-premise Data Lake to an AWS-hosted LakeHouse.
The role involves pipeline migration, SQL/Spark code conversion, data migration, data reconciliation, data modelling, and stakeholder coordination to ensure migrated data and consumption patterns meet business requirements.
Key Responsibilities
1. Pipeline Migration
- Refactor and migrate extraction logic and job scheduling from legacy frameworks to the new LakeHouse environment.
- Execute physical migration of datasets while maintaining data integrity.
- Coordinate with data owners for technical hand-off and sign-off.
2. Consumption Pattern Migration
- Convert and optimize legacy SQL and Spark consumption patterns for Snowflake and Apache Iceberg.
- Analyze usage patterns and deliver required data products.
- Coordinate with stakeholders for hand-off and sign-off.
3. Data Reconciliation & Quality
- Perform rigorous data validation and reconciliation.
- Use reconciliation frameworks to establish functional equivalence between migrated and production data.
- Identify and troubleshoot data discrepancies.
4. Engineering & Collaboration
- Work with internal data management platform teams.
- Learn and adapt to new workflows, tools, and language constructs.
- Follow SDLC and CI/CD best practices.
- Support Kubernetes (K8s) deployments.
Mandatory Technical Skills
- 3–5 years hands-on Data Engineering experience
- SQL
- Data Modelling
- Pipeline/Data Migration
- Python OR Java
- Strong SQL troubleshooting
- SDLC & CI/CD
- Kubernetes/K8s deployment experience
- Temporal Data Modelling – SCD Type 2
- Schema Evolution & Schema Management
- Data Partitioning & Clustering
- Normalization vs. Denormalization
- Natural vs. Surrogate Keys
Technical Stack
Extraction & Processing
- Kafka
- ANSI SQL
- FTP
- Apache Spark
Data Formats
- JSON
- Avro
- Parquet
Platforms
- Hadoop / HDFS / Hive
- Snowflake
- Apache Iceberg
- Sybase IQ
Candidate Profile
Candidates should demonstrate:
- Strong analytical and troubleshooting ability.
- Ownership and delivery focus.
- Clear communication and stakeholder management.
- Ability to collaborate with global and cross-functional teams.
- Ability to identify risks and resolve issues constructively.
- Willingness to learn new technologies and workflows.
More Info
Key Skills
Apache Iceberg
Schema Evolution
SQL troubleshooting
Parquet
Data Partitioning
Pipeline Data Migration
Schema Management
Temporal Data Modelling
Denormalization
CI CD
Surrogate Keys
