

Search by job, company or skills

Get to Know Indium Tech
Indium Software is a leading provider of Digital Engineering Services, helping clients drive measurable business
value through technology. We provide services across Application Engineering, Data & Analytics, Cloud
Engineering, Digital Assurance, and Low-Code Development.
Indium is an AI-driven digital engineering company with 5,000+ associates globally and more than 25 years in
business. Its expertise spans Generative AI, Product Engineering, Intelligent Automation, Data & AI, Quality
Engineering, and Gaming.
Indium partners with Fortune 500, Global 2000, and leading technology firms across multiple industries and
geographies.
Learn more: https://www.indium.tech/#
About the Role
We are looking for a Data Engineer to build and maintain scalable, reliable data pipelines and data platforms. The
role involves working across data ingestion, transformation, processing, warehousing, and analytics while
leveraging modern big-data and cloud technologies.
You will work closely with data analysts, data scientists, and engineering teams to ensure high-quality, accessible,
and production-ready data.
Key Responsibilities
• Design, develop, and maintain scalable data pipelines using Python and PySpark.
• Build robust ETL/ELT workflows for ingesting, transforming, validating, and integrating data from multiple sources.
• Write and optimize advanced SQL queries for large-scale data processing and analytics.
• Develop and maintain data warehouse solutions and dimensional data models.
• Work with Hadoop and other big-data technologies to process high-volume datasets.
• Implement data quality, validation, monitoring, and error-handling mechanisms across pipelines.
• Collaborate with cross-functional teams to understand data requirements and translate them into scalable
engineering solutions.
• Optimize data pipelines for performance, reliability, scalability, and cost efficiency.
• Contribute to modern data-platform and AI-enabled data engineering initiatives.
Required Skills
• Strong hands-on experience with Python.
• Strong proficiency in PySpark and distributed data processing.
• Advanced SQL, including complex queries, joins, CTEs, window functions, and query optimization.
• Strong understanding of Data Engineering principles and best practices.
• Hands-on experience with ETL/ELT pipeline development.
• Strong understanding of Data Warehousing concepts and data modeling.
• Experience with Hadoop and Big Data technologies.
• Strong debugging, problem-solving, and communication skills.
Preferred Skills
• Experience with Apache Airflow for workflow orchestration.
• Understanding of Vector Embeddings and their application in modern data and AI systems.
• Exposure to Agentic Frameworks and LLM workflows.
• Familiarity with MCP Servers and modern data platforms.
• Understanding of Semantic Search and retrieval-oriented data systems.
• Experience with cloud platforms such as AWS, GCP, or Azure.
What We Value
• Strong analytical and problem-solving ability.
• Ability to work independently and collaboratively in a fast-paced environment.
• Good understanding of scalable and production-grade data engineering practices.
• Strong communication and stakeholder collaboration skills.
• Curiosity and willingness to learn modern data, cloud, and AI technologies.
Why Join Indium
Work on modern Data Engineering, Big Data, Cloud, and AI-enabled data platform initiatives while contributing
to scalable solutions that create measurable business impact.
Job ID: 153617403
Skills:
Apache Airflow, Spark, Apache Beam, Python, Data Platform Management, Data Layer Design, dbt, Google Cloud Ecosystem, Google Cloud Services, Data Pipeline Development
Skills:
snowflake , Data Transformation, ELT, Sql Queries, Data Cleansing, MS SQL, Views, Etl, complex joins, data processing pipelines, stored procedures, temporary tables, Data Validation, CTEs
Skills:
Pl Sql, Ddl, Informatica, Bodi, Data Integration, Kinesis, Odi, RDBMS, Datastage, Data Governance, Python, AWS, Hadoop, Scala, Emr, Sparksql, SSIS, Sql, Hive, Hiveql, Spark, Data Warehousing, Amazon, Etl, DataCraft, R, Redshift Spectrum, MDX, Data Lakes, Cradle, Glue, KornShell
Skills:
snowflake , Sql Queries, Pyspark, Python Programming, Databricks, AWS, HCP and HCO data management, Azure blob storage, CDW Commercial Data Warehouse
Skills:
tokenization , Spark SQL, Apache Spark, Git, Ocr, Python, Hugging Face Transformers, Text Extraction, PDF Parsing, Vector Databases, Document Chunking, Data pipelines, CI CD practices, Text Processing, NLP concepts, ETL workflows, HTML Parsing