- Posted 6 hours ago
- Be among the first 10 applicants
Job Description
Job Responsibilities
Own the end-to-end construction and maintenance of the company's real-time data ecosystem, from ingestion to warehouse architecture to streaming pipeline stability, providing reliable, low-latency data support for business teams, analysts and global data products.
Real-Time Data Ingestion System Construction
Own the construction and maintenance of the company's business data ingestion system. Responsible for the full lifecycle development of real-time collection, cleaning and warehousing of business logs, user behavioural data and database operational data, ensuring the completeness, timeliness and stability of end-to-end data ingestion pipelines.
Real-Time Data Warehouse Architecture & Optimisation
Lead the implementation and iterative optimisation of the real-time data warehouse architecture. Design multi-layer real-time data models, develop and tune streaming ETL pipelines, and support the stable output of real-time metrics, dashboards and business monitoring data.
Cross-Functional Requirement Delivery
Collaborate closely with business teams, data analysts and data product managers. Fully understand business logic and statistical calibres, and independently complete data requirement decomposition, logical sorting, development implementation, release verification and continuous iteration.
Kafka Streaming Pipeline Operation & Stability Governance
Maintain and optimise Kafka streaming pipelines for real-time data ingestion and consumption. Troubleshoot and resolve core online issues including message backlog, data skew, data loss and duplication, and ensure high availability and stability of real-time data pipelines.
Spark Streaming Task Development & Performance Tuning
Develop, tune and operate Spark real-time tasks. Continuously optimise computing resource utilisation, data latency and throughput, and improve the overall performance and stability of streaming data pipelines.
Data Standardisation, Quality Governance & Downstream Data Service Support
Participate in data standardisation, real-time data quality governance and pipeline redundancy optimisation. Summarise and standardise data ingestion specifications and development best practices to improve the team's overall data development efficiency. Provide stable real-time data support for downstream scenarios including BI dashboards, data services, user portrait platforms and business effectiveness measurement, ensuring reliable online data service delivery.
Job Requirements
Basic Qualifications
- Bachelor's degree or above in Computer Science, Software Engineering, Data Science or a related field. 3+ years of professional big data development experience in the internet industry, with proven experience in end-to-end real-time data warehouse implementation and data ingestion system construction.
- Proficient in English reading, writing and verbal communication, capable of collaborating with overseas teams and business stakeholders. Overseas study or work experience is highly preferred.
Technical Competency
- Expert-level knowledge of Kafka. Familiar with real-time data ingestion, partition strategy and consumer mechanisms, and capable of diagnosing and optimising online risks such as message backlog, data skew, and data loss and duplication; experienced in large-scale streaming pipeline stability governance.
- Hands-on experience with Spark, Spark SQL, Flink and Flink SQL. Solid capability in streaming task development, resource tuning, and latency and throughput optimisation, and online troubleshooting. Familiar with real-time data warehouse layered architecture and dimensional modelling methodologies.
- Deep understanding of big data ingestion systems. Experienced in full-process real-time processing of multi-source data, including user behaviour logs and business database data; proficient in real-time collection, CDC synchronisation, data cleaning and exception handling with standardised implementation capability.
- Practical experience with cloud-native data tools including BigQuery (BQ), dbt and Databricks. Capable of data warehouse modelling, pipeline orchestration, metric iteration and dashboard development in cloud data environments.
- Proficient in at least one programming language among Python, Scala and Java, with solid hands-on big data development capability. Familiar with Linux commands and Shell scripting, able to independently complete task deployment, log analysis and day-to-day operations and maintenance.
- Familiar with mainstream relational databases and Binlog/CDC real-time synchronisation principles. Experienced in business database real-time ingestion implementation. Knowledgeable in real-time data SLA governance, data quality monitoring, alert configuration and incident review mechanisms.
Preferred Qualifications
- Experience with Google Cloud, Azure or other mainstream cloud big data platforms.
- Prior experience collaborating with overseas teams and governing real-time data quality across global business units.
More Info
Key Skills
Flink SQL
Metric iteration
Streaming ETL pipelines
Flink
Data quality monitoring
Real-time data ingestion
Cloud data environments
Pipeline orchestration
Incident review mechanisms
Binlog CDC
Data warehouse modelling
Real-time data SLA governance
Alert configuration
Dashboard development



