Job Description:
In a data-driven and evaluation-driven manner, build an efficient closed loop for data iteration and establish an end-to-end data system spanning data sourcing, labeling, processing, synthesis, and evaluation. Continuously build high-quality datasets and evaluation sets to keep improving foundation model capabilities and to drive the development of AI models and applications.
Responsibilities cover one or more of the following directions:
- Design and implement high-performance, scalable, and distributed data infrastructure covering the full lifecycle - data storage, ingestion, cleaning, labeling, management, and analysis - and continuously improve data engineering efficiency.
- Design audio-visual multimodal training data strategies develop efficient data processing, synthesis, and optimization operators and pipelines, and build a multimodal data asset repository to meet the data needs of large model development.
- Build a data-model-evaluation closed loop together with Agents, using data to drive rapid iteration of large models.
- Track cutting-edge techniques and methods in the large-model data domain, explore innovative approaches such as data augmentation, data distillation, and high-quality data filtering, and land them in real business scenarios to increase data value.
Requirements:
- Bachelor's degree or above in Computer Science, Software Engineering, Data Science, Statistics, or a related field.
- Solid programming fundamentals proficient in Python and Java, competent in SQL strong command of common data structures and algorithms.
- Familiar with at least one big data processing framework (any of Spark / Flink / Ray), or strong self-learning ability backed by relevant coursework / projects.
- Familiar with the fundamentals of large models / multimodal / AIGC (LLM, Diffusion, T2V, CLIP / VLM, etc.), or having relevant side projects.
- Familiar with the fundamentals of Agents (Tool Use, trajectory, Memory, multi-turn interaction), or having worked on Agent-related projects.
- Good data sense: able to identify data quality issues and willing to be rigorous about data accuracy.
- Good communication skills and a collaborative mindset.
Good To Have:
- Experience with large-scale data processing (from coursework / internships / competitions), having handled TB-PB scale data. Familiarity with data warehouse dimensional modeling (Kimball / OneData approach).
- Experience with model evaluation / benchmarking (exposure to VBench, ELO, GSB, Langfuse, etc. is a plus).