Job Description:
- Working around AI Agent architecture and Harness / Loop Engineering / Post-Training / Agentic RL, you will participate in or lead the following:
- AI Agent Architecture and Harness (Agent Architecture and Runtime Framework)
- Design and iterate on AI Agent architectures (Planning / Tool Use / Memory / Multi-Agent coordination, etc.), and build toolsets, context engineering, SOP/knowledge integration, and other Harness components to fully utilize foundation model capabilities in customer service scenarios.
- Improve Agent system efficiency through model layering, parallel scheduling, context compression, and cache reuse, reducing end-to-end latency and inference cost.
- Loop Engineering (Data and Training Self-Evolution Loop)
- Build a self-evolution loop: real online feedback (CSAT, escalation to human, human rewrites, etc.) automated evaluation training data generation Post-Training / RL gray-scale validation, enabling the system to continuously improve with real traffic.
- With the goal of reaching and surpassing human customer service, establish an automated evaluation and LLM-as-a-Judge framework to measure performance using both offline metrics and business metrics.
- Post-Training and Agentic RL
- Design post-training strategies for customer service scenarios, including SFT / DPO / GRPO / PPO, constructing preference and reward signals based on explicit and implicit feedback (user ratings, CSAT, escalations, human rewrites, etc.).
- Design Agentic RL objectives and training methods around key Agent decision points (action selection, tool calling, clarification/follow-up strategies, etc.), exploring the application of RLHF / RLAIF in dialogue policies.
- Cutting-Edge Practices
- Continuously track the latest research advances in Post-Training, RL, Agents, and Harness / Loop Engineering, and drive their adoption in business applications.
Requirements:
- Master's degree or above in Computer Science, Artificial Intelligence, or a related field
- At least 3 years of Algorithm Engineering / Machine Learning experience are welcome
- Strong computer science and mathematics foundation, with solid algorithm and data structure skills.
- Systematic knowledge of machine learning and deep learning fundamentals, with a deep understanding of Transformer and LLM architectures.
- Substantial research or project experience in at least one or two of the following areas:
- Post-Training (SFT / DPO / GRPO, etc.)
- Reinforcement Learning / Agentic RL
- Agent direction (Planning / Reasoning / Tool Use / Memory / Multi-Agent, etc.)
- Harness / Loop Engineering (Agent runtime frameworks, automated evaluation, and data self-evolution loops).
- Proficient in Python and familiar with mainstream deep learning frameworks such as PyTorch / TensorFlow.
- Interest in cross-disciplinary problems spanning research, engineering, and business, with willingness to validate and refine methods in real-world scenarios.
Preferred Qualifications:
- Experience working on Agent or LLM projects at leading AI teams (OpenAI, Google DeepMind, Anthropic, Meta, ByteDance Seed, Alibaba Tongyi, Tencent Hunyuan, Baidu Wenxin, etc.).
- Experience designing and scaling AI Agent systems, or practical experience improving Agent inference efficiency (model distillation, model routing, parallel scheduling, context engineering, etc.).
- Publications at top-tier conferences/journals in NLP, LLM, RL, Agent, or related fields (e.g., ACL, EMNLP, NeurIPS, ICLR, ICML), or complete research work.
- Hands-on experience participating in LLM Post-Training / Reinforcement Learning (including Agentic RL) projects.
- Practical experience applying RLHF / Agentic RL in dialogue/customer service/Agent scenarios.
- Familiarity with distributed training and large model optimization (e.g., DeepSpeed, FSDP, Megatron, etc.).