AI Engineer – LLM Algorithm Engineer (Agentic Commerce)
Job Description
Job Description:
- Lead continuous pre-training and post-training of large language models (Dense and MoE architectures) for vertical business domains, including domain-specific data synthesis, multi-stage data curation, and monthly iteration to support downstream business integration.
- Design and build autonomous agents (e.g., Plan-and-Execute agents) that automatically retrieve, extract, and synthesize domain knowledge from large-scale multilingual/unstructured data sources, producing high-confidence corpora for continued pre-training.
- Develop and iterate RL-based post-training methods (e.g., CPT, SFT, DAPO) to improve model performance on domain-specific QA, reasoning, and knowledge-modeling tasks.
- Build and optimize multimodal LLM applications for content generation and understanding, including product copywriting, semantic search recall, and business potential/CTR prediction.
- Design end-to-end agent workflows integrating query understanding, content generation, validation, and downstream product/search systems.
- Build generative representation-learning pipelines (e.g., token-based generative pre-training on behavioral sequences) to produce reusable embeddings for downstream ranking and recommendation models.
- Collaborate cross-functionally to translate business requirements into scalable model training and agent system design continuously monitor production metrics to guide model and agent optimization.
Requirements:
- Master's degree or above in Computer Science, Natural Language Processing, Artificial Intelligence, Electrical and Electronics Engineering, Signal Processing or a related field.
- Minimum 5 years of full-time industry experience in LLM/ML algorithm engineering, including hands-on experience with large-scale model continuous pre-training (Dense and/or MoE architecture) for vertical business domains.
- Hands-on experience designing and building autonomous agent systems (e.g., using LangGraph or similar frameworks) for automated knowledge retrieval, extraction, and synthesis pipelines, combined with hands-on experience in LLM post-training techniques (SFT, RL-based methods such as DAPO) and data-mixture optimization techniques for pre-training data curation.
- Experience fine-tuning and deploying multimodal large models for content generation, semantic recall, or business potential prediction, with demonstrated production impact on key business metrics (e.g., CTR, conversion).
- Good programming and engineering skills solid foundation in classic ML techniques (e.g., LightGBM, XGBoost, Bayesian modeling) is a plus.
- Good problem-solving skills able to independently drive projects from research through to large-scale production deployment.
- Prior experience in multimodal retrieval/recommendation systems (e.g., cross-modal contrastive learning for content matching) will be a strong plus.
More Info
Job Type:
Industry:
Function:
Employment Type:
Key Skills
DAPO
domain-specific data synthesis
SFT
data-mixture optimization techniques
multilingual unstructured data sources
LightGBM
token-based generative pre-training
Plan-and-Execute agents
multi-stage data curation
MoE architectures
LangGraph
Bayesian modeling
CTR prediction
generative representation-learning pipelines
RL-based post-training methods
large language models
multimodal LLM applications
autonomous agents
