Search by job, company or skills

Senior AI Engineer (Computer Vision)

Senior AI Engineer (Computer Vision)

Masan Group
3-5 Years
Not Disclosed
  • Posted 16 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

Masan Group's AI team is extending our multi-agent AI platform — currently text-based and in production — into vision. You will own the visual understanding layer of the agentic system: enabling agents to see, interpret, and act on images and video from Masan's retail and consumer operations (thousands of WinCommerce stores, manufacturing, logistics, consumer touchpoints).

The platform is established and in production; vision is its next frontier. You will shape the foundational technology choices for how vision integrates into our agent stack — new problems and fast project cycles, on top of a proven foundation.

What You Will Do

  • Build vision capabilities into the multi-agent platform: integrate and fine-tune vision-language models (VLMs) so agents can reason over images and video as naturally as text.
  • Solve classic and modern CV problems in real business contexts: OCR/document understanding, product/shelf recognition, detection and tracking, visual QA — chosen per project, not a fixed menu.
  • Design the pipeline from raw visual data to agent-consumable signals: preprocessing, model serving, cost/latency optimization at retail scale.
  • Prototype fast, then productionize: feasibility study → POC → deployed, monitored service.
  • Contribute to the shared agentic platform (evaluation, observability, tools) beyond the vision specialty.

Must-Have

  • 3+ years in ML/AI engineering, with strong hands-on computer vision experience: training/fine-tuning and deploying CV models (detection, OCR, classification, segmentation) and familiarity with modern VLMs (e.g., open-source multimodal LLMs or commercial multimodal APIs).
  • Experience building on LLM/agentic systems (function calling, RAG, multi-agent frameworks) or strong evidence of ability to ramp quickly.
  • Strong Python and production ML engineering: model serving, GPU inference optimization, data pipelines for image/video.
  • Comfort with ambiguity: able to take a new capability from concept to production without a playbook.

Strong Plus

  • Video understanding, edge deployment, or large-scale visual data pipelines.
  • Retail/e-commerce CV applications (shelf analytics, planogram compliance, visual search).
  • Data engineering fundamentals (image/video dataset curation and labeling ops at scale).
  • Contributions to open-source CV or multimodal projects.

More Info

Key Skills

detection

multi-agent frameworks

vision-language models (VLMs)

multimodal LLMs

large-scale visual data pipelines

multimodal APIs

GPU inference optimization

RAG

edge deployment

LLM agentic systems

function calling

model serving

video understanding

About Company

Similar Jobs

5-7 yrs
Ho Chi Minh, Vietnam
Skills:
CI/CD tools (Jenkins, GitHub Actions), Bash Shell Scripting, Docker, Linux, Python, GitLab CI, Computer Vision evaluation metrics