Search Jobs

Search by job, company or skills

AI Research Engineer, Computer Vision & VLMs

AI Research Engineer, Computer Vision & VLMs

zapdos labs
3-5 Years
  • Posted 13 hours ago
  • Be among the first 10 applicants

Job Description

Zapdos Labs builds AI for the physical world, starting with industrial sites. Zapdos runs on the CCTV a site already owns.

Understanding a busy warehouse, factory floor, loading dock or marine yard means making sense of people, vehicles, objects, activities and events as they change over time, despite occlusion, dust, changing lighting, night-time IR, fisheye and low-resolution feeds, and cameras that were mounted for security, not for AI.

We are looking for an AI Research Engineer with a strong research background in computer vision and vision-language models (VLMs) to develop the visual intelligence behind Zapdos. You will work on image and video understanding, spatiotemporal reasoning, and multimodal models that connect what a camera sees to a useful alert, a corrective action, or a record in a real industrial environment.

This role combines research depth with ownership of working systems. You will formulate research questions, build datasets, train and evaluate models, and partner with product and engineering to bring successful approaches into production across customer sites in Singapore and the United States. Researchers and engineers from autonomous driving, robotics, embodied AI, and related perception fields are especially encouraged to apply.

What you'll own

  • Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and understanding events across video: forklift-pedestrian proximity, zone and access breaches, loading dock activity, housekeeping and blocked egress, PPE, falls and unsafe behaviour, and the security, quality and operations scenarios that run on the same cameras.
  • Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction grounded in observable evidence, so every alert can be explained by what is in the clip.
  • Design training and adaptation strategies, including supervised fine-tuning, representation learning, distillation, and domain adaptation, based on measurable product needs.
  • Build representative image and video datasets, annotation workflows, and evaluation sets that capture difficult edge cases across sites, camera types and shifts, while protecting sensitive customer footage.
  • Create rigorous experiments and benchmarks that measure perception quality, temporal consistency, hallucinations, false-alert rate, robustness, latency, and cost across locations and operating conditions.
  • Diagnose failures caused by occlusion, lighting changes, camera placement, rare events, and domain shift between sites; use those findings to improve data and models.
  • Partner with infrastructure and product engineers to deploy efficient inference pipelines, on-premise and in the cloud, with monitoring, quality gates, staged rollouts, and rollback paths.
  • Translate advances in computer vision, VLMs, and embodied AI into practical product capabilities, and communicate the evidence and tradeoffs behind your decisions.
  • Raise research and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.

Requirements

  • 3+ years of research or applied development experience in computer vision, multimodal learning, or a closely related field; relevant graduate research counts toward this experience.
  • A demonstrated research track record in computer vision or vision-language modeling, through publications, substantial research projects, open-source contributions, or research delivered in industry.
  • Strong foundations in deep learning, visual representation learning, and experimental design, with depth in areas such as video understanding, detection and tracking, visual grounding, or multimodal reasoning.
  • Hands-on experience training, fine-tuning, or adapting computer vision models, and developing or evaluating VLMs beyond basic API integration.
  • Strong Python skills and experience with PyTorch or an equivalent deep learning framework, along with modern training and evaluation tooling.
  • Experience building datasets, designing reliable evaluations, analyzing model failures, and using ablations to understand what drives improvements.
  • Strong software engineering judgment and the ability to turn research code into reproducible, tested systems that other engineers can use.
  • Ability to connect modeling choices to product constraints including latency, cost, privacy, reliability, and the supervisor's experience on the floor.
  • Comfort working through ambiguity in an early-stage team, and collaborating across research, engineering, product and customers.

Especially relevant experience

  • A PhD or research-focused master's degree in computer vision, machine learning, robotics, or a related field, or equivalent research experience.
  • Industry research or engineering experience in autonomous driving, robotics, embodied AI, video analytics, or other applications of perception in the physical world.
  • Publications at venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, CoRL, ICRA, or RSS.
  • Experience with monocular or multi-camera video perception, spatial understanding, long-video reasoning, or learning from limited and noisy labels.
  • Experience shipping vision models under real-time constraints on edge or on-premise hardware, including model compression, distillation, quantization, or inference optimization (NVIDIA GPUs, TensorRT, or similar).
  • Experience with existing CCTV and VMS ecosystems (RTSP, ONVIF, Hikvision, Dahua, Genetec, or similar).

When applying, please include links to relevant publications, research projects, or code, and briefly describe your own contribution.

Benefits

  • Competitive salary and employee stock option plan.
  • Medical, dental, and vision benefits as applicable.
  • Paid time off and public holidays.
  • Learning and development support, including conference attendance.

Work directly with the founders and with live customer sites; see your models running on real cameras within weeks.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

NVIDIA GPUs

ONVIF

model compression

Genetec

Dahua

inference optimization

image and video understanding

quantization

visual representation learning

Hikvision

activity recognition

About Company