Fizibox is looking for a VLM Inference Optimization Specialist to work with our engineering team in Da Nang.
This role is suitable for someone who has hands-on experience with AI inference, model deployment, CUDA, OpenVINO, TensorRT, ONNX Runtime, LLM/VLM serving, or edge AI systems.
What You'll Work OnYou will work directly with the Fizibox engineering team on:
- Reviewing the current VLM inference pipeline.
- Benchmarking performance across different runtime configurations, including OpenVINO, CUDA, and hybrid backend setups.
- Measuring and analyzing:
- End-to-end latency.
- Time to first token
- Token throughput.
- RAM / VRAM usage
- CPU / GPU utilization.
- Stability during long-running inference.
- Identifying bottlenecks in model loading, memory transfer, batching, queueing, and runtime scheduling.
- Improving serving architecture, including persistent workers, backend routing, fallback behavior, and observability.
- Checking output consistency across model versions, precision modes, and backend configurations.
- Helping the team define practical benchmark scripts, runtime checklists, and release validation steps.
What We're Looking ForYou do not need to know every technology listed below. We are looking for someone with strong hands-on experience in AI inference and the ability to reason deeply about runtime performance.
Relevant experience includes one or more of:
- CUDA
- OpenVINO
- TensorRT / TensorRT-LLM
- ONNX Runtime
- PyTorch inference
- vLLM
- llama.cpp
- Triton Inference Server
- LLM / VLM / multimodal model deployment
- Model conversion, quantization, profiling, or production inference serving
Nice to Have- Experience with vision-language models or multimodal AI systems.
- Experience with edge AI, camera streams, or real-time video processing.
- Experience optimizing models for limited compute environments.
- Familiarity with profiling tools such as Nsight, OpenVINO Benchmark Tool, nvidia-smi, perf, or similar.
- Experience comparing model output after conversion or quantization.
Engagement Model- Part-time, contract, or project-based.
- Based in Da Nang.
- Hybrid working arrangement: remote work is possible, with onsite sessions when hardware testing or direct collaboration is needed.
- Scope can start with a short technical engagement and expand based on mutual fit.
Why Join This Project- The system is already running; this is not a from-scratch research project.
- You will work on a practical VLM inference pipeline involving real hardware, real latency constraints, and real deployment requirements.
- The work is focused, technical, and outcome-driven.
- Flexible engagement model: part-time, milestone-based, or consulting arrangement.
How to ApplyPlease apply with:
- CV, LinkedIn, or GitHub profile.
- A short note about relevant AI inference, model deployment, or runtime optimization work you have done.
- Technologies you have worked with, such as CUDA, OpenVINO, TensorRT, ONNX Runtime, vLLM, llama.cpp, or similar.
- Your preferred engagement model: part-time, project-based, or consulting.