T
Senior AI Engineer Document Intelligence LLM Infrastructure USA Shift
T
Senior AI Engineer Document Intelligence LLM Infrastructure USA Shift
tomorrow world technologies (twt)6-8 Years
Early Applicant
- Posted 10 days ago
- Be among the first 10 applicants
Job Description
Job Title – Senior AI Engineer Document Intelligence & LLM Infrastructure
Nature: Contract
Time Zone : US Shift
Working Hours : 5.30pm – 2.30am
Work Location :Remote
Exp : 6 to 8 +Years
Contract Duration : 6 Months
No of Position : 1
Role Overview
We are hiring a senior individual contributor to own and scale large-scale, OCR-driven document intelligence systems powered by self-hosted LLMs.
This is a deeply hands-on engineering role focused on production systems that process long-form documents (200+ pages), extract structured data deterministically, and run on optimized GPU-backed inference infrastructure.
You will work closely with the AI leadership team but will independently own architecture, performance, and reliability of document processing pipelines.
Core Responsibilities
Own PDF ingestion, layout-aware parsing, and multi-page document assembly
Implement robust chunking, segmentation, and metadata tracking across long documents
Handle exception detection, retries, and deterministic failure handling
Optimize systems to reliably process 200+ page documents at scale
Build layout-aware extraction systems using bounding boxes and structural metadata
Implement deterministic schema validation and cross-field consistency checks
Reduce reliance on manual QA through rule-based validation layers
Ensure traceability from extracted field back to source span
Own production reliability of LLM serving infrastructure
Implement schema enforcement, rule engines, invariants, and rejection logic
Build automated exception routing without default human review
Ensure auditability and reproducibility of extraction results
Create measurable correctness guarantees for high-stakes use cases
Collaborate with cross-functional teams to ship production-grade AI systems
Required Experience
6+ years of hands-on Python engineering
Proven production experience building OCR-driven document pipelines
Experience handling long-form PDFs (100+ pages)
Strong Experience With
Strong debugging and systems-level thinking
Ability to clearly articulate system trade-offs and business impact
Strongly Preferred
Experience with layout-aware models (LayoutLM, DocFormer, vision-language models)
Experience optimizing GPU cost and inference performance
Experience in regulated domains (healthcare, finance, compliance)
Familiarity with document-heavy workflows such as loan processing, underwriting, or claims
Nature: Contract
Time Zone : US Shift
Working Hours : 5.30pm – 2.30am
Work Location :Remote
Exp : 6 to 8 +Years
Contract Duration : 6 Months
No of Position : 1
Role Overview
We are hiring a senior individual contributor to own and scale large-scale, OCR-driven document intelligence systems powered by self-hosted LLMs.
This is a deeply hands-on engineering role focused on production systems that process long-form documents (200+ pages), extract structured data deterministically, and run on optimized GPU-backed inference infrastructure.
You will work closely with the AI leadership team but will independently own architecture, performance, and reliability of document processing pipelines.
Core Responsibilities
- Large-Scale Document Intelligence Pipelines
Own PDF ingestion, layout-aware parsing, and multi-page document assembly
Implement robust chunking, segmentation, and metadata tracking across long documents
Handle exception detection, retries, and deterministic failure handling
Optimize systems to reliably process 200+ page documents at scale
- OCR & Structured Extraction Systems
Build layout-aware extraction systems using bounding boxes and structural metadata
Implement deterministic schema validation and cross-field consistency checks
Reduce reliance on manual QA through rule-based validation layers
Ensure traceability from extracted field back to source span
- Self-Hosted LLM Inference (Production Ownership)
- vLLM
- Hugging Face TGI
- GPU-backed serving stacks
- KV cache management
- Batching
- Context window control
- Throughput vs latency trade-offs
Own production reliability of LLM serving infrastructure
- Deterministic Validation & Control Systems
Implement schema enforcement, rule engines, invariants, and rejection logic
Build automated exception routing without default human review
Ensure auditability and reproducibility of extraction results
Create measurable correctness guarantees for high-stakes use cases
- Production Engineering & Scale
- Large document volumes
- Concurrency
- Failure states
- Observability and monitoring
Collaborate with cross-functional teams to ship production-grade AI systems
Required Experience
6+ years of hands-on Python engineering
Proven production experience building OCR-driven document pipelines
Experience handling long-form PDFs (100+ pages)
Strong Experience With
- vLLM or Hugging Face TGI
- GPU-based LLM serving
- Open-source LLMs (LLaMA, Qwen, Mistral, etc.)
Strong debugging and systems-level thinking
Ability to clearly articulate system trade-offs and business impact
Strongly Preferred
Experience with layout-aware models (LayoutLM, DocFormer, vision-language models)
Experience optimizing GPU cost and inference performance
Experience in regulated domains (healthcare, finance, compliance)
Familiarity with document-heavy workflows such as loan processing, underwriting, or claims
