Role Overview
We are seeking a Senior Data Pipeline / Graph DB Engineer to design and build the core data layer powering our trade capture and analytics platform. In this role, you will architect high-throughput data ingestion pipelines, real-time transformations, and graph-based data models within a cloud-native Linux environment.
You will be instrumental in mapping complex financial relationships—such as counterparty networks, entity hierarchies, trade dependencies, and reference data—into performant graph and analytical stores.
Key Responsibilities
- Architect and build scalable real-time streaming and batch ingestion pipelines capable of processing high-volume financial transaction feeds and market datasets.
- Design, implement, and maintain graph database models (nodes, edges, properties) to represent complex trading relationships, legal entity structures, risk exposure networks, and reference metadata.
- Develop low-latency microservices and query APIs in Python or Rust for efficient downstream consumption of graph and analytical datasets.
- Profile and optimize graph traversal queries, streaming jobs, and storage engines for maximum throughput, sub-second latency, and strict data reliability.
- Implement continuous data quality checks, schema evolution management, automated data lineage tracking, and real-time pipeline monitoring.
- Actively learn and internalize the structural context and metadata of our trade capture, financial product, and counterparty domain to inform model design.
Required Skills & Experience
- Strong proficiency in Python and/or Rust for scalable data engineering, script automation, and microservice development.
- Proven track record of designing, operating, and maintaining large-scale, fault-tolerant data pipelines in production.
- Hands-on experience with native graph database technologies (e.g., Neo4j, Memgraph, Amazon Neptune, or similar Cypher/Gremlin-based engines).
- Deep understanding of real-time streaming and event-driven data architectures (e.g., Apache Kafka, Redpanda, NATS).
- Strong expertise with Linux environments, cloud-native infrastructure, and containerized deployment (Docker, Kubernetes).
- Proven ability to rapidly grasp complex data structures, schemas, and metadata within specialized domains.
Preferred Qualifications
- Direct experience handling financial market data feeds, trade capture, counterparty reference data, or risk analytics platforms.
- Practical experience with high-performance time-series stores or analytical engines (e.g., ClickHouse, TimescaleDB, InfluxDB).
- Exposure to cloud-native data ecosystem services (e.g., GCP Dataflow, BigQuery, AWS Kinesis, AWS Neptune).