Key job responsibilities
AI-Native Infrastructure & Real-Time Processing
· Design and architect AI-native infrastructure supporting real-time data processing for AI/ML inference, training, and continuous learning at scale
· Lead the development of semantic layers and knowledge graphs enabling intelligent query routing and context-aware data access across the organization
· Architect infrastructure for agentic AI systems with multi-agent orchestration, defining patterns and best practices for the team
· Drive GenAI-powered data quality, entity resolution, and metadata management strategies that raise the bar for data integrity
Data-as-a-Product Delivery
· Own end-to-end accountability for complex data products from ingestion to consumption, defining SLAs and driving adoption across stakeholders
· Lead the delivery of data products with measurable quality metrics, customer satisfaction targets, and continuous improvement mechanisms
· Design and build self-service platforms with embedded governance, lineage, and discovery — enabling teams to independently access and trust data
· Define data contracts and API standards for reliable, versioned data consumption across downstream consumers
AWS Infrastructure & Pipeline Engineering
· Architect and optimize AWS infrastructure: EC2, Lambda, S3, Redshift, EMR — balancing performance, reliability, and cost
· Design high-throughput, fault-tolerant pipelines supporting analysts, data scientists, and AI agents at global scale
· Lead implementation of CDC and event-driven architectures for sub-minute data availability with end-to-end observability
· Drive infrastructure-as-code best practices using CDK, establishing reusable patterns and deployment standards
Technical Leadership & Operational Excellence
· Mentor Data Engineer I team members, conducting code reviews and elevating engineering standards
· Own operational health of critical pipelines, driving root cause analysis and long-term prevention strategies
· Author and maintain technical design documents, influencing architectural decisions across the team
· Lead peak readiness efforts and incident response, ensuring system reliability during high-traffic events
· Drive automation initiatives that reduce operational toil and enable non-linear scaling
Basic qualifications
- 3+ years of data engineering experience- Experience with data modeling, warehousing and building ETL pipelines
- Experience with SQL
- Bachelor's degree
Preferred qualifications
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions- Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases)
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.