Data Engineer II, CMT
Amazon
As a Data Engineer II on the Data Engineering team, you will own and drive the design of scalable data systems that power Amazon's pricing and catalog optimization. You'll lead technical initiatives at the intersection of data engineering and AI, architecting solutions that process large-scale data across Amazon's global operations. You'll mentor junior engineers, influence technical direction, and deliver high-impact data products with measurable business outcomes.
Key job responsibilities AI-Native Infrastructure & Real-Time Processing
- Design and architect AI-native infrastructure supporting real-time data processing for AI/ML inference, training, and continuous learning at scale
- Lead the development of semantic layers and knowledge graphs enabling intelligent query routing and context-aware data access across the organization
- Architect infrastructure for agentic AI systems with multi-agent orchestration, defining patterns and best practices for the team
- Drive GenAI-powered data quality, entity resolution, and metadata management strategies that raise the bar for data integrity
Data-as-a-Product Delivery
- Own end-to-end accountability for complex data products from ingestion to consumption, defining SLAs and driving adoption across stakeholders
- Lead the delivery of data products with measurable quality metrics, customer satisfaction targets, and continuous improvement mechanisms
- Design and build self-service platforms with embedded governance, lineage, and discovery — enabling teams to independently access and trust data
- Define data contracts and API standards for reliable, versioned data consumption across downstream consumers
AWS Infrastructure & Pipeline Engineering
- Architect and optimize AWS infrastructure: EC2, Lambda, S3, Redshift, EMR — balancing performance, reliability, and cost
- Design high-throughput, fault-tolerant pipelines supporting analysts, data scientists, and AI agents at global scale
- Lead implementation of CDC and event-driven architectures for sub-minute data availability with end-to-end observability
- Drive infrastructure-as-code best practices using CDK, establishing reusable patterns and deployment standards
Technical Leadership & Operational Excellence
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

