Requirements
- •Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- •2–4 years of experience in data engineering or a backend data-related role
- •Strong skills in Java, Scala, or another backend programming language
- •Python (PySpark/pandas) skills
- •Experience with SQL and distributed data systems (e.g., Spark, Kafka, SQS)
- •Familiarity with NoSQL stores like Cassandra, HBase, or similar
- •Understanding of data modeling for analytics and reporting
- •Proficient in English with strong communication skills — able to explain or demo work to non-engineers
- •Self-driven — picks up new technologies and frameworks with little guidance
- •Debugs methodically; breaks complex problems into smaller steps
- •Open to feedback — iterates and improves through code reviews
- •Uses AI-assisted engineering tools (Claude Code, Codex, Cursor, etc.) as part of daily workflow
- •Reliable internet connection to sustain video, audio, and screen sharing
What You'll Do
- •Build and maintain reliable data pipelines and ETL/ELT workflows
- •Develop and optimize data models for analytics and internal tools
- •Work with team members to deliver clean, trusted datasets
- •Support core data platform tools like Spark and AWS (S3, SNS, SQS, ECS/Fargate, EMR)
- •Monitor data pipelines for quality, performance, and reliability
- •Write clear documentation and contribute to test coverage and CI/CD processes
- •Help shape our data lakehouse architecture and platform roadmap
Nice to Have
- •Experience with dbt, Databricks, or real-time data pipelines
- •Familiarity with cloud infrastructure tools like Terraform or CloudFormation
- •Interest in data governance, ML pipelines, or compliance standards
- •Personal projects or open source contributions demonstrating initiative
Benefits
- •Parental leave
- •Diversity and inclusion working groups
- •Flexible working practices
