Actively recruiting / 6 applicants
We’re here to help you
Juliana Torrisi is in direct contact with the company and can answer any questions you may have. Email
Juliana Torrisi, RecruiterRole Overview
We are looking for a skilled Senior Data Engineer to transform detection designs into robust production pipelines. You will collaborate closely with researchers to implement techniques into scalable Spark jobs, ensuring they operate effectively over extensive data lake tables. Your role is crucial in maintaining data integrity and quality, and in explaining analytical outcomes.
Responsibilities
- Develop and maintain Spark jobs that process large volumes of data, transforming prototype techniques into reliable production systems.
- Ensure data quality by verifying incoming data against specified schemas and identifying anomalies in output data.
- Manage open table formats like Delta Lake, Iceberg, or Hudi, focusing on merges, incremental processing, compaction, and table maintenance.
- Implement event-time correctness, understanding the differences and impacts of processing time versus event time.
- Convert prototypes into supervised-free systems, ensuring they run smoothly in production environments.
- Utilize Infrastructure as Code and CI/CD practices with tools like Crossplane, Argo CD, and Argo Workflows, although prior experience with these specific tools is not required.
Required Skills
- Extensive experience with production-level Spark, including tuning for skew, partitioning, shuffle behavior, and join strategy.
- Proficiency in handling open table formats and managing their maintenance.
- Strong understanding of event-time versus processing-time correctness.
- Proven ability to transform prototypes into unsupervised, production-ready systems.
- Familiarity with Infrastructure as Code and CI/CD pipelines.
Nice to Have
- Comfort in environments where correctness is nuanced and not binary.
- Security background is not required but can be beneficial.
Clarification