Databricks Developer with AWS
Role description
Databricks Developer with AWS
Location: Toronto, ON- Work Mode: Hybrid – Tuesday, Wednesday & Thursday onsite
Contract: 12 Months contract
Experience: 7+ Years
Job Summary
We are seeking an experienced Databricks Developer with strong AWS expertise to join a data engineering team supporting enterprise-scale data platforms for a leading financial services/investment management organization.
The ideal candidate will have 7+ years of experience in data engineering , with strong hands-on expertise in Databricks, Apache Spark, PySpark, Python, SQL, and AWS cloud services . The candidate will be responsible for designing, developing, optimizing, and maintaining scalable data pipelines and data processing solutions.
This role requires someone who can work effectively in a large enterprise environment , collaborate with data architects, business stakeholders, developers, and QA teams, and follow strong standards around data quality, security, performance, and governance.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Databricks and Apache Spark .
- Develop complex data transformations using PySpark, Python, and SQL .
- Build and optimize data pipelines on AWS cloud platforms .
- Work with AWS services such as S3, Glue, Lambda, EMR, Redshift, IAM , and related data services.
- Develop data processing solutions using Databricks notebooks, workflows, jobs, clusters, and Delta Lake .
- Implement reliable and reusable data ingestion frameworks for structured and semi-structured data.
- Perform data cleansing, transformation, validation, aggregation, and enrichment.
- Work with Delta Lake and implement efficient storage and data management strategies.
- Optimize Databricks/Spark jobs for performance, scalability, and cost efficiency .
- Troubleshoot production data pipeline failures and resolve data quality or performance issues.
- Implement appropriate error handling, logging, monitoring, and recovery mechanisms.
- Collaborate with data architects to implement enterprise data architecture and engineering standards.
- Participate in data modeling and development of analytical data structures.
- Work with large-volume datasets and implement distributed data processing solutions.
- Ensure data pipelines meet data quality, security, availability, and governance requirements.
- Participate in code reviews and follow established development standards.
- Develop and maintain technical documentation for data pipelines and processes.
- Work closely with QA teams to support data validation and testing.
- Support deployment of Databricks and AWS solutions across development, test, and production environments.
- Participate in Agile ceremonies including sprint planning, daily stand-ups, reviews, and retrospectives.
Required Technical Skills
Databricks
- Strong hands-on experience with Databricks .
- Databricks notebooks, jobs, workflows, clusters, and job scheduling.
- Experience with Delta Lake .
- Experience optimizing Databricks/Spark workloads.
- Understanding of Databricks data engineering best practices.
AWS
Strong hands-on experience with AWS data services, including:
- AWS S3
- AWS Glue
- AWS Lambda
- AWS EMR
- AWS Redshift
- AWS IAM
- Experience with AWS-based data architecture and cloud data pipelines.
Programming & Data Processing
- Strong Python / PySpark development experience.
- Strong Apache Spark knowledge.
- Advanced SQL skills.
- Experience working with large-scale datasets and distributed processing.
Data Engineering
- Strong experience developing ETL/ELT pipelines .
- Experience with batch and/or near-real-time data processing.
- Data ingestion from multiple source systems.
- Data transformation and data quality validation.
- Experience working with structured and semi-structured data.
- Understanding of data warehousing and modern data lake/lakehouse architectures.
Preferred Skills
- Experience with Databricks Unity Catalog .
- Experience with AWS Glue Data Catalog .
- Knowledge of Delta Live Tables / Lakeflow .
- Experience with CI/CD for Databricks.
- Git/GitHub or similar source-control systems.
- Experience with Terraform or Infrastructure as Code .
- Knowledge of DevOps practices.
- Experience with data governance and data security.
- Experience with enterprise data platforms in the financial services, banking, or investment management domain .
- Familiarity with Agile/Scrum development methodologies.
Financial Services / Enterprise Experience
Experience working in a large financial services, banking, investment management, or capital markets environment is highly desirable.
Candidates should be comfortable working with enterprise-level data, security requirements, regulatory considerations, data governance standards, and production-critical data pipelines.
Education & Experience
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline preferred.
- 7+ years of overall IT/data engineering experience.
- Strong recent hands-on experience with Databricks and AWS .
- Proven experience developing and supporting enterprise-scale data pipelines.