Senior /Master Data Developer in Brazil at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior /Master Data Developer based in Brazil.
This senior role is central to a large-scale modernization of data pipelines, helping transform legacy Databricks workloads into a modern, cloud-native architecture on Google Cloud. You will work hands-on with complex data processing logic, refactoring legacy code into scalable ELT patterns while maintaining critical downstream integrations. The position combines advanced Python, PySpark, SQL, and Google Cloud data engineering with strong data quality and validation practices. You will build scalable ingestion pipelines and help maintain reliable data contracts across Raw, Trusted Core, and Gold layers. Close collaboration with Data Stewards and business stakeholders will be essential to validate migrations and manage dependencies. This is an opportunity to play a key technical role in a major data transformation initiative within a collaborative, engineering-focused environment.
- Legacy Code Refactoring: Translate complex Databricks notebook logic into modern ELT patterns, using declarative SQL-based transformations for standard use cases and Python-based distributed processing for more intricate logic, including RDD-based operations.
- Data Contract Preservation: Maintain backward compatibility throughout migrations by implementing trusted-core layers and reverse-view strategies that protect downstream systems, integrations, and production dashboards from breaking changes.
- Ingestion Pipeline Development: Build scalable ingestion pipelines by parameterizing YAML configurations that feed automated DAG-generation processes, using workflow orchestration and standardized processing templates.
- Data Quality & Testing: Implement synchronous and unit-level assertions within transformation layers to identify null values, enforce key uniqueness, and validate domain-specific data requirements.
- Migration Validation: Conduct technical validation exercises by reconciling data between legacy environments and the new platform, performing record-level and field-level parity checks to confirm migration accuracy.
- GitOps & CI/CD: Follow structured GitOps development practices, including feature branches, code reviews, pull requests, and automated deployment pipelines across development and production environments.
- Stakeholder Collaboration: Work closely with business Data Stewards and technical teams to understand integration dependencies, negotiate refactoring scope, and validate migrated data and outcomes.
- Architecture Modernization: Contribute to the evolution of data processing practices by applying cloud-native engineering principles and helping transition tightly coupled legacy workloads into modular, maintainable components.
- Data Engineering Expertise: Strong hands-on technical expertise in Python, PySpark, and advanced SQL, particularly for large-scale and distributed data processing.
- Migration & Refactoring: Proven experience migrating data lakes and refactoring legacy data-processing code, including the ability to reverse-engineer tightly coupled logic from platforms such as Databricks or Azure Data Factory.
- Google Cloud: Practical mastery of the Google Cloud data engineering ecosystem, particularly BigQuery, Dataform, and Dataproc Serverless.
- Data Architecture: Strong understanding of Medallion Architecture, including Raw/Bronze, Silver/Trusted Core, and Gold layers, as well as analytical modeling approaches such as Star Schema and One Big Table.
- Orchestration & Automation: Hands-on experience with modern GitOps development practices, including code reviews and pull requests, combined with orchestration using Airflow or Composer.
- Data Quality: Mature understanding of data quality practices and experience implementing testing directly within data transformation pipelines.
- Performance Optimization: Ability to analyze and optimize distributed data workloads, with strong attention to scalability, reliability, and processing efficiency.
- Collaboration: Consultative and collaborative communication style, with the ability to align cross-team dependencies, negotiate technical scope, and validate outcomes with business Data Stewards.
- Generative AI: Experience using Generative AI tools such as Gemini, Vertex AI, or Claude to support PySpark-to-SQL refactoring, dependency analysis, documentation, or other data engineering automation is a plus.
- Additional Data Engineering Expertise: Experience with Change Data Capture (CDC), event-driven ingestion, BigQuery performance tuning, partitioning, execution cost optimization, and Data Vault versus Star Schema modeling approaches is advantageous.
- Location Requirement: Candidates residing in the Campinas Metropolitan Region are expected to work from the local offices in accordance with the applicable attendance policy.
- Health Coverage: Health and dental insurance.
- Food Allowance: Food and meal allowances.
- Family Support: Childcare assistance and extended parental leave.
- Wellness: Partnerships with gyms and health and wellness professionals through Wellhub (Gympass) and TotalPass.
- Profit Sharing: Participation in a profit-sharing and results program (PLR).
- Life Insurance: Life insurance coverage.
- Continuous Learning: Access to a dedicated continuous learning platform and professional development resources.
- Online Learning: Partnerships with online course platforms to support ongoing skill development.
- Language Learning: Access to a dedicated language-learning platform.
- Discounts: Employee discount club with partner offers.
- Wellbeing Resources: Free access to an online platform focused on physical health, mental wellbeing, and overall wellness.
- Family Development: Pregnancy and responsible parenting courses.
- Inclusive Workplace: Dedicated health and wellbeing teams, inclusion specialists, and affinity groups providing support throughout the employee journey.
- Professional Growth: Opportunities to work on large-scale cloud and data modernization initiatives while developing expertise across modern data engineering technologies.