Databricks Data Engineer SME - Clearance Required in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Databricks Data Engineer SME - Clearance Required based in the United States.
This is a senior technical leadership opportunity supporting a mission-critical healthcare revenue-cycle transformation within a federal environment.
You will lead the design and implementation of a scalable Databricks data foundation that connects complex healthcare and financial data across the full revenue lifecycle.
The role combines enterprise data architecture, engineering, governance, data quality, and visualization to create trusted, reusable data products.
You will transform fragmented source data into governed Bronze, Silver, and Gold layers designed for operational dashboards, KPI reporting, and executive decision-making.
Working closely with data scientists, analytics engineers, healthcare SMEs, and technical stakeholders, you will translate operational workflows into actionable data architectures.
You will also establish engineering standards, quality controls, security practices, and reusable patterns while mentoring other data engineers.
This role is ideal for a Databricks expert who enjoys solving complex data challenges in highly regulated, security-conscious environments.
- Lead the technical design and implementation of the revenue-cycle data architecture within Databricks and the War Data Platform.
- Assess and remediate fragmented or poorly structured data environments, addressing duplicate entities, inconsistent definitions, missing relationships, weak keys, data-quality issues, and reconciliation gaps.
- Design a canonical revenue-cycle data model connecting patients, encounters, appointments, eligibility, authorization, providers, facilities, documentation, diagnoses, procedures, coding, charges, claims, payers, adjudication, payments, denials, accounts receivable, appeals, follow-up, and recovery opportunities.
- Design and implement scalable Bronze, Silver, and Gold medallion architectures using Databricks and Delta Lake.
- Build Gold data products and analytical marts optimized for operational dashboards, executive reporting, financial analysis, revenue recovery, and drill-down workflows.
- Develop and maintain Databricks SQL and AI/BI dashboards that provide actionable insights across operational, financial, revenue-recovery, and executive use cases.
- Create reusable semantic datasets and optimized SQL structures supporting consistent KPI calculations across enterprise, network, facility, department, provider, encounter, and claim levels.
- Optimize dashboard queries, Gold datasets, storage, partitioning, clustering, compute usage, refresh schedules, and concurrent-user performance.
- Configure data access and dashboard security in accordance with Unity Catalog permissions, row-level and column-level security, PHI/PII requirements, and user roles.
- Profile government-provided Bronze datasets, develop source-to-target mappings, and normalize and conform data from approved healthcare and financial sources.
- Establish reusable Silver-layer entities with common keys, standardized timestamps, enterprise reference dimensions, and consistent business definitions.
- Develop Gold products supporting revenue-cycle KPIs, coding-audit financial impact, charge completeness, claims readiness, days-to-bill, denial management, payer performance, payment and remittance reconciliation, underpayment detection, AR aging, revenue recovery, and audit traceability.
- Design data structures supporting a Revenue Opportunity Ledger and associated recovery work queues.
- Implement deterministic and, where appropriate, probabilistic entity-resolution approaches for connecting encounters, claims, charges, payments, providers, payers, and related entities across disparate systems.
- Develop and maintain machine-readable data contracts based on the Open Data Contract Standard (ODCS).
- Implement automated quality controls covering completeness, uniqueness, referential integrity, schema conformity, temporal integrity, code validity, financial reconciliation, and source-to-target consistency.
- Establish automated Bronze-to-Silver and Silver-to-Gold quality gates and reconciliation controls across billed, allowed, paid, adjusted, patient-responsibility, AR, and recovery values.
- Configure and manage Unity Catalog, including catalog, schema, and table structures, ownership, classification, tagging, lineage, security, and access controls.
- Support enterprise metadata federation and data-governance requirements.
- Develop production-grade Databricks jobs, workflows, orchestration, monitoring, alerting, and error-handling processes.
- Use GitLab for version control of pipeline code, transformations, data contracts, configurations, and infrastructure-as-code, while supporting automated CI/CD and controlled production promotion.
- Partner with Data Scientists and Analytics Engineers to support predictive modeling, payer analytics, recovery scoring, anomaly detection, and visualization requirements.
- Collaborate with Revenue Cycle and Coding SMEs to translate healthcare workflows into executable data structures, business rules, and analytical products.
- Produce source-to-target mappings, data dictionaries, lineage documentation, architecture diagrams, runbooks, and sustainment materials.
- Mentor Data Engineers and establish reusable Databricks engineering, data modeling, governance, and visualization patterns across the environment.
Requirements:
- Active SECRET security clearance is required.
- Bachelor’s degree with 10+ years of experience in enterprise data engineering, data architecture, data platform development, or a related discipline.
- Senior or SME-level hands-on expertise with Databricks.
- Demonstrated experience building Databricks visualization and dashboard solutions, including Databricks SQL and/or AI/BI dashboards.
- Strong ability to design the Gold and semantic data structures required for enterprise-scale dashboards, operational reporting, and interactive drill-down.
- Advanced SQL skills, including the development and optimization of queries, datasets, views, and structures supporting interactive visualization.
- Strong experience with Apache Spark/PySpark, SQL, Python, Delta Lake, ETL/ELT, data pipeline orchestration, and large-scale data transformation.
- Demonstrated experience implementing Bronze/Silver/Gold medallion architectures.
- Strong background designing normalized and analytical data models for complex enterprise environments.
- Proven ability to redesign and remediate poorly structured existing data environments rather than simply adding new pipelines.
- Experience with canonical data models, conformed dimensions, enterprise keys, entity resolution, and reusable semantic structures.
- Strong knowledge of data-quality engineering, financial reconciliation, metadata, lineage, and enterprise data governance.
- Experience with Databricks Unity Catalog or an equivalent enterprise data governance and catalog platform.
- Experience building production-grade pipelines with automated testing, monitoring, logging, failure handling, and CI/CD.
- Ability to translate operational workflows into logical and physical data architectures and visualization-ready data products.
- Experience working with highly regulated healthcare, financial, government, or similarly controlled environments.
- Ability to work with PHI, PII, CUI, and other controlled data in accordance with applicable security and privacy requirements.
- Strong collaboration skills and the ability to work across engineering, analytics, visualization, cybersecurity, product, architecture, and business SME teams.
- Ability to meet applicable federal security, privacy, access-control, and data-handling requirements.
- Prior experience with Advana and/or a War Data Platform is preferred.
- Direct experience developing Databricks data products and dashboards within a DoD enterprise environment is desirable.
- Experience with DHA, MHS, or other DoD healthcare data is preferred.
- Familiarity with MHS GENESIS, Oracle Health/Cerner Millennium, and/or Abacus is advantageous.
- Healthcare revenue-cycle experience involving encounters, coding, charge capture, claims, denials, adjudication, remittances, payments, AR, and revenue recovery is highly desirable.
- Familiarity with healthcare EDI transactions, including 837, 835, 270/271, 276/277, and 278, is a plus.
- Experience with ODCS or comparable machine-readable data contracts is preferred.
- Experience with GitLab-based DevSecOps, infrastructure-as-code, automated testing, and secure production promotion is advantageous.
- Familiarity with Collibra or comparable enterprise data catalogs is a plus.
- Experience supporting financial auditability, reconciliation, lineage, and audit-remediation initiatives is desirable.
- Familiarity with DoD RMF, NIST, IL4/IL5 environments, and federal data-governance requirements is preferred.
Benefits:
- Target salary range of $123,000–$160,000.
- 100% remote working arrangement.
- Opportunity to contribute to a high-impact federal healthcare data transformation.
- Work on complex Databricks, data engineering, visualization, governance, and analytics initiatives at enterprise scale.
- Exposure to advanced healthcare revenue-cycle data, predictive analytics, data quality, and revenue-recovery use cases.
- Opportunity to collaborate with experienced engineers, data scientists, analytics professionals, healthcare SMEs, and technical leaders.
- Opportunities to mentor other engineers and establish reusable technical standards and patterns.
- Mission-focused environment supporting critical government and healthcare outcomes.
- Compensation may vary based on location, internal equity, business considerations, contract requirements, qualifications, experience, skills, and security clearance.