Staff Data Engineer in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Data Engineer based in United States.
This is a foundational opportunity to help build an end-to-end data platform from the ground up.
You will design the near real-time and AI-ready infrastructure that brings together data from CRM, financial, HR, and operational systems.
Your work will establish the pipelines, streaming architecture, vector infrastructure, and reliability practices that become the backbone of the organization’s data ecosystem.
You will work closely with senior data leadership, Analytics Engineering, BI, Product, Engineering, Finance, Marketing, Customer Success, and Operations.
The role combines deep hands-on engineering with the opportunity to shape architecture, standards, governance, and technical direction.
You will also help enable AI-powered reporting, semantic search, retrieval-augmented generation, and agentic use cases through production-grade data infrastructure.
This is an ideal role for a senior builder who thrives in lean environments and wants meaningful ownership over a modern data platform.
- Design, build, and own near real-time data pipelines using CDC, streaming ingestion, and event-driven architectures to power the platform’s core data flows.
- Evaluate, implement, and maintain vector database infrastructure and embedding pipelines supporting semantic search, retrieval-augmented generation, AI agents, and other AI-enabled applications.
- Build scalable ELT and ETL pipelines that ingest data from internal platforms, financial systems, CRM, HRIS, and other sources for both real-time and batch use cases.
- Partner with senior data leadership to architect the warehouse or lakehouse as the supporting system of record beneath streaming and AI infrastructure.
- Develop lightweight transformation layers using technologies such as dbt to help Analytics Engineering teams turn raw data into reliable, business-ready datasets and metrics.
- Own data pipeline reliability and observability, including monitoring, automated failure alerting, lineage tracking, and operational troubleshooting across streaming and batch environments.
- Establish the technical foundation for self-service and AI-powered reporting while partnering with BI, Product, and Engineering teams on executive and departmental reporting needs.
- Implement and maintain data governance practices covering documentation, access controls, lineage, and data quality standards.
- Collaborate with Finance, Marketing, Customer Success, Operations, and other stakeholders to translate business requirements into reliable, low-latency data products.
- Leverage AI-augmented development tools such as Claude Code or comparable solutions to accelerate engineering, testing, documentation, and workflow automation.
- Contribute to the evolution of the broader data architecture and establish scalable engineering practices as the platform grows.
- 7+ years of hands-on data engineering experience, including substantial depth in streaming and event-driven architectures rather than exclusively batch processing.
- Proven experience designing and building production-grade near real-time pipelines from the ground up using technologies such as Kafka, Kinesis, Flink, Debezium, CDC, or similar tools.
- Hands-on production experience with vector databases and embedding infrastructure, such as Pinecone, Weaviate, pgvector, Milvus, Zilliz, or comparable technologies.
- Experience developing embedding strategies and chunking approaches for retrieval and AI-powered use cases is strongly valued.
- Advanced proficiency in SQL and Python.
- Working knowledge of cloud warehouse or lakehouse platforms such as Snowflake, BigQuery, or Databricks, as well as dbt.
- Proven experience building or materially contributing to an end-to-end production data environment, ideally as an early, founding, or highly autonomous data engineering hire.
- Familiarity with B2B SaaS data models, including customer lifecycle, sales pipeline, conversion, recurring revenue, ARR, CAC, and churn concepts.
- Working knowledge of BI and reporting tools such as Looker, Tableau, Power BI, or similar platforms.
- Familiarity with ELT/ETL tools such as Fivetran or Airbyte for batch data integration.
- Strong understanding of data observability and reliability practices, including automated alerting, lineage tracking, monitoring, and failure recovery.
- Experience using AI-augmented engineering and development tools such as Claude, Copilot, or similar technologies.
- Ability to work independently, manage complex multi-stakeholder initiatives, and make sound technical decisions in a lean, fast-moving environment.
- Comfortable operating with ambiguity and building new infrastructure, processes, and standards where limited existing foundations are in place.
- Competitive compensation package based on skills, experience, qualifications, and work location.
- Health benefits for eligible U.S.-based employees.
- Flexible paid time off.
- Parental leave.
- Fertility and adoption assistance.
- 401(k) benefits.
- Educational reimbursement.
- Opportunities to work with modern streaming, AI, vector database, and data platform technologies.
- Meaningful ownership over the development of a new end-to-end data platform.
- Support for reasonable accommodations throughout the application and interview process.
- A collaborative environment focused on building scalable technology and continuously improving data capabilities.