Senior DevOps engineer. in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer based in the United States.
This is a senior-level opportunity to help operate and evolve a globally distributed cloud infrastructure supporting millions of connected devices and users. You will play a key role in improving reliability, scalability, performance, automation, and cloud cost efficiency across complex production environments. The role spans cloud infrastructure, networking, databases, messaging, observability, CI/CD, and infrastructure as code. You will collaborate with engineering leaders and technical teams across countries and time zones while contributing to both immediate operational priorities and long-term infrastructure strategy. The position offers significant influence over architecture, automation, capacity planning, and cloud optimization. This is a fully remote role with occasional travel, supporting a global technology platform.
- Manage and continuously improve highly scalable and reliable systems supporting millions of devices and users across geographically distributed platform instances.
- Drive infrastructure reliability, scalability, performance, automation, and cloud cost-efficiency initiatives.
- Collaborate with engineering leaders and peers across multiple countries to establish repeatable, scalable infrastructure objectives and practices.
- Partner with DevOps leadership on short- and long-term infrastructure initiatives, including architecture, budgeting, implementation, capacity planning, and cloud cost optimization.
- Identify opportunities to automate operational processes and improve engineering workflows.
- Design, deploy, and operate cloud applications and infrastructure across AWS and GCP environments.
- Support networking, messaging, databases, observability, and high-availability services within production environments.
- Participate in a scheduled 24/7 on-call rotation to support production systems and maintain service reliability.
- Contribute to the long-term technical direction of infrastructure while balancing operational requirements, scalability, resilience, and efficiency.
Requirements:
- 8+ years of professional experience supporting Docker-based microservices, PaaS environments, and cloud technologies.
- 5+ years of networking experience supporting high-availability applications on AWS and GCP, with experience across both platforms preferred.
- Experience with HTTP- and MQTT-based ingestion and messaging solutions.
- 5+ years of hands-on Linux administration experience, including system, user, and machine administration, package management, and scripting with Python and Bash.
- 5+ years of experience deploying and operating cloud applications, particularly across AWS and GCP; Azure experience is a plus.
- 3+ years of experience managing high-availability Kafka clusters, with strong knowledge of brokers, producers, consumers, partitions, and Kafka internals.
- 3+ years of database administration experience with technologies such as MySQL, Cassandra, or other NoSQL databases.
- 3+ years of experience with log collection and analysis, performance monitoring, and tuning using platforms such as Coralogix, Prometheus, OpenTelemetry, New Relic, Datadog, or similar tools.
- 3+ years of experience with infrastructure automation, particularly Terraform; Ansible experience is valuable.
- Strong knowledge of modern CI/CD tools, automation practices, and deployment workflows.
- Hands-on experience with AI-assisted development and automation tools such as Claude, Cursor, Gemini, or similar technologies.
- Strong understanding of DevOps architecture and operations, including infrastructure strategy, capacity planning, budgeting, cost optimization, implementation, reliability, and scalability.
- Excellent communication and collaboration skills, with the ability to work effectively across teams, functions, countries, and time zones.
- Ability to work effectively in a collaborative engineering environment and solve complex infrastructure and distributed-systems challenges.
Benefits:
- 100% remote position.
- Candidates should be located in the Eastern Time Zone of the United States or Canada, or in Belfast, Northern Ireland, to support collaboration with the broader engineering team.
- Occasional travel may be required.
- Opportunity to work on large-scale distributed systems supporting millions of connected devices and users.
- Exposure to complex challenges across cloud infrastructure, networking, messaging, databases, observability, security, automation, and high availability.
- Significant opportunity to influence infrastructure architecture, technical direction, scalability, resilience, automation, and cloud efficiency.
- Collaboration with an experienced international engineering team.