JobTarget Logo

Member of Technical Staff | Observability & Reliability in New York at Jobgether

NewJob Function: Information Technology
Jobgether
New York, 10455, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Member of Technical Staff | Observability & Reliability

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | Observability & Reliability based in Brazil.

This is a high-impact platform engineering role focused on making distributed systems observable, reliable, and operationally resilient.
You’ll own observability across cloud and customer-hosted environments, ensuring teams can understand system health wherever workloads run.
The role spans logs, metrics, traces, alerting, SLOs, incident response, and reliability engineering.
You’ll work closely with Kubernetes-based infrastructure and software operating across environments you may not fully control.
Your work will directly support high availability, faster incident resolution, and consistent deployment health.
You’ll also help reduce telemetry costs by improving the quality and efficiency of the signals collected.
This is an autonomous, hands-on position where you’ll build, operate, and continuously improve critical platform capabilities.

Accountabilities
  • Evolve and maintain the observability platform covering logs, metrics, traces, alerting, and system health across cloud and customer-hosted dataplanes.

  • Ensure every environment reports critical operational information, including active releases, health status, heartbeats, logs, metrics, and usage to the central control plane.

  • Implement telemetry collection within customer Kubernetes environments using outbound-only connectivity models.

  • Detect and investigate differences between desired infrastructure or deployment state and what is actually running in each environment.

  • Monitor the health and availability of deployment and runtime agents, including ephemeral workloads such as Ray clusters supporting batch inference.

  • Define and maintain Service Level Objectives (SLOs), establish actionable alerting, and contribute to error-budget practices.

  • Lead or participate in incident response and postmortems, identifying improvements that reduce recurring failures and mean time to recovery (MTTR).

  • Coordinate incident resolution across internal teams and customers when fixes involve customer-managed environments.

  • Optimize telemetry pipelines to reduce redundant data, control infrastructure costs, and improve the signal-to-noise ratio of operational information.

  • Write production-quality code, review technical changes, and take operational ownership of the systems you build.

Requirements
  • Deep professional experience with OpenTelemetry and modern observability platforms or backends.

  • Hands-on experience defining SLOs, working with error budgets, designing actionable alerts, and managing production incidents.

  • Strong experience with Kubernetes and infrastructure-as-code tools such as Terraform and Helm.

  • Experience operating software across distributed or customer-hosted environments where infrastructure and connectivity may not be fully under your control.

  • Strong software engineering fundamentals, with experience producing maintainable, production-ready code and conducting effective code reviews.

  • Willingness to participate in operational ownership, troubleshooting, incident response, and continuous reliability improvements.

  • Strong analytical and problem-solving abilities, with an ability to investigate complex distributed-system behavior.

  • Experience communicating clearly across engineering teams and, when required, working directly with external customers or stakeholders.

  • Experience with GCP/GKE or AWS/EKS is an advantage.

  • Familiarity with multi-node or multi-cluster ML workloads in production is a plus.

  • Experience deploying software to customer-hosted Kubernetes environments, including Helm-based deployments and outbound-only connectivity, is valuable.

  • Experience in financial services or other regulated environments is an additional advantage.

Benefits
  • Fully remote role based in Brazil.

  • Full-time position within an engineering-focused environment.

  • Opportunity to own critical observability and reliability systems with direct impact on production availability.

  • Work across cloud and customer-hosted Kubernetes environments, providing broad exposure to distributed infrastructure.

  • Opportunity to work with modern observability, Kubernetes, infrastructure-as-code, and ML infrastructure technologies.

  • High degree of technical ownership, with responsibility for both building and operating the systems you develop.

  • Exposure to complex reliability challenges involving real-time services, batch workloads, customer environments, and distributed systems.

  • Opportunity to contribute to incident management, platform architecture, and long-term reliability practices.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

New York, 10455, United States

Frequently asked questions about this position

Similar Jobs In Other / Non-US, New York

NewHot Job

Director of Nutrition Services

BMS Family Health and Wellness Centers
Brooklyn, New York
NewHot Job

1st Shift Counter Salesperson/Material Handler

Alro Steel Corporation
Rochester, New York

Maintenance Technician

DIMARCO STAFFING LLC
Hudson Falls, New York
New

Legal Operations and Compliance Associate

Jobgether
Other / Non-US, New York
New

Senior Compliance Coordinator

Rochester Regional Health
ROCHESTER, New York
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.