JobTarget Logo

Shelton - Infrastructure - Engr- Observability Sr in Shelton, Connecticut at Subway

NewSalary: $119200 - $149000
Subway
Shelton, Connecticut, 06484, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Shelton - Infrastructure - Engr- Observability Sr

Sr. Engineer, Observability

Why Join Us?

We are Subway Headquarters a dedicated team of professionals supporting thousands of franchisees around the globe. Ready for a fresh, new career? Look no further, because one of the world's most iconic brands can help you get there.

At Subway, better is baked into our DNA. We are a brand that believes in continued improvement in our lives, our businesses, and our planet. From the handshake that started our very first sandwich shop to earning our position as one of the world's leading restaurant brands, we've always embraced change and the path ahead. And today, we're making better living way easier.

Our purpose is about more than the food we serve in our restaurants. It's centered on fueling healthy businesses and healthier lives. It is one of the most exciting times to join the Subway team and contribute to our transformational journey.

About the Role

We have an exciting opportunity to support our Technology team as a Senior Engineer, Observability, based in Shelton, CT.

This is a senior individual-contributor role. The Senior Engineer, Observability is responsible for implementing and managing observability across Subway's infrastructure and applications servers, networks, cloud services, databases, restaurant point-of-sale and kiosk estates, mobile apps, and web platforms. The role owns configuration, maintenance, and documentation of our observability platforms, and partners with technology stakeholders through instrumentation planning, service reviews, and operational enablement so that monitoring keeps pace with the business.

Subway's observability estate is large and actively changing: business-event telemetry now runs in production across middleware, order, payment, and restaurant-device flows; a network and security logging domain has recently been onboarded; and incident routing is moving onto PagerDuty. Alongside building coverage, this role carries real responsibility for signal quality and for the cost of the telemetry we collect and query.

If you feel that this is the role for you, and you are successful with your application, be ready to be Bold, Empowered, Accountable, and ready to have Fun in a fast-paced and agile working environment.

Responsibilities include but are not limited to:
  • Design, configure, and maintain observability and alerting platforms Dynatrace (primary), SolarWinds, PagerDuty, Fluent Bit / Calyptia log pipelines, AWS CloudWatch, Azure Monitor, and Splunk. This covers instrumentation, dashboards and notebooks, alerting and anomaly detection, business-event processing and creation, log pipeline configuration and analysis, and synthetic and real-user monitoring.
  • Practice detection engineering, not just alert configuration select the right detection pattern for each failure mode, including silence and dead-man's-switch coverage for feeds that can stop, expected-versus-actual and seasonal-baseline comparison, and success-rate and ratio controls for partial degradation. Backtest new alerts against historical incidents before enabling them and prefer one dimensioned alert over many per-geography copies.
  • Improve signal quality and reduce noise review and tune alert configurations, thresholds, dwell times, and routing in collaboration with stakeholders; identify and retire duplicate, unowned, and non-actionable alerts; and make sure every production alert has an owner and a documented response.
  • Own telemetry cost and consumption governance understand how each observability capability is billed, attribute consumption to owning teams through cost and product tagging, monitor ingest and query volumes against commitments, and identify optimization opportunities in retention, sampling, synthetic frequency, and query efficiency.
  • Build and maintain automation author Dynatrace workflows and automation apps and manage observability configuration as code in source control with a promotion path from lower environments to production.
  • Collaborate with developers, architects, and platform teams to define observability strategies aligned to business and operational goals. This includes integrating observability into CI/CD and infrastructure-as-code workflows and supporting business observability through end-to-end transaction tracing and event instrumentation.
  • Provide technical leadership and enablement guide the work of internal engineers and external partners engaged on observability, set technical direction and quality expectations, and coach teams through knowledge transfer, documentation, and training. This role leads through technical influence rather than direct people management.
  • Support incident and problem management perform proactive monitoring and incident response for production systems; participate in major-incident reviews and problem records; correlate anomalies against planned change windows before treating them as faults; and translate post-incident findings into concrete detection improvements.
  • Manage platform vendor relationships day to day raise and drive vendor support cases to resolution, evaluate release notes and advisories for impact, and coordinate agent and platform upgrades with change management.
  • Develop, document, and maintain observability processes and response procedures, ensuring operational readiness and participating in service reviews so that coverage continues to meet stakeholder needs.
  • Bachelor's degree in Computer Science or a related field.
  • 6+ years of related technical experience.
Qualifications

Observability expertise

  • Expert in Dynatrace, with demonstrable skills across distributed tracing, log management, business events, metrics, DQL, OpenPipeline, dashboards and notebooks, alerting, anomaly detection, and automation workflows.
  • Working familiarity with additional observability and alerting tooling such as SolarWinds, PagerDuty, Fluent Bit / Calyptia, Azure Monitor, AWS CloudWatch, and Splunk.
  • Practical understanding of observability for distributed retail and digital estates point-of-sale and in-restaurant devices, third-party delivery and ordering integrations, payment flows, and mobile and web front ends.

Detection and alert design

  • Track record of designing alerts that catch real failures with a low false-positive rate, and of retiring alerts that do not earn their keep.
  • Comfortable reasoning about missing data, low-volume and seasonal signals, and the difference between an outage and a degradation.

Cost and data governance

  • Experience managing observability consumption against a commercial commitment ingest, retention, and query cost and making evidence-based decisions about what telemetry is worth collecting.
  • Familiarity with tagging and attribution models that let cost be allocated to owning teams.

Cloud proficiency

  • Hands-on experience with AWS and/or Azure services to support instrumentation and management of observability for cloud-hosted applications and services, including serverless and container workloads.

Collaboration and communication

  • Strong interpersonal and communication skills, written and verbal, with the ability to work effectively across technology teams, development groups, architecture stakeholders, and non-technical business partners.
  • Able to explain what telemetry does and does not show, and to hold a clear position in a post-incident review.

Technical skills

  • Proficient in Windows and Linux environments; scripting in PowerShell and Bash, and comfort with at least one general-purpose language for automation.
  • Skilled in administering observability platforms and their integrations, including ServiceNow for incident creation, PagerDuty for routing and on-call, and cloud-native monitoring services.
  • Working knowledge of SNMP, WMI, ICMP, NetFlow, syslog, and REST APIs for data collection and integration.
  • Version control and configuration-as-code practice (Git-based workflows); familiarity with Azure DevOps for backlog and delivery tracking.

Mindset and approach

  • Naturally curious and committed to continuous learning and improvement.
  • Proactive in identifying opportunities for optimization and innovation in monitoring strategy.
  • Rigorous about evidence willing to verify a number before acting on it, and to correct the record when it turns out to be wrong.
  • Comfortable working with ambiguity and competing priorities across concurrent platform changes.

What do we Offer?
  • Insurance Plans (Medical/Life)
  • 401K
  • Competitive Bonus
  • Mobility Allowance
  • Tuition Reimbursement
  • Company Holidays
  • Volunteering time
  • And Many More

Actual pay is determined based on a number of job-related factors including skills, education, training, credentials, qualifications, scope and complexity of role responsibilities, geographic location, performance, and working conditions.

The Company is only considering applicants who are currently authorized to work in the country the position is based. AA/EOE/D/V

Job Location

Shelton, Connecticut, 06484, United States

Frequently asked questions about this position

Similar Jobs In Shelton, Connecticut

Hot Job

Cleanroom Operator - Manufacturing Technician

Central Semiconductor LLC
Hauppauge, New York
Hot Job

Doubles Driver

DICARLO DISTRIBUTORS, INC.
HOLTSVILLE, New York

2nd Shift Production Operator

sbhpp
Manchester, Connecticut

Crane Operator

CentiMark Corporation
Rocky Hill, Connecticut
Urgently Hiring

Equipment Operator

Peckham Industries
Brewster, New York
Apply For This Position

Apply Now