JobTarget Logo

Senior DevOps Engineer, AI Platform in Canada Creek, Nova Scotia at Jobgether

NewJob Function: Information Technology
Jobgether
Canada Creek, Nova Scotia, B0P 1V0, Canada
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Senior DevOps Engineer, AI Platform

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer, AI Platform based in Canada.

As a Senior DevOps Engineer, you’ll build and operate the cloud infrastructure powering AI platforms, web applications, APIs, and backend services.
You’ll translate technical designs into secure, scalable, reliable, and production-ready environments across modern cloud platforms.
The role combines deep Kubernetes expertise with cloud networking, infrastructure automation, CI/CD, and observability.
You’ll support AI workloads including LLM gateways, agent runtimes, RAG pipelines, and asynchronous processing services.
You’ll work closely with AI engineers, application developers, and architects while independently owning infrastructure delivery and operations.
Production reliability, incident response, scalability, security, and cost optimization will be central to your impact.
This is an opportunity to help create reusable platform capabilities that enable engineering teams to deliver sophisticated services faster and more consistently.

Accountabilities
  • Translate application and platform technical designs into reliable, secure, scalable, and production-ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily using Azure Kubernetes Service and Oracle Kubernetes Engine.
  • Build and support infrastructure for AI workloads, including LLM gateways, Python-based agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage cloud networking, including virtual networks, subnets, routing, NAT, load balancers, DNS, TLS, private connectivity, firewalls, network policies, ingress, egress, and service-to-service communication.
  • Operate infrastructure supporting web applications, APIs, databases, caches, queues, scheduled jobs, microservices, and event-driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and Infrastructure as Code practices.
  • Implement comprehensive observability across infrastructure and applications through metrics, logs, distributed tracing, dashboards, alerts, health checks, and service-level objectives.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns, templates, and operational practices that enable engineering teams to launch services efficiently and consistently.
  • Determine required cloud resources, Kubernetes configurations, namespaces, scaling models, identities, secrets, and supporting services for new workloads.
  • Provision and operate dependencies such as PostgreSQL, Redis, RabbitMQ, storage systems, and other shared platform services.
  • Establish CI/CD workflows covering builds, testing, container publishing, deployment, validation, and rollback.
  • Define operational runbooks, capacity monitoring, dashboards, alerts, and health checks before production launches.
  • Own infrastructure delivery through UAT and production while partnering with architects and engineers to resolve technical design trade-offs.
Requirements
  • 7+ years of professional experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or a closely related discipline.
  • Strong hands-on experience operating production Kubernetes environments, including expertise in networking, scheduling, storage, autoscaling, security, and troubleshooting.
  • Strong Microsoft Azure experience, particularly with AKS, networking, identity, storage, and monitoring; Oracle Cloud Infrastructure experience is preferred.
  • Deep understanding of cloud networking concepts, including virtual networks, subnets, routing, NAT, load balancing, private networking, DNS, TLS, firewalls, ingress, and egress.
  • Proven experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
  • Experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
  • Hands-on knowledge of databases, caching, and messaging technologies such as PostgreSQL, Redis, RabbitMQ, or equivalent platforms.
  • Experience implementing production observability using tools such as OpenTelemetry, Grafana, Prometheus, Sentry, or cloud-native monitoring solutions.
  • Strong Linux, systems administration, and production troubleshooting capabilities.
  • Working knowledge of Python, particularly backend services built with frameworks such as FastAPI.
  • Familiarity with at least one additional programming language such as C#, Java, Go, JavaScript, or TypeScript.
  • Strong understanding of HTTP/HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead-letter queues, and asynchronous processing.
  • Ability to read application logs and stack traces and diagnose infrastructure and application issues involving latency, memory, CPU, connections, and dependencies.
  • Strong cross-functional communication skills and the ability to independently execute technical designs while engaging architects and application engineers when needed.
  • Experience supporting AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services is highly desirable.
  • Familiarity with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry would be an advantage.
  • Experience building reusable infrastructure platforms for high-scale SaaS or customer-facing applications is a plus.
  • Strong experience operating distributed systems using technologies such as RabbitMQ, Redis, and PostgreSQL is beneficial.
Benefits
  • Full-time, fully remote position available across Canada.
  • Opportunity to work on infrastructure supporting modern AI platforms, agent systems, RAG workloads, and cloud-native applications.
  • High-impact role with significant ownership over production infrastructure, reliability, scalability, and platform engineering.
  • Opportunity to work across Microsoft Azure and Oracle Cloud Infrastructure.
  • Exposure to modern technologies including Kubernetes, Terraform, Helm, ArgoCD, Docker, OpenTelemetry, and GitOps.
  • Collaborative environment working closely with AI engineers, application engineers, architects, and platform teams.
  • Opportunity to create reusable infrastructure capabilities that accelerate engineering delivery.
  • Professional growth through hands-on work with large-scale distributed systems and emerging AI infrastructure.
  • Flexible remote-first working environment designed to support collaboration across distributed teams.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

Canada Creek, Nova Scotia, B0P 1V0, Canada

Frequently asked questions about this position

Similar Jobs In Canada Creek, Nova Scotia

New

Software Engineer, AI Agents

Jobgether
Canada Creek, Nova Scotia
New

Backend Engineer, Core APIs

Jobgether
Canada Creek, Nova Scotia
New

Software Engineer - Data Insights

Jobgether
Canada Creek, Nova Scotia
New

App Developer - FlutterFlow AI

Jobgether
Canada Creek, Nova Scotia
New

Staff Developer - Cloud Engineering

Jobgether
Canada Creek, Nova Scotia
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.