Senior Software Engineer (Serverless) in Ireland, Scotland at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer (Serverless) based in Ireland.
This role offers the opportunity to build the next generation of AI cloud infrastructure powering advanced machine learning workloads worldwide. You will work on a high-impact serverless platform designed to help developers deploy and scale AI applications without managing complex infrastructure. As a senior engineer, you will take ownership of critical distributed systems challenges, from GPU scheduling and runtime performance to customer-facing APIs and platform reliability. Working in a highly technical environment, you will influence architecture decisions, mentor engineers, and help define engineering standards. This position is ideal for an experienced software engineer passionate about large-scale systems, cloud technologies, and solving complex infrastructure challenges at the forefront of AI innovation.
- Design, develop, and maintain core components of a GPU-native serverless AI platform, including control planes, schedulers, runtimes, autoscaling systems, and customer-facing APIs.
- Solve complex engineering challenges related to cold-start optimization, GPU scheduling, multi-tenant isolation, fair resource allocation, request routing, and platform scalability.
- Own technical architecture decisions for key platform areas by creating design documents, evaluating solutions, and aligning engineering teams around effective approaches.
- Establish and maintain high engineering standards through code reviews, design reviews, technical guidance, and active collaboration with team members.
- Operate services with an SRE mindset by defining reliability objectives, improving observability, supporting incident response, and driving continuous platform improvements.
- Collaborate directly with customers on architecture discussions, performance optimization, and complex production issues requiring deep technical expertise.
- Partner with product, infrastructure, and go-to-market teams to translate customer needs into scalable technical roadmaps.
- Contribute to improving platform performance, reliability, and developer experience through innovative engineering solutions.
- Support knowledge sharing and technical mentorship to help raise the overall engineering capability of the team.
- 7+ years of professional software engineering experience with a proven track record of building and operating large-scale distributed systems.
- Strong programming experience with Golang, or the ability and willingness to quickly become proficient in the language.
- Deep experience with Kubernetes and container orchestration systems, including real-world operation of production environments.
- Strong understanding of distributed systems concepts such as consistency, availability trade-offs, queueing, backpressure, retries, idempotency, and multi-tenancy.
- Experience designing and operating high-throughput, low-latency services with a strong focus on performance optimization and reliability.
- Proven ability to take ownership of complex technical challenges, lead design discussions, unblock teams, and deliver impactful solutions.
- Strong software engineering fundamentals with the ability to write reliable, maintainable code and investigate complex technical problems.
- Collaborative mindset with excellent communication skills and the ability to work effectively across engineering and business teams.
- Experience with serverless platforms or function-as-a-service technologies such as Knative, AWS Lambda, GCP Cloud Run, Cloudflare Workers, or similar solutions is a strong advantage.
- Familiarity with GPU scheduling technologies, including Kubernetes device plugins, MIG, MPS, time-slicing, or NVIDIA GPU Operator, is highly desirable.
- Experience with ML inference technologies such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, or similar platforms is a plus.
- Knowledge of runtime optimization techniques including cold-start reduction, image streaming, checkpoint/restore, or sandboxing technologies is beneficial.
- Experience developing Kubernetes operators using Go and frameworks such as controller-runtime or kubebuilder is advantageous.
- Contributions to open-source projects related to serverless infrastructure, scheduling, or AI inference are a plus.
- Competitive compensation package.
- Flexible working environment with hybrid opportunities.
- High level of ownership and autonomy over technical decisions.
- Career development opportunities and continuous learning support.
- Opportunity to work on impactful AI infrastructure projects shaping the future of machine learning.
- Collaborative and innovative culture with highly skilled international teams.
- Exposure to cutting-edge technologies across cloud infrastructure, distributed systems, GPUs, and AI workloads.
- Opportunity to contribute to ambitious projects in a fast-moving, high-growth environment.