HPC Infrastructure & Cluster Engineer in Springfield, Virginia at INflow
Explore Related Opportunities
Job Description
We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environmen. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.
Here, your work is more than a job- it's a journey in innovation. With opportunities to work on high-impact projects, access to the latest technologies, and a culture that thrives on creativity and collaboration, INflow Federal is where your expertise can truly make a difference.
Specific Duties and Responsibilities:- Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
- Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
- Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
- Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
- Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
- Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
- Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
- Technical Skills:
- Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
- Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
- Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
- Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
- Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
- Familiarity with parallel file systems and high-throughput storage architectures.
- Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
Other Notes
- Some travel may be required: Must have valid driver’s license and transportation. This is subject to change at the direction of the customer.
- If accommodation is needed with your application or the interview process for applicants with disabilities, please contact Human Resources at 703-594-8601.
- Candidate must have the ability to lift up to 50 lbs.
- Must have willingness to perform duties not listed in the job description as required by INflow and our customer.
$140,000 - $185,000 a year