Posted August 01, 2026
HPC Linux Administrator
ITR
Oak Ridge, TN, US
Full Time
Job Description
Job Description
HPC Linux Administrator
Basic Qualifications:
- A BS degree in computer science, computer engineering, information technology, business, science, or a related field of study and eight (8) to twelve (12) years of proven experience is required. An overall combination of equivalent experience may be considered.
- Five (5) or more years with managing UNIX/Linux Systems.
- Three (3) or more years of proven experience with configuration management and automation tools such as Git, Jenkins, Ansible, or Puppet.
- Moderate proficiency in at least one scripting language such as Bash, Python, or others.
- Experience performing advanced troubleshooting and system administration with Linux Servers.
- Experience supporting large data systems.
- A strong desire to innovate and identify new technologies and opportunities and be able to communicate the potential benefits of those choices to others within the team and our research partners.
- A collaborative and upbeat approach to thrive on the opportunity to build trust and credibility, and ultimately become a trusted advisor to our research teams.
- The ability to obtain and maintain a Department of Energy "Q" clearance is required. This requires US Citizenship.
Preferred Qualifications:
- Solid understanding of multiple operating systems and cluster technologies.
- Experience with Centos/RHEL, Ubuntu, and VMware
- Understanding of HPC platforms to support users with SLURM job submissions and troubleshooting.
- Experience building and running containerized applications in an HPC environment.
- Experience with multiple deployment mechanisms like Diskless, Warewulf, and traditional deployment (Cobbler, PXEboot, and/or Bright).
- Experience managing systems utilizing GPU (NVIDIA and AMD) clusters for AI/ML and/or image processing.
- Knowledge of networking fundamentals including TCP/IP, traffic analysis, common protocols, and network diagnostics.
- Experience with InfiniBand networks and diagnostics.
- Extensive experience with High Performance Parallel File Systems (Lustre, WEKA, GPFS, etc).
- Experience with performance and diagnostic tools for benchmarking, analysis, and tuning of systems, networking, and storage.
- Experience with Grafana, CheckMK, Nagios, Zabbix, SolarWinds, Ganglia, or other network and device monitoring systems.
- Previous experience working in a government, scientific, or other highly technical environment.
- Good documentation skills, including the ability to prepare simple documentation web pages.
