We are seeking a Site Reliability Engineer (SRE) to support and improve the reliability, availability, security, and performance of applications and infrastructure hosted within an Amazon Web Services (AWS) cloud environment. The SRE will work closely with application development, cloud infrastructure, cybersecurity, and operations teams to automate deployments, maintain containerized workloads, support Red Hat Enterprise Linux (RHEL) systems, and improve CI/CD pipelines and orchestration capabilities.
The successful candidate thrives in a fast-paced, results-driven environment, has a strong systems administration and cloud engineering foundation with an automation-first mindset and an interest in improving system reliability through monitoring, infrastructure as code, automated testing, and repeatable deployment processes.
This position offers the ability to work a hybrid remote schedule with the requirement to perform on-site support at the Pentagon, Arlington, VA up to 80%.
Job Responsibilities:
As a Site Reliability Engineer, you'll have the opportunity to work with the AWS GovCloud and Classified Cloud environments. Your responsibilities will include:
- Maintain and improve the reliability, availability, performance, and security of AWS-hosted applications and infrastructure.
- Support containerized applications using technologies such as Docker, Kubernetes, Amazon ECS, and Amazon EKS.
- Administer and troubleshoot Red Hat Enterprise Linux (RHEL) servers and associated services.
- Develop, maintain, and troubleshoot CI/CD pipelines used to build, test, scan, and deploy applications and infrastructure.
- Support container orchestration platforms and automated application deployments.
- Automate infrastructure provisioning, configuration, deployment, and operational tasks.
- Implement and maintain monitoring, logging, alerting, and observability solutions.
- Troubleshoot application, container, operating system, networking, and AWS infrastructure issues.
- Participate in incident response, root-cause analysis, and corrective-action activities.
- Identify repetitive operational tasks and develop automation to reduce manual effort and operational risk.
- Work with development and security teams to integrate security controls and automated security testing into CI/CD pipelines.
- Develop and maintain operational documentation, runbooks, and standard operating procedures.
- Support system patching, vulnerability remediation, configuration management, and security-hardening activities.
Basic Qualifications:
We seek individuals who bring the following qualifications and skills to our team:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field
- 5+ years of experience in systems administration, cloud engineering, DevOps, SRE, or a related technical role.
- Hands-on experience working with AWS services and cloud infrastructure.
- Experience administering Linux, preferably Red Hat Enterprise Linux (RHEL).
- Experience with Docker container technologies.
- Experience supporting or working with a container orchestration platform, such as Kubernetes, Amazon EKS, or Amazon ECS.
- Experience building and managing container images and container registries.
- Experience developing, maintaining, or troubleshooting CI/CD pipelines.
- Working knowledge of Git and source-code management practices.
- Experience with scripting or automation using Python, Bash, and/or similar technologies.
- Understanding of networking fundamentals, including TCP/IP, DNS, HTTP/HTTPS, load balancing, routing, and firewalls/security groups.
- Experience troubleshooting complex technical issues across applications, operating systems, networks, and infrastructure.
- Ability to work collaboratively with development, infrastructure, cybersecurity, and operations teams.
- Must have an active Interim Secret or Secret Clearance
- Must meet DoD 8140 foundational compliance requirements through qualifying degree or certification
Preferred qualifications
- Strong experience with AWS services such as EC2, VPC, IAM, S3, CloudWatch, CloudTrail, Route 53, Elastic Load Balancing, ECR, ECS, and EKS.
- Knowledge of PKI Certificate based authentication / Single Sign-On
- Strong experience with Kubernetes administration and troubleshooting.
- Experience with GitLab CI/CD, Jenkins, AWS CodePipeline, or similar CI/CD platforms.
- Experience with Infrastructure as Code using Terraform, AWS CloudFormation, or AWS CDK.
- Experience implementing observability solutions using technologies such as CloudWatch.
- Experience performing root-cause analysis and developing corrective actions following production incidents.
- Knowledge of AWS security best practices, IAM policies and roles, encryption, secrets management, vulnerability management, and least-privilege principles.
- Experience with DevSecOps practices, including automated vulnerability scanning, container scanning, static code analysis, and security gates within CI/CD pipelines.
- Experience supporting systems subject to DISA STIGs, NIST controls, RMF, or other federal cybersecurity requirements.
- AWS certification such as AWS Certified Solutions Architect, SysOps Administrator, or DevOps Engineer.
- Red Hat, Kubernetes, or related industry certifications.
Pay: $70,000.00 - $120,000.00 per year
Benefits:
• 401(k) matching
• Dental insurance
• Health insurance
• Life insurance
• Paid time off
• Referral program
• Vision insurance
Work Location: Hybrid remote in Pentagon, DC 20301