Title: Senior Site Reliability Engineer Duration: 06 Months Location: Austin, TX - Hybrid 4 days weekly onsite
Our Opportunity: We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications — including AI/ML-driven approaches to observability and reliability.
What you’ll do: • Evangelize SRE mindset and solve problems through systematization. • Identify opportunities to build innovative tools and solve unique operations problems on large enterprise and mission-critical applications. • Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions that measurably reduce manual toil and improve operational throughput. • Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems - including anomaly detection and predictive alerting to improve platform reliability. • Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms. • Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability. • Triage alerts and diagnose/resolve critical issues; manage implementation of changes with clear communication and minimal risk. • Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility and rollout validation at scale. • Champion AIOps platform adoption and ML-assisted observability practices across the team. • Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting. • Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps concepts and AI-assisted pipeline optimization. • Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development. • Participate in on-call support.
What do you have: Required Skills: • 6-8 years of experience with enterprise-level administration and support. • 6-8 years of experience writing automation scripts, building application dashboards for proactive monitoring, and setting up alerts for early issue determination. • 6-8 years practicing SDLC, process improvements. • Hands-on enterprise systems administration, monitoring, and deployment activities. • Experience with Windows 2019/2022 and Linux hosted via Virtual Machine. • Experience in Cloud application configuration, deployment, support, and migration - GCP/PCF is a plus. • Knowledge of IP networking including DNS, DHCP, firewalls, IP routing, etc. • Familiarity with large-scale distributed systems and high-availability architecture. • Linux and Windows system administration, troubleshooting, and tuning. • Development experience in one or more programming languages: .NET, PowerShell, Java, Python, Bash. • Knowledge of one or more of SQL, Oracle, MongoDB databases. • Working knowledge of Actimize. • Knowledge of one or more Message Brokers: Solace, RabbitMQ, IBM MQ, Kafka. • Knowledge of Splunk, AppDynamics, or similar observability tools. • Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments. • Bachelor's degree in computer science or related discipline.
Helpful Skills: • Financial services industry experience. • Agile methodologies. • Hands-on experience with AIOps platforms or ML-driven observability tooling. • Experience integrating AI/ML capabilities into CI/CD or operational automation workflows. • Familiarity with CI/CD tools (Harness, Jenkins, GitHub Actions) or GitOps concepts. • Exposure to container orchestration (Kubernetes, OpenShift) or cloud platforms (AWS, Azure, GCP).
Personal Skills: • Strong customer orientation with an affinity to proactively own, communicate, and follow through on projects and issues. • Extreme sense of ownership to resolve problems in a distributed environment. • Gritty resolve to dig deeper into technical issues in a complex login ecosystem. • A self-starter with the ability and confidence to independently resolve issues and bring results back to the team.
Applicant Notices & Disclaimers
For information on benefits, equal opportunity employment, and location-specific applicant notices, click here
At SPECTRAFORCE, we are committed to maintaining a workplace that ensures fair compensation and wage transparency in adherence with all applicable state and local laws.This position's pay range is $65.00/hr - $71.00/hr.