Sr. Site Reliability Engineer - Observability
ExpiredDenverLast seen 15 days ago
Summary
Senior Site Reliability Engineer focused on observability, responsible for building and improving shared infrastructure, monitoring, alerting, and disaster recovery across AppFolio's real estate platform. The role requires diagnosing performance and reliability issues across the full stack, developing Service Level Indicators/Objectives, and collaborating with engineering teams to enhance system reliability and quality. Must have strong coding skills (Go, Ruby, or Python preferred), deep expertise with Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Pulumi), and AWS services. On-call responsibilities included; 5+ years industry experience or equivalent required.
DevOps / InfrastructureAWSKubernetesPythonTerraformCapacity PlanningCloudformationConfiguration ManagementContainer OrchestrationDdos ProtectionDisaster RecoveryDynamodbEbsEksGoInfrastructure As CodeLambdaLinuxLoad BalancersMySQLObservabilityPulumiRds AuroraRoute53RubyRuby On RailsRunbooksS3Service Level IndicatorsService Level ObjectivesSREVPC