Site Reliability Engineer responsible for monitoring system health, detecting issues before they escalate, and owning incident response and debugging end-to-end. You'll build logging and observability tooling, automate deployment pipelines, manage capacity planning, and drive platform reliability at scale.… The role requires 5+ years in SRE/DevOps, hands-on production incident response, and strong communication during live troubleshooting. You'll work with AWS, Terraform, Datadog, Python/Bash scripting, and modern infrastructure-as-code practices.