Senior Site Reliability Engineer
United Kingdom, United States
£145,000 - £185,000 / per year
About this role
Phone numbers and emails in this ad are masked until you log in.
auto_translated_note
About the Role
SunCore Digital is seeking a hands-on Senior Site Reliability Engineer to assess the current reliability and scalability of our systems, identify risks, and implement the technical changes required to address them.This is not a monitoring-only or advisory position. The SRE will investigate existing applications and infrastructure, establish reliability baselines, develop observability and testing capabilities, and directly implement reliability mitigations within the SRE domain. When a mitigation requires application-specific code or business-workflow changes, the system-owning team will implement and maintain those changes with SRE guidance.Because the platform has not yet been validated under real customer traffic, capacity and scalability will be treated as open risks until testing provides evidence otherwise.ResponsibilitiesReliability assessment and remediationAssess the reliability of web applications, Flutter mobile services, APIs, backend systems, infrastructure, databases, queues, and third-party integrations.Identify single points of failure, fragile dependencies, manual operational processes, and failure modes.Distinguish confirmed issues from suspected risks and areas that have not yet been evaluated.Design and implement shared reliability capabilities and improvements within the SRE domain.Define application-specific reliability changes and work with system-owning teams to implement them; those teams retain responsibility for their code and services.Transfer service-specific instrumentation, runbooks, and ongoing maintenance responsibilities to the appropriate system owners after the solution is tested and hardened.Track identified risks through implementation and validation rather than stopping at recommendations.Performance, load, and capacityEstablish a practical performance and capacity-testing program.Work with QA and product stakeholders to identify critical workflows and realistic usage scenarios.Establish baseline response times, throughput, concurrency, and resource consumption.Design and execute load, stress, endurance, scalability, and failure tests.Identify bottlenecks involving applications, databases, networks, queues, caches, infrastructure, and external services.Implement shared SRE mitigations and coordinate application-specific mitigation work with system-owning teams, which retain responsibility for their code and services.Repeat testing after changes to verify results.Document tested capacity, observed constraints, and remaining unknowns.Define safe operating limits and early-warning indicators.Performance and load testing have not yet been completed across the platform.
Establishing this capability will be an early priority.ObservabilityAssess current logging, metrics, tracing, health checks, dashboards, and alerting.Establish consistent observability standards across services.Implement and harden shared health-checking and observability capabilities, transfer shared components to the designated long-term owner, and work with system-owning teams on service-specific instrumentation that they will maintain after handoff.Define meaningful service-level indicators and initial reliability objectives with system owners and engineering leadership.Build shared reliability dashboards and initial service-specific views, then train system owners to maintain their service-specific dashboards and alerts.Ensure alerts are actionable, routed to accountable owners, and tested.Identify monitoring blind spots.Improve application instrumentation in collaboration with developers.Ensure logs and telemetry do not expose sensitive information.Incident readinessHelp establish the initial on-call and escalation model.Create and test incident-response and troubleshooting runbooks.Define severity levels and technical escalation paths.Lead or support incident investigation.Improve detection, diagnosis, mitigation, and restoration capabilities.Facilitate technically focused post-incident reviews.Track corrective actions and recurring failure patterns.Conduct controlled failure exercises where appropriate.SRE automationAutomate repetitive reliability-engineering work, including health checks, diagnostics, alert enrichment, incident triage, capacity checks, and evidence collection.Develop safe automated remediation or self-healing for clearly defined and well-tested failure conditions.Reduce manual diagnostic, maintenance, incident-response, and reliability-validation steps within the SRE function.Define and implement the reliability checks, test logic, and SRE automation that should run through CI/CD. Work with DevOps to integrate them into the shared delivery framework, and with system-owning teams to maintain application-specific configuration after handoff.Create reusable reliability tools and patterns, harden and document them, transfer shared components to the designated long-term owner, and train application teams to operate the service-specific portions they own.Document automation ownership, safeguards, limitations, rollback behavior, and conditions requiring human intervention.CollaborationWork with the Disaster Recovery and Resilience Engineer on service and data recovery dependencies.Work with Security Operations on the security review of SRE implementations before production adoption.Provide reliability
Requirements
for deployment safeguards and environment health, and work with DevOps and platform staff on shared integration. SRE does not own deployment automation or routine application deployments.Work with QA to define realistic user scenarios and post-mitigation validation.Work with application teams on system-specific code changes and instrumentation.Provide factual technical findings to engineering leadership without presenting unverified assumptions as confirmed conclusions.Initial PrioritiesInventory critical services and their dependencies.Assess reliability risks and potential single points of failure.Establish system-health and latency visibility.Define critical workflows for performance testing.Build the initial performance, load, and capacity-testing capability.Establish technical baselines.Identify and implement the highest-priority reliability and scaling mitigations.Create initial runbooks and on-call recommendations.Identify SRE processes that should be automated, including diagnostics, alert handling, capacity validation, incident response, and safe remediation.Document tested behavior, unresolved risks, and unknowns.Required
Qualifications
Five or more years of experience in site reliability, production engineering, infrastructure engineering, DevOps, or a comparable role.Hands-on experience supporting cloud-hosted applications and distributed systems.Experience implementing reliability improvements rather than only producing assessments.Experience with performance, load, stress, or capacity testing.Experience diagnosing application, database, network, and infrastructure bottlenecks.Strong knowledge of monitoring, logging, metrics, alerting, and incident response.Experience with reliability automation and scripting.Experience working with CI/CD systems.Experience troubleshooting APIs, backend services, databases, queues, caches, and external integrations.Understanding of timeout, retry, rate-limit, circuit-breaking, and graceful-degradation patterns.Ability to work directly in unfamiliar codebases and environments.Strong technical documentation and communication skills.Ability to distinguish confirmed findings, suspected risks, and unknowns.Comfort working in a small organization where operational practices are still being established.Bonus PointsExperience preparing a platform for its first external users.Experience establishing an SRE capability in a startup or small engineering organization.Experience with financial, digital-asset, telemetry, equipment-monitoring, or operational platforms.Experience working with Flutter-backed mobile services.Experience with chaos or fault-injection testing.Experience with data-intensive or event-driven systems.Experience in a remote, asynchronous environment.
What Success Looks Like
Critical services and dependencies are inventoried.Reliability risks are visible and prioritized.Performance and capacity baselines exist for critical workflows.High-priority reliability and scaling mitigations are implemented and tested.Critical services have meaningful health checks, dashboards, alerts, and runbooks.The company can identify and respond to failures without relying entirely on one person.Application teams understand and maintain the reliability improvements made to their systems.Remaining capacity and reliability risks are documented clearly.Compensation• Negotiable based on location• Deferred completed Demo• Role transitions into a permanent hire thereafterWhy Join SunCore Digital• High-impact role shaping the security foundation of a rapidly scaling digital company• Fully remote team with async flexibility and optional collaboration in London or Bellevue• Direct partnership with executive leadership and top-tier engineering, marketing, and design partners• Competitive compensation with performance upsideOriginally posted on Himalayas
Community Q&A
Anyone worked here? Ask before you apply.
No threads yet for this job or company.