Back to all jobs
S

VP, Site Reliability Engineer

SGX

Singapore · sg 2h ago

Job description

BUILD MARKETS. SHAPE ECONOMIES. At SGX Group, we create markets. We turn ideas into products and expand access to new opportunities. We are building the future of the exchange in-house: from architecture and platforms to the critical systems that power markets. The biggest decisions are still open to people who join us now. The opportunity As Vice President, Site Reliability Engineering, you will lead reliability across our platforms and services and build SRE into a discipline the wider engineering organisation works to. You will own the standards for observability, automation, incident response, and resilience, along with the accountability for whether they hold when it matters. Most engineering roles here manage the tension between agility and reliability. This one arbitrates it. Service levels, error budgets, and release decisions run through your remit, which means saying no sometimes and, harder, saying yes when the instinct is to hold. In market infrastructure, an outage is not an internal inconvenience. Everyone sees it. There are strong foundations to build on, and a significant amount still to do. The engineering model, tooling, and ways of working will keep changing. If you want to inherit a mature SRE practice, this is not that one yet. You would be building it. It will suit someone who enjoys complexity, is comfortable leading through change, and wants to build capability that lasts. The Site Reliability Engineering team Site Reliability Engineering keeps the platforms behind SGX Group's business and market infrastructure available, observable, and recoverable. The team sets the standards other engineering teams work to across service levels, error budgets, observability, incident response and automation. It works across engineering, infrastructure, security, and product, and it owns the practice as well as seen as the SME for ensuring resilience and proactive maintenance and continuous improvement of the estate. The next phase is about establishing SRE as a discipline rather than a function, moving reliability decisions upstream into design, and reducing the manual work that currently sits behind keeping services up. What you will do Own the outcome Define and lead the enterprise SRE strategy, establishing reliability engineering as a core discipline across critical products, platforms, and services. Establish and govern SLO, SLI, and error budget practices, so reliability trade-offs are made on evidence rather than instinct. Take accountability for availability, recovery readiness, and how services behave under load and under failure. Set the engineering standard Set the direction for observability, automation, self-healing, and resilience platforms, improving service stability, incident detection, and operational efficiency at scale. Lead incident management, recovery readiness, post-incident learning, and chaos engineering, so resilience is tested rather than assumed. Drive toil reduction and platform standardisation programmes to improve engineering productivity, reduce operational risk and embed reliability by design. Build the function Lead, coach and develop SRE managers and engineers, building strong technical depth and leadership capability. Influence engineering teams outside your reporting line, since most reliability decisions are made long before anything reaches you. What we are looking for Essentials Deep expertise in SRE principles: automation-first operations, observability-driven engineering, error budget management, and resilience engineering. Experience shaping SRE operating models, defining service-level objectives, governing automation standards, and leading platform and architecture reviews. A track record of strengthening resilience through incident response, performance optimisation, capacity planning, and continuous improvement. Experience building SRE capability, including coaching emerging leaders and influencing senior stakeholders across functions. Experience designing or scaling observability, automation, self-healing or resilience platforms in a complex environment. Credibility with engineering teams you do not manage. A large part of this role is influence rather than authority. Client confirmation required: Confirm the observability, incident-management, service-mesh and chaos-engineering stack. Do not list competing platforms as though all are in use. You should be comfortable with observability and cloud-native tooling such as OpenTelemetry, Prometheus, Grafana, Istio, Linkerd, and PagerDuty, and with chaos engineering tools such as Chaos Mesh or Litmus. What may set you apart You have built an SRE practice from a thin base rather than inheriting a mature one. You have led incident response somewhere downtime was visible outside the company. You combine technical depth with curiosity about the business, its products, and what drives its revenue. A degree in computer science, engineering, information systems, or a related field is useful. Equivalent practical experience is valued just as highly. Why this role matters You will work on technology that underpins critical market infrastructure, where reliability is not a quality attribute of the product. It is the product. When these platforms work, participants trade and capital moves. When they do not, everyone knows within seconds. The mandate is real. You will decide what reliability means here, how it is measured, and what the organisation is willing to trade for it. The work is demanding and the practice is still being built, which is exactly where the opportunity sits. There are not many chances in a career to establish a discipline rather than inherit one. About SGX Group SGX Group is one of the world's most trusted international marketplaces, known for its stability and openness. Anchored in Singapore, we enable price discovery, capital formation and risk management across asset classes, supported by resilient infrastructure and robust clearing. We convene issuers, investors and intermediaries to create and grow markets that stand the test of time. Find out more at www.SGXGroup.com .

Similar open jobs