Site Reliability Engineer (Application Reliability)
Hyderabad, IN
ABOUT US:
As a world leading provider of integrated solutions for the alternative investment industry, Alter Domus (meaning “The Other House” in Latin) is proud to be home to 90% of the top 30 asset managers in the private markets, and more than 6,000 professionals across 24 jurisdictions.
With a deep understanding of what it takes to succeed in alternatives, we believe in being different - in what we do, in how we work and most importantly in how we enable and develop our people. Invest yourself in the alternative, and join an organization where you progress on merit, where you can speak openly with whoever you are speaking to, and where you will be supported along whichever path you choose to take.
Find out more about life at Alter Domus at careers.alterdomus.com
Site Reliability Engineer (Application Reliability)
Location: Hyderabad
Working Hours: GMT/BST, aligned with London
About the Role
We are growing our Site Reliability Engineering team with a Senior Officer, Site Reliability Engineer who will help make the applications and services behind our products measurably more reliable. Our SRE team brings together engineers who care about instrumentation, incident response and the health of production systems, all working toward a common mission: giving product and engineering teams the observability, guardrails and operational practices they need to run their services with confidence.
This role is for an SRE who comes from an application engineering background. You will spend your time close to the code and the telemetry it produces: instrumenting services with OpenTelemetry, building the dashboards, SLOs and alerts that tell us when something is wrong, and digging into application behaviour when it is. Expect to read and debug code in a stack such as .NET, Java, Node.js or Python, follow a request through HTTP, gRPC, queues and databases, and work out where latency, errors or resource exhaustion are actually coming from, then work with the owning team to fix it properly rather than restart it again.
Our platform runs around the clock and so does the team behind it. On-call coverage moves naturally across time zones as the day progresses. This position anchors the London (GMT/BST) window of that follow-the-sun rotation, and as part of it you should expect to step in on weekends from time to time.
This is a strong fit for an experienced individual contributor who has built and operated production software, enjoys the discipline of measuring reliability rather than guessing at it, and wants to deepen their expertise in site reliability engineering within a regulated financial services environment.
On-Call & Coverage Expectations
On-call is a core, ongoing part of this role and not an occasional add-on. Please read this section carefully before applying.
- Your standard working hours align with London (GMT/BST) and cover the European business day from our Hyderabad office.
- You will take part in a shared on-call rotation with the wider SRE team, following a follow-the-sun model in which coverage hands over between regions as each working day ends.
- While on rotation you are expected to be reachable, to acknowledge and respond to pages within the agreed response targets, and to drive or support incidents through to resolution and handover.
- Weekend and public holiday coverage is part of the rotation and comes around periodically; you should expect to step in from time to time.
- Escalation and paging run through PagerDuty. You will help keep escalation policies, schedules and alert routing accurate and low-noise.
- After significant incidents you will take part in blameless postmortems and follow through on the resulting action items.
What You'll Do
Build: Observability & Reliability Engineering
- Instrument applications and services with OpenTelemetry (traces, metrics and logs), including SDK-based and auto-instrumentation, correct context propagation and sensible attribute and cardinality choices.
- Build and maintain Grafana dashboards and Prometheus-based metrics, recording rules and alerts that reflect real user-facing behaviour rather than raw infrastructure noise.
- Define and implement SLIs, SLOs and error budgets with product engineering teams, and make burn-rate alerting the basis of paging where it fits.
- Improve alert quality: remove duplicate and low-value alerts, tighten thresholds, add runbook links and make sure every page is actionable.
- Build and maintain PagerDuty integrations, services, escalation policies and schedules so that alerts reach the right people with the right context.
- Write code and automation (e.g., Python, Go, Bash, or your primary application language) to reduce toil, automate recurring operational tasks and codify observability configuration.
- Use AI-assisted tools (e.g., coding assistants, observability and incident copilots) in your own day-to-day work and help other developers adopt AI tooling effectively.
Operate: Incident Response & Production Ownership
- Monitor service health, respond to incidents and take part in our on-call rotation, including occasional weekend coverage.
- Act as incident responder and, where needed, incident commander: triage, coordinate responders, communicate status and drive to mitigation.
- Debug production issues in the application layer, including latency and error-rate regressions, memory and garbage-collection pressure, thread and connection pool exhaustion, retry storms, timeouts and cascading failures.
- Investigate data-layer problems alongside application code: slow queries and query plans, missing indexes, lock contention and deadlocks, connection pooling and replication lag.
- Run blameless postmortems, identify contributing causes and track corrective actions to completion.
- Contribute to release and change safety: readiness and liveness behaviour, graceful shutdown, canary and progressive rollout signals, rollback criteria.
- Support disaster recovery and resilience testing, including failover and restore exercises for the services you support.
Engineering: Reliability in the Application Stack
- Review application changes from a reliability standpoint: timeouts, retries with backoff and jitter, idempotency, circuit breakers, bulkheads and graceful degradation.
- Work with product teams on performance and capacity: profiling, load and soak testing, and understanding how services behave under saturation.
- Help teams design for observability from the start, so new services ship with useful telemetry, dashboards and alerts on day one.
- Contribute reliability improvements directly to application repositories rather than working around them from the outside.
Collaboration & Continuous Improvement
- Partner with product engineering teams to understand their services, their failure modes and their reliability goals, and translate those into SLOs and operational practice.
- Partner with the Cloud Platform and DevEx teams so that observability and reliability tooling is available as a self-service golden path rather than a bespoke effort per team.
- Help developers across the organization onboard to our observability tooling, troubleshoot issues, get support when needed and review their instrumentation and reliability-related PRs.
- Contribute to internal documentation, runbooks and best practices for reliability and incident response.
What You Bring
Required Skills
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- 4+ years of professional experience in software engineering, site reliability engineering, production support engineering, or a closely related technical domain, including hands-on production ownership.
- Strong background in at least one application stack (e.g., .NET, Java, Node.js, Python or Go), with the ability to read, debug, profile and change production code.
- Hands-on experience with OpenTelemetry for application instrumentation (traces, metrics, logs) and trace context propagation.
- Practical experience with Grafana and Prometheus (PromQL, dashboards, recording and alerting rules).
- Experience with PagerDuty or an equivalent incident and on-call platform, including escalation policies and schedules.
- Solid grasp of core SRE concepts: SLIs, SLOs, error budgets, toil reduction, blameless postmortems, change and release safety.
- Working knowledge of core protocols and their failure modes: HTTP/1.1 and HTTP/2, REST, gRPC, TLS, DNS and TCP, plus messaging or streaming (e.g., Kafka, SQS, Service Bus, RabbitMQ).
- Practical database skills: SQL fluency, reading query plans, indexing basics, connection pooling, and caching with something like Redis.
- Comfortable at the command line and with Linux fundamentals (processes, memory, file descriptors, networking tools).
- Experience working with CI/CD pipelines and Git-based workflows.
- Comfort operating in a regulated environment, with an understanding of security and compliance fundamentals (access control, least privilege, audit logging).
- Clear written and verbal communication in English, and the ability to stay calm and structured during incidents.
- Willingness and ability to work London (GMT/BST) hours from Hyderabad and to take part in the on-call rotation described above, including occasional weekend coverage.
Nice to Have
- Experience with distributed tracing and log backends such as Grafana Tempo, Loki, Jaeger, Elastic or Splunk.
- Working knowledge of Kubernetes from an application perspective (probes, resource requests and limits, HPA behaviour, debugging a workload in-cluster).
- Experience with load and performance testing tools such as k6, JMeter or Gatling, or with chaos and resilience testing.
- Exposure to infrastructure-as-code (e.g., Terraform) and observability-as-code practices.
- Experience with AWS and/or Azure application and monitoring services.
- Familiarity with compliance frameworks relevant to financial services (e.g., SOC 2, ISO 27001).
- Relevant certifications (e.g., Google/Coursera SRE, AWS or Azure developer or DevOps certifications, CKAD).
WHAT WE OFFER
We are committed to supporting your development, advancing your career, and providing benefits that matter to you.
Our industry-leading Alter Domus Academy offers six learning zones for every stage of your career, with resources tailored to your ambitions and resources from LinkedIn Learning.
Our global benefits also include:
- Support for professional accreditations such as ACCA and study leave
- Flexible arrangements, generous holidays, plus an additional day off for your birthday!
- Continuous mentoring along your career progression
- Active sports, events and social committees across our offices
- 24/7 support available from our Employee Assistance Program
- The opportunity to invest in our growth and success through our Employee Share Plan
- Plus additional local benefits depending on your location
Alter Domus is an Equal Opportunity Employer: Equity Statement
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, or protected veteran status.
(Alter Domus Privacy notice can be reviewed via Alter Domus webpage: https://alterdomus.com/privacy-notice/)
#LI-HYBRID