
- B2B
- Scale StageRapidly increasing operations
- Valuation $1B+This company has a valuation of $1B or more
Site Reliability Engineer
- |5 years of exp
- |Full Time
In office - WFH flexibility
Not Available
About the job
TITLE: Site Reliability Engineer
DEPARTMENT: Product Engineering / Operational Readiness
REPORTING TO: Senior Manager, System Engineering and Integrated Product Support
OFFICE LOCATION: New York, NY
ROLE TYPE: Hybrid/Full-time
IPC is a global fintech company that puts people at the center of innovation. With a strong global footprint, we empower financial institutions and capital markets with advanced cloud-based trading communications and managed connectivity solutions.
Through our portfolio of communications and connectivity solutions, we focus on solving business challenges and adapting to regulatory changes in the fast-paced global financial markets. This enables our clients to maintain consistent market access, a strong competitive advantage, and enhanced operational efficiency.
Join a team that is dedicated to delivering groundbreaking products and making a significant impact on our clients' success.
Overview of the Team
The Network Services and OSS Engineering team is responsible for the reliability, monitoring, automation, and operational support of IPC's global technology platforms. The team supports mission-critical voice, network, infrastructure, and application services that power communication and connectivity solutions for the world's leading financial institutions.
Working across multiple regions and time zones, the team partners closely with Infrastructure Engineering, Product Engineering, Information Security, Network Operations, Service Delivery, and third-party technology providers to ensure highly available, scalable, and secure services.
Role Overview:
The Site Reliability Engineer II (SRE) is responsible for the architecture, development, implementation, administration, and operational support of IPC's enterprise monitoring, observability, event management, and automation platforms. The role ensures the reliability, scalability, performance, and operational visibility of IPC's global infrastructure, applications, cloud environments, and customer-facing services.
The position combines traditional Site Reliability Engineering responsibilities with software development, automation engineering, observability architecture, and ITSM integration expertise. The engineer serves as a technical leader responsible for monitoring transformation initiatives, event correlation, ServiceNow integrations, dashboard development, operational automation, and continuous platform modernization. The role partners closely with Product Engineering, Cloud Engineering, Infrastructure Engineering, Service Operations, Security, and Architecture teams to ensure operational readiness across all IPC products and platforms.
The ideal candidate possesses strong Linux administration expertise, experience with enterprise monitoring platforms, automation technologies, and a passion for improving reliability through engineering best practices.
Monitoring & Observability Engineering
- Own and improve the monitoring and observability capabilities for assigned services and platforms, helping teams detect issues earlier, understand system health, and operate reliable customer-facing services.
- Maintain and enhance monitoring across IBM Netcool, Splunk, Datadog, Grafana, Nagios, HPE NNMi, OpenTelemetry, and event-management platforms.
- Apply established standards and patterns for metrics, logs, traces, synthetic checks, alerting, dashboards, and operational reporting across cloud, on-premises, and hybrid environments.
- Partner with engineering teams to define practical monitoring and production-readiness requirements for new and changing services.
- Use known reliability factors, service history, and precedent to resolve diverse observability issues, escalating or seeking review at key decision points.
Software Development & Automation
- Build dependable tools and automation that reduce operational effort and make production services easier to support.
- Develop and maintain Python automation, REST API integrations, reusable monitoring components, and event-processing workflows.
- Create automated remediation and self-service capabilities for well-understood operational scenarios, with appropriate safeguards and review.
- Integrate monitoring and operational checks into CI/CD pipelines and use Infrastructure as Code to support repeatable deployments.
ServiceNow & ITSM Integration
- Own assigned integrations between observability platforms and ServiceNow, connecting alerts, incidents, configuration data, and operational workflows.
- Implement and support Event Management, IT Operations Management, CMDB, service-mapping, ticket-generation, and incident-lifecycle integrations.
- Advise colleagues on difficult integration and workflow issues and recommend solutions grounded in platform capabilities and operational precedent.
Platform Reliability & Service Ownership
- Take end-to-end responsibility for the reliability and day-to-day health of assigned monitoring services, systems, or workstreams, working independently with review at key points.
- Administer and improve Netcool ObjectServer, Event Gateway, probes, dashboards, correlation rules, suppression, deduplication, enrichment, and related platform components.
- Improve availability, performance, capacity, security, and maintainability through tuning, upgrades, lifecycle management, patching, and technical debt reduction.
- Define and track meaningful service indicators, objectives, alerts, and error budgets for assigned systems.
Incident Response & Continuous Improvement
- Play an active role in production support and incident response, restoring service quickly while turning operational learning into lasting reliability improvements.
- Troubleshoot and resolve complex infrastructure, application, and monitoring incidents; coordinate recovery for assigned systems and contribute effectively during major incidents.
- Lead or contribute to root-cause analysis, identify recurring failure patterns, and deliver corrective actions that reduce repeat incidents.
- Maintain clear runbooks, support procedures, service documentation, and operational readiness evidence.
Cloud & Container Observability
- Support AWS, Kubernetes, container, distributed-tracing, and OpenTelemetry observability for defined environments and services.
- Help teams interpret service health, choose effective telemetry, and adopt reliable operating practices across hybrid platforms.
Collaboration & Technical Influence
- Work closely with engineering, operations, security, service management, and senior partners to deliver pragmatic reliability outcomes.
- Communicate recommendations clearly, using evidence and operational context to persuade stakeholders when priorities or approaches differ.
- Provide practical guidance to colleagues on difficult technical matters and contribute to standards, roadmaps, evaluations, and proof-of-value initiatives.
- Balance delivery, operational risk, security requirements, and long-term maintainability when planning improvements.
Experience & Professional Level
This Site Reliability Engineer II opportunity is suited to a fully qualified professional with approximately five years of relevant experience in reliability engineering, production operations, monitoring, observability, platform engineering, or automation. You will have room to own meaningful systems and improvements, work independently, influence experienced partners, and continue developing your technical depth with support and review at important points.
How You Will Make an Impact:
As an OSS Site Reliability Engineer II, you will take meaningful ownership of mission-critical systems and help make them more reliable, observable, and ready to scale. You will solve real operational challenges, improve the customer experience, and grow your expertise alongside experienced technical partners.
- Own the day-to-day reliability and performance of mission-critical services, following issues through to durable improvements.
- Build and improve modern observability through actionable metrics, logs, traces, dashboards, and alerts.
- Automate repetitive operational work to reduce risk, increase consistency, and give the team more time for higher-value engineering.
- Improve incident detection and recovery by sharpening alerts, supporting effective response, and applying lessons learned.
- Strengthen operational readiness for new and changing services through practical reviews, testing, documentation, and launch support.
- Collaborate with senior engineers and technical partners to prioritize reliability work, protect the customer experience, and expand your SRE skills.
- Success means more dependable services, faster detection and recovery, smoother launches, and measurable growth in your ability to own and improve complex systems.
Essential Skills and Experience to be Successful in this Role:
- Bachelor’s degree in a relevant field, or equivalent practical experience.
- Approximately 5 years of experience in site reliability, platform engineering, or production infrastructure, with independent ownership of assigned systems or workstreams.
- Strong Linux skills and working knowledge of Windows environments.
- Hands-on experience with monitoring, alerting, logging, and observability for production services.
- Ability to automate operational work using Python or a similar language, APIs, and configuration-management tools.
- Proven troubleshooting across applications, infrastructure, and networks in highly available environments.
- Clear communication, familiarity with IT service management practices, and the ability to manage competing priorities.
Desired Skills and Experience:
- Experience applying SRE practices, including SLIs, SLOs, and error budgets.
- Experience with modern observability, APM, and distributed tracing tools (for example, Datadog, Dynatrace, New Relic, Splunk, Prometheus, Grafana, or OpenTelemetry).
- Experience with cloud, containers, or hybrid environments, plus Infrastructure as Code and CI/CD practices.
- Experience integrating ServiceNow or comparable ITSM platforms with monitoring, event management, or automated workflows.
- Experience in fintech, telecom, or another mission-critical environment, with responsible use of AI-assisted engineering tools.
What’s in It for You?
At IPC, your compensation is only part of the package. We are committed to investing in a range of programs and initiatives to improve the overall experience of our employees.
In addition to a collaborative, high-performing team environment, we’re pleased to offer benefits including:
- Competitive Base Salaries
- Medical, Dental and Vision
- 100% Employer Paid Short/Long Term Disability, AD&D and Life Insurance Coverage
- Limited & HealthCare Flexible Spending Accounts
- Dependent Care Flexible Spending Account
- Health Savings Accounts with Employer Contributions
- Pet Insurance
- Legal Insurance
- Critical Illness, Hospital Indemnity and Accident Coverage
- Commuter Benefits
- Medicare Services Education program
- Financial Wellness Account Fiduciary and Training
- Identity Theft Insurance
- 401(k) and Roth plan with matching contributions
- Flexible PTO, Sick Pay and Public Holidays
- Additional Time off for Charity Work and Volunteering
- Tuition Reimbursement
- Certification Bonus Program
- Access to “IPC University” our Internal E-Learning Platform
- Structured Onboarding Training and Peer Mentor Support
- Parental Leave Policy
- Free Mental Health Wellness Programs, Tools, Coaching and Free therapy sessions
- Employee Referral Scheme
Further information about your benefits will be provided during your onboarding process.
Additional Information:
At IPC, we believe that hybrid working creates an inclusive, flexible environment where employees can perform at their best, and teams can collaborate, innovate, and celebrate successes together. We spend around 60% of our time in the office and around 40% of our time working remotely. Some employees may be required to work from the office or client sites more than 60% of the time, if required by their role and/or client needs.
Your precise work schedule will be determined by you and your Line Manager before commencement of employment with IPC.
You can explore more about our culture, offerings and commitment on www.ipc.com/careers/ and www.ipc.com/about-us/about-ipc/.
IPC’s Work Culture:
The IPC work culture is one that fosters inclusion, prioritizes innovation, and maximizes potential. We are a global ecosystem, full of diverse people that together made IPC what it is today.
Our strength as an organization is the sum of our different backgrounds, perspectives, skills and geographies; supported by an ironclad commitment to constructive dialogue and open-mindedness.
We live and breathe our commitment to innovation by embracing bold ideas, seizing new opportunities and striving for excellence. Our people have continued to deliver ground-breaking solutions to our clients for over 50 years.
About the company

IPC Systems
- B2B
- Scale StageRapidly increasing operations
- Valuation $1B+This company has a valuation of $1B or more
Similar Jobs









