Avatar for MetAntz
MetAntz
Actively Hiring
Connecting top talents with global opportunities

Site Reliability Engineer

  • $130k – $150k
  • |
  • |8 years of exp
  • |Full Time
Posted: 3 weeks ago
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationNot Allowed
Skills
Python
Java
Distributed Systems
Backend Development
Go (Golang)
Go

About the job

Title: Site Reliability Engineer (Application SRE)
Duration: Full-time Perm role.
Location: Palo Alto, California

About the Role
Committed to delivering best-in-class service reliability and performance. As part of this commitment, we are expanding our Site Reliability Engineering (SRE) team to ensure the reliability, performance, and availability of our software applications. We are looking for a highly motivated and technically talented Senior Application SRE to support our 24x7 FX trading environment. This role will focus on application monitoring, automation, and optimization to enhance system stability, minimize downtime, and improve overall user experience. The ideal candidate will bring strong problem-solving skills, experience in large-scale distributed systems, and a deep understanding of software and infrastructure reliability principles.

Responsibilities
Ensure the reliability, performance, and availability of applications through proactive monitoring and automation.
Develop and maintain real-time monitoring, alerting, and logging systems to detect and resolve issues before they impact customers.
Automate manual operations, including application deployment, configuration, scaling, and recovery.
Collaborate with software engineering teams to integrate reliability best practices into the development lifecycle.
Conduct root cause analysis (RCA) and implement preventive measures to mitigate recurring issues.
Support a 24x7 distributed enterprise environment across multiple global data centers.
Work closely with Support to enhance incident response processes, ensuring fast and effective resolution of technical escalations.
Participate in on-call rotations to support critical application issues and outages.
Maintain and optimize CI,CD pipelines to ensure fast and reliable application releases.
Enhance system security by managing SSL certificates, encryption, and authentication mechanisms.
Foster a culture of continuous improvement by evaluating new tools, frameworks, and methodologies to enhance system reliability.

Requirements

Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
3+ years of experience in a similar role, focusing on application reliability, automation, and performance optimization.
Strong expertise in Linux and Windows system administration.
Proficiency in at least one scripting language (e.g., Python, Shell, Perl, JavaScript).
Experience with Docker, Kubernetes, or containerization technologies.
Familiarity with CI,CD tools like Jenkins and deployment automation frameworks.
Hands-on experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack, New Relic, Datadog).
Understanding of networking concepts (TCP,IP, DNS, load balancing, firewalls).
Experience with configuration management tools like Ansible, Salt, or Puppet.
Strong debugging and troubleshooting skills across application, database, and infrastructure layers.
Ability to work in a fast-paced, high-pressure environment with multiple priorities.
Excellent communication and collaboration skills to work effectively with engineering and support teams.

Nice-to-Have Skills
Experience in the financial services or trading industry.
Knowledge of distributed computing, cloud platforms (AWS, GCP, Azure).
Exposure to security best practices and compliance standards.
Familiarity with incident management frameworks (ITIL, SRE best practices, or similar methodologies).

Why Join Us?
Be a key player in shaping SRE strategy and improving mission-critical trading systems.
Work in a collaborative, fast-paced environment with top engineering talent.
Enjoy career growth opportunities in an organization that values technical excellence and innovation.
Competitive compensation and benefits package.
If you are passionate about site reliability, automation, and scaling highly available applications, we would love to hear from you! Apply now and help us build the future of reliable trading technology.

Roles and responsibilities
Ensure the reliability, performance, and availability of applications through proactive monitoring and automation. Develop and maintain real-time monitoring, alerting, and logging systems to detect and resolve issues before they impact customers. Automate manual operations, including application deployment, configuration, scaling, and recovery. Collaborate with software engineering teams to integrate reliability best practices into the development lifecycle. Conduct root cause analysis (RCA) and implement preventive measures to mitigate recurring issues. Support a 24x7 distributed enterprise environment across multiple global data centers. Work closely with Support to enhance incident response processes, ensuring fast and effective resolution of technical escalations. Participate in on-call rotations to support critical application issues and outages. Maintain and optimize CI,CD pipelines to ensure fast and reliable application releases. Enhance system security by managing SSL certificates, encryption, and authentication mechanisms. Foster a culture of continuous improvement by evaluating new tools, frameworks, and methodologies to enhance system reliability.

Experience and education
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

About the company

MetAntz company logo

MetAntz

Actively Hiring
Connecting top talents with global opportunities51-200 Employees
Learn more about MetAntz image

Similar Jobs

Sponsor a Pet company logo
Sponsor a Pet
We are a fundraising company for animal non-profits
Assured company logo
Assured
Automated claims are now a reality
Caro company logo
Caro
Collaborative workforce planning to help companies hire faster without overspending
Mudflap company logo
Mudflap
Leading fintech innovation in trucking for independent owner operators & fleets
Avoma company logo
Avoma
AI Meeting Assistant, Collaboration, and Intelligence platform
Eudia company logo
Eudia
Eudia is revolutionizing legal work with AI-powered Augmented Intelligence