Avatar for Vimo
Vimo
Actively Hiring
SaaS platform for health & human service programs

Site Reliability Engineer

  • $120k – $170k
  • |
  • |2 years of exp
  • |Full Time
Posted: 2 weeks ago• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Linux
Bash
DNS
Jenkins
Unix
TCP/IP
Load Balancing
Docker
PagerDuty
ec2
S3
AWS
Elk
Kubernetes
Terraform
GitLab CI
IAM
Prometheus
Grafana
Go
RDS
CloudWatch
Tls/ssl
Vpc
GitHub Actions
VictoriaMetrics
Loki
HTTP/HTTPS

About the job

Vimo® started as the “Expedia” of health insurance and has evolved into a leader in transforming government IT infrastructure with its proven SaaS and AI technology. Our innovative approach to health insurance shopping and enrollment has expanded beyond exchanges, and we are now reinventing how states administer safety net programs such as Medicaid, SNAP (food stamps), child care, and unemployment insurance. With our cutting-edge technology, we are helping agencies serve more people, faster, and transforming healthcare service delivery as we know it.

We are looking for a Site Reliability Engineer (SRE) to join our Vimo team.

About The Role

As a Site Reliability Engineer, you will help ensure the reliability, availability, and performance of Vimo’s production platform. Our systems power health insurance exchanges, Medicaid enrollment, and other safety-net programs for state governments—the services you support directly impact millions of people’s access to critical benefits. You will work alongside senior SREs, developers, and infrastructure engineers to build automation, improve observability, respond to incidents, and reduce operational toil. This role is ideal for an engineer who is passionate about systems thinking, enjoys solving problems at scale, and wants to grow their career in reliability engineering within a mission-driven environment.

Responsibilities

  • Monitor, maintain, and troubleshoot production services to ensure high availability and performance across Vimo’s SaaS platform.
  • Respond to production incidents as part of the on-call rotation; triage alerts, coordinate with engineering teams, and drive issues to resolution.
  • Contribute to blameless postmortem processes by documenting incidents, identifying root causes, and tracking follow-up action items.
  • Build and maintain CI/CD pipelines to support safe, repeatable, and efficient application deployments.
  • Write automation scripts and tools (Python, Bash, or Go) to reduce manual operational work and improve reliability.
  • Support and improve observability infrastructure including monitoring dashboards, log aggregation, distributed tracing, and alerting using tools such as Datadog, Prometheus, Grafana, ELK/OpenSearch, and PagerDuty.
  • Manage cloud infrastructure on AWS (EC2, EKS, RDS, S3, VPC, CloudFront, Route 53, Lambda) following established standards and best practices.
  • Work with infrastructure-as-code tools (Terraform) and container orchestration (Docker, Kubernetes/EKS) to provision and manage environments.
  • Assist with capacity planning, load testing, and performance analysis to prepare for peak traffic periods such as open enrollment seasons.
  • Collaborate with application development teams to improve service reliability through architecture reviews, production readiness checklists, and resilience patterns.
  • Support disaster recovery procedures including backup validation, failover testing, and documentation of recovery runbooks.
  • Help maintain compliance with security and regulatory requirements (HIPAA, FedRAMP, SOC 2) by ensuring infrastructure controls are properly implemented and documented.
  • Continuously improve on-call processes, runbooks, and operational documentation to reduce mean time to detection (MTTD) and mean time to resolution (MTTR).

Qualifications

Basic Qualifications/Skills

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or a related operations-focused role.
  • Proficiency in at least one programming or scripting language (Python, Bash, Go, or similar) for writing automation and tooling.
  • Hands-on experience with AWS cloud services (EC2, RDS, S3, VPC, IAM, CloudWatch, or equivalent).
  • Familiarity with container technologies (Docker) and orchestration platforms (Kubernetes).
  • Experience with infrastructure-as-code tools such as Terraform, or similar.
  • Understanding of CI/CD concepts and experience with at least one pipeline tool (Jenkins, GitLab CI, GitHub Actions, or similar).
  • Familiarity with monitoring and observability tools (Loki, Prometheus, VictoriaMetrics, Grafana, CloudWatch, ELK, or PagerDuty).
  • Solid understanding of Linux/Unix systems administration, including process management, file systems, and networking basics.
  • Understanding of networking fundamentals: TCP/IP, DNS, HTTP/HTTPS, load balancing, and TLS/SSL.
  • Strong troubleshooting and analytical skills with the ability to diagnose issues across the application and infrastructure stack.
  • Good communication skills and the ability to work collaboratively in a team-oriented environment.

Preferred Qualifications/Skills

  • Experience in healthcare technology, government IT, or benefits administration platforms.
  • Familiarity with compliance frameworks such as HIPAA, FedRAMP, or SOC 2 and their operational implications.
  • Experience with PostgreSQL, Mongo or such relational & NoSQL database systems in a production environment.
  • Exposure to incident management frameworks and blameless postmortem practices.
  • Experience with GitOps workflows and tools (ArgoCD or similar).
  • Familiarity with configuration management tools (Puppet, Ansible, Chef or similar).
  • Experience with log management and analysis at scale.
  • Exposure to load testing or performance benchmarking tools (k6, Locust, JMeter).
  • AWS certifications (Cloud Practitioner, Solutions Architect Associate, or SysOps Administrator) are a plus.
  • Familiarity with SLO/SLI concepts and error budget-driven development practices.

Compensation and Benefits

Competitive compensation - All In range of ($120,000-$165,000). (Please note that compensation may vary based on factors such as skills, experience, performance and location.)

We offer a comprehensive benefits package, including but not limited to:

  • Health, Dental, Life, Disability, and Vision insurance
  • Healthcare spending or reimbursement accounts (HSA/FSA)
  • Retirement benefits (401k)
  • Paid time off
  • Holidays: 13 paid days per year
  • Education assistance or tuition reimbursement
  • Employee discounts for Gym memberships & commuting/travel assistance

Our Values

  • We believe that working hard, when it is imbued with purpose, can and should be fun.
  • You'll find we are a "can do" place where people work together and roll up their sleeves to get the job done.
  • Everyone has a voice; everyone's ideas count, and everyone is respected.
  • We have built a company, as well as a community of friends and colleagues, with respect for each other.

About the company

Vimo company logo

Vimo

Actively Hiring
SaaS platform for health & human service programs501-1000 Employees
Learn more about Vimo image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Scale AI company logo
Scale AI
Accelerate the development of AI applications
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
C3 AI company logo
C3 AI
C3 AI is a leading enterprise AI software provider for accelerating digital transformation
Mercor company logo
Mercor
Mercor is at the intersection of labor markets and AI research