- Growth StageExpanding market presence
Principal DevOps Engineer
- Remote ()
- |5 years of exp
- |Full Time
Remote only
Not Available
About the job
Don’t miss your chance to join a global team that's transforming the way people travel as a Principal DevOps Engineer.
About iVisa
At iVisa we believe that traveling should be simple. That’s why 2.6M+ travelers have chosen us to facilitate their visas, passports, and other travel documents. We are the easiest, fastest, and simplest solution in the market. Our company is growing 80% year on year. We know our biggest strength is our people and we’re looking for the right new team members to help propel our culture and achieve our goals. Above all else, we always have fun!
About The Role
Own the reliability, security, cost and evolution of iVisa's production and staging infrastructure end to end. This is the primary infrastructure role: you are the technical owner of our AWS accounts, Kubernetes clusters, GitOps and CI pipelines, Cloudflare edge, observability stack and internal DevOps tooling; the escalation point for every engineering team; and the manager and mentor of the Site Reliability Engineer who reports to you.
Why iVisa
- Collaborative, friendly, and diverse culture: We foster an inclusive and vibrant atmosphere, featuring a dynamic and international environment with flat hierarchies and exceptionally amiable colleagues.
- Truly remote-first work environment: work from anywhere or everywhere - we encourage global travel.
- Access to continuous learning and cutting-edge AI tools, supported by dedicated budgets for professional development.
- Extended Family Leave policy: Our policy covers all birthing parents, non-birthing parents, and adopting parents.
- Thrive in a highly tech-savvy company equipped with cutting-edge tools and the power to make a substantial impact.
- Join us in our commitment to the planet and sustainability: For every iViser, we plant one tree, allowing you to contribute to our environmental initiatives.
- Rest and Relaxation: We offer flexible PTO for all team members.
As a Principal DevOps Engineer at iVisa, you will make an impact by:
Platform and cloud
- Run, patch and upgrade EKS clusters, node groups and add-ons with minimal downtime; plan and communicate maintenance windows.
- Own the Terraform estate as the single source of truth: eliminate drift, review and apply infrastructure merge requests.
- Operate managed data services: RDS MySQL upgrades and parameter tuning, ElastiCache scaling, Redshift users/groups/WLM, S3 lifecycle and access logging, snapshot replication to the backup account. Lead their roadmap items (for example TLS enforcement and auth-plugin migration ahead of MySQL 9).
- Manage cloud cost: find idle or over-provisioned resources, right-size, retire obsolete distributions, domains and services.
Delivery and GitOps
- Own GitLab CI/CD and the self-hosted runners; keep pipelines fast and reliable across all product repositories; maintain the shared CI images.
- Own Argo CD, the Helm chart library and the review-environment lifecycle; unblock developers whose deployments fail.
- Provision infrastructure for new services (buckets, CDNs, IAM, DNS, routing, environment configuration) with the web, AI/ML, data, mobile and growth teams.
Reliability and incident response
- Primary on-call for infrastructure in a rotation shared with the SRE: respond to alerts, lead incidents, restore service, write root-cause analyses and turn them into alerts, runbooks or fixes.
- Own the observability stack: metrics, log retention and performance, alert rules, dashboards as code, uptime checks, Sentry organization.
- Own disaster-recovery posture: backups, restore testing, multi-replica and multi-node scheduling, capacity for traffic spikes.
Security and access
- Own the edge: Cloudflare WAF rules, bot, credential-stuffing and enumeration mitigation, IP blocking, Under Attack decisions, Worker deployments.
- Manage identity and access end to end: onboarding and offboarding for VPN, Vault, Keeper, GitLab and database users; least-privilege IAM; SSO; image vulnerability scanning and policy enforcement.
- Handle DNS, registrar and certificate operations (Route53, MarkMonitor, BIMI/SPF/DKIM records, wildcard certificates).
Platform engineering and enablement
- Maintain and extend the in-house Go DevOps tooling and the Runway service catalog.
- Keep runbooks and the wiki current (DDoS, VPN, environment configuration, Kubernetes operations); answer requests in #devops-comms; review infrastructure-touching merge requests from other teams.
- Manage, mentor and grow the Site Reliability Engineer; set priorities for the DevOps board.
- Use AI tooling (Claude, MCP servers, agent-friendly repository docs) to automate toil and speed up investigations.
What makes you a great fit for this role:
- 5+ years in DevOps, SRE or platform engineering, including 2+ years as the primary or lead owner of a production Kubernetes platform.
- Expert Kubernetes on AWS EKS: upgrades, node lifecycle, networking/CNI, ingress (Traefik or similar), RBAC, policy engines, debugging node and dataplane failures.
- Strong Terraform: modules, multi-environment layouts, state management, drift remediation.
- Broad AWS: EKS, EC2/Auto Scaling, IAM (including IRSA), VPC, RDS MySQL, ElastiCache, S3, CloudFront, ECR, Route53, ALB.
- GitOps with Argo CD and Helm chart authoring; GitLab CI with self-hosted Kubernetes runners.
- Cloudflare in production: DNS, WAF and custom rules, caching, Workers.
- Observability with Prometheus, Grafana and Loki; alert design; incident command and blameless RCAs.
- Working proficiency in Go (our internal tooling is Go) and fluent Linux/Bash; comfortable reading PHP, Python and TypeScript.
- Secrets and identity: HashiCorp Vault or equivalent, rotation, least-privilege design.
- Networking fundamentals: DNS, TLS, VPN/WireGuard, load balancing, HTTP.
- Clear written communication: maintenance notices, incident updates and RCAs read by non-infrastructure stakeholders.
- Available for on-call and occasional off-hours maintenance windows with US Central overlap.
Preferred Qualifications
- External Secrets Operator, Kyverno, Crossplane, Backstage.
- NATS or a similar event bus; webhook-driven automation.
- Operating Laravel/PHP-FPM with Horizon and Redis queue workloads.
- Airflow, OpenMetadata, n8n, Tableau and Redshift data-platform operations.
- Ansible, Packer, Taskfile; Trivy or Harbor image scanning.
- FinOps experience; multi-account AWS (we run iVisa and HelloGov accounts).
- Experience managing or mentoring an engineer.
- Basic Spanish.
iVisa is committed to building a diverse, inclusive, and respectful workplace. We provide equal employment opportunities to all and do not discriminate on the basis of race, gender identity, age, sexual orientation, religion, national origin, disability, or any other protected characteristic.

