
Raas Infotek
Actively Hiring
Source for IT Talent
Job Location
Remote Work Policy
In office
Visa Sponsorship
Not Available
RelocationAllowed
Skills
Python
Bash
JSON
Jenkins
DynamoDB
Splunk
LAMBDA
YAML
Docker
ec2
S3
AWS
Kubernetes
DataDog
Terraform
IAM
Infrastructure as code
RDS
helm
CI/CD
AWS CloudWatch
Vpc
GitOps
GitHub Actions
Amazon EKS
ArgoCD
OpenTelemetry
Secrets Manager
Cribl
EventBridge
Harness.IO
Kinesis Firehose
OpenTelemetry Collectors
About the job
Senior SRE / DevOps Engineer – AWS
Job Summary
We are seeking a highly experienced Senior Site Reliability Engineer / DevOps Engineer with 12+ years of experience in cloud infrastructure, SRE, DevOps automation, Kubernetes, and observability. The ideal candidate will have strong hands-on experience with AWS, EKS, Terraform, CI/CD, Datadog, CloudWatch, Splunk, OpenTelemetry, and Python automation.
Key Responsibilities
- Design, implement, and maintain highly available and scalable AWS cloud infrastructure.
- Manage and optimize Amazon EKS/Kubernetes environments across development, staging, and production.
- Develop reusable Terraform modules and implement Infrastructure as Code for AWS services.
- Build and maintain CI/CD pipelines using GitHub Actions, Jenkins, and Harness.io.
- Implement blue/green and canary deployment strategies with automated rollback and approval controls.
- Establish and maintain SLIs, SLOs, and error budgets for distributed applications and microservices.
- Develop enterprise-wide observability and monitoring solutions using Datadog, AWS CloudWatch, Splunk, and OpenTelemetry.
- Design centralized log collection, routing, filtering, and transformation using Cribl and Kinesis Firehose.
- Deploy and manage OpenTelemetry Collectors for metrics, logs, and distributed traces.
- Develop automation using Python, Bash, YAML, and JSON to reduce operational toil and improve engineering efficiency.
- Support incident response, troubleshooting, root-cause analysis, runbook development, and reliability improvements.
- Implement cloud security, RBAC, IAM, compliance controls, and security logging across AWS environments.
- Collaborate with Azure teams to extend observability and monitoring across hybrid-cloud environments.
- Mentor junior engineers and participate in architecture reviews, code reviews, and technical design discussions.
Required Skills
- 12+ years of experience in DevOps, SRE, Cloud Infrastructure, or Platform Engineering
- Strong hands-on experience with AWS
- Expert-level experience with Kubernetes and Amazon EKS
- Strong Terraform / Infrastructure as Code experience
- Experience with CI/CD automation using GitHub Actions, Jenkins, or Harness
- Strong experience with Datadog and AWS CloudWatch
- Experience with Splunk / Splunk Enterprise Security
- Hands-on experience with OpenTelemetry
- Experience with Cribl and enterprise log management
- Strong understanding of SLI, SLO, SLA, and Error Budget concepts
- Strong scripting/automation skills with Python and Bash
- Experience with Docker, Helm, GitOps, and ArgoCD
- Knowledge of AWS services including EC2, EKS, S3, RDS, Lambda, VPC, IAM, Secrets Manager, DynamoDB, Kinesis, and EventBridge
- Strong incident management and production troubleshooting experience
Preferred Skills
- Microsoft Azure / Azure Databricks
- Snowflake
- Security data lake architecture
- OpenTelemetry Collector
- Blue/green and canary deployments
- Cloud governance and multi-account AWS environments
- Agile/Scrum methodology
- ITIL practices
Domain / Experience
- Enterprise Cloud Infrastructure
- Site Reliability Engineering
- Financial Services / Investment Management
- Security & Observability
- Hybrid Cloud
- Enterprise DevOps and Platform Engineering
About the company
Similar Jobs

Archesys
Improving the government services that impact everyday lives