Avatar for Spotline Software Solutions
IT Consulting

Data Center SRE

  • |8 years of exp
  • |Full Time
Posted: 1 week ago• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
SQL
MySQL
PostgreSQL
Splunk
Shell Scripting
RESTful APIs
Prometheus
Grafana
NetBox
IPMI

About the job

Job Title: Data Center SRE

Location: Austin, TX (Onsite 5 days/week)

Duration: 12+ Months (with possible extension)

Job Description

We are seeking an experienced Sr. Data Center Site Reliability Engineer to automate operations and maximize the uptime, efficiency, and scalability of data center, facility power/cooling infrastructure, and software automation. In this role, you will manage, monitor, and optimizing both server reliability and the critical power and cooling infrastructure that sustains our distributed production systems.

Key Responsibilities

  • Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.
  • Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.
  • Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.
  • Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.
  • Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.
  • Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.
  • Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.
  • Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.

Qualifications

  • Experience: 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.
  • Education: Bachelor’s Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.
  • Bare-Metal & Hardware Expertise: Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.
  • Inventory & Asset Management: Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.
  • Data & API Capabilities: Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.
  • Monitoring & Tooling: Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.
  • Infrastructure Protocols: Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.
  • Facilities Knowledge: Practical understanding of data center physical infrastructure, specifically power feed distribution systems and cooling infrastructure (e.g., HVAC, liquid cooling, hot/cold aisle containment, air handling units).
  • Automation: Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

About the company

Founders

Vinod Kumar
Founder
Delhi
image
View the team image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Take-Two Interactive Software company logo
Take-Two Interactive Software
Game development needs creativity, but creativity needs an environment to be nurtured in
Neuralink company logo
Neuralink
Ultra-high bandwidth brain-machine interfaces to connect humans and computers
T-Rex Solutions company logo
T-Rex Solutions
We solve our clients’ critical challenges by leveraging our innovative technical expertise
Boom company logo
Boom
Modern rental financial services for property managers and renters
LogicMonitor company logo
LogicMonitor
We expand what’s possible for businesses by advancing the technology behind them
IgniteTech company logo
IgniteTech
AI-first enterprise software that helps organizations grow revenue and transform