Avatar for The Ksquare Group
The Ksquare Group
Actively Hiring
Provides tech solutions & digital transformation via consulting

Senior Technical Lead

Posted: 2 weeks ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
SQL
DNS
TCP/IP
API Integration
Splunk
Shell Scripting
New Relic
PowerShell
Ansible
routing
Kubernetes
DataDog
Firewalls
Terraform
Tls
Dynatrace
Elastic
Prometheus
Grafana
Ports
Load Balancers
helm
Proxies
Mutual TLS
OpenTelemetry
Certificates
VPNs
Bindplane

About the job

The Senior Enterprise Observability Lead will be responsible for designing, implementing, configuring, and operationalizing Health’s enterprise observability platform across cloud, on-premises infrastructure, networks, applications, integrations, data platforms, and AI workloads.

The current ecosystem includes BindPlane as the collector layer, ClickHouse as the observability datastore, and Langfuse for AI and LLM monitoring.

This is a hands-on technical leadership role. The individual must be capable of owning the architecture while directly configuring collectors, building telemetry pipelines, writing queries, implementing dashboards and alerts, automating deployments, and troubleshooting production issues.

Key Responsibilities

  • Own the overall enterprise observability architecture, standards, implementation roadmap, and technical governance.
  • Assess the current environment and identify gaps in coverage, scalability, reliability, performance, security, and operations.
  • Install, configure, upgrade, and troubleshoot BindPlane and OpenTelemetry collectors.
  • Configure telemetry sources, destinations, processors, filters, transformations, routing rules, and enrichment.
  • Build telemetry pipelines for logs, metrics, traces, events, and infrastructure data.
  • Develop ClickHouse schemas, queries, materialized views, aggregations, retention policies, and performance optimizations.
  • Create infrastructure, network, application, API, middleware, database, and business-service dashboards.
  • Configure alerts, thresholds, anomaly-detection rules, event correlation, and escalation workflows.
  • Troubleshoot collector failures, missing telemetry, dropped data, ingestion delays, schema issues, and query-performance problems.
  • Develop automation using Python, PowerShell, shell scripting, Terraform, Ansible, Helm, or comparable tools.
  • Integrate observability configuration, instrumentation, testing, and validation into CI/CD processes.
  • Design observability across physical servers, virtual machines, Windows and Linux environments, VMware, cloud platforms, containers, and Kubernetes.
  • Work directly with networking teams on firewall rules, DNS, routing, proxies, load balancers, VPNs, ports, TLS, mutual TLS, and certificates.
  • Establish service and dependency mapping across infrastructure, applications, integrations, databases, and networks.
  • Support observability for AI, LLM, RAG, and agentic applications using Langfuse and related platforms.
  • Implement monitoring for prompt and response traces, model latency, token consumption, costs, failures, retrieval performance, and agent execution.
  • Ensure PHI, PII, credentials, prompts, responses, tokens, and other sensitive data are masked, filtered, or excluded.
  • Review existing Splunk use cases and determine which dashboards, alerts, searches, and monitoring capabilities should be migrated, retained, redesigned, or retired.
  • Lead technical workshops, architecture reviews, implementation planning, and client discussions.
  • Review the configurations, queries, dashboards, automation scripts, and deliverables produced by the LATAM and India engineers.
  • Participate directly in deployments, production support, incident resolution, and root-cause analysis.
  • Prepare architecture diagrams, implementation guides, operational runbooks, and knowledge-transfer documentation.

Required Qualifications

  • Ten or more years of experience in infrastructure, cloud, DevOps, SRE, platform engineering, or enterprise monitoring.
  • At least five years of hands-on experience implementing enterprise observability solutions.
  • Demonstrated experience configuring telemetry collectors, building pipelines, creating dashboards, writing queries, and implementing alerts.
  • Strong hands-on experience with OpenTelemetry.
  • Strong infrastructure knowledge across physical servers, virtual machines, Windows, Linux, storage, VMware, containers, Kubernetes, and cloud platforms.
  • Strong networking knowledge covering TCP/IP, DNS, firewalls, routing, proxies, load balancers, VPNs, ports, certificates, TLS, and mutual TLS.
  • Experience with BindPlane, Dynatrace, Splunk, Grafana, Prometheus, Elastic, Datadog, New Relic, or comparable technologies.
  • Experience with high-volume telemetry pipelines and analytical datastores.
  • Strong SQL, scripting, API integration, and automation capabilities.
  • Ability to lead technical client discussions while independently performing implementation activities.
  • Strong troubleshooting, communication, and documentation skills.

Preferred Qualifications

  • Hands-on experience with BindPlane and ClickHouse.
  • Experience with Langfuse and AI or LLM observability.
  • Experience monitoring RAG and agentic AI solutions.
  • Experience migrating or rationalizing Splunk monitoring use cases.
  • Healthcare or regulated-industry experience.
  • Understanding of HIPAA, PHI protection, and healthcare security controls.

About the company

The Ksquare Group company logo

The Ksquare Group

Actively Hiring
Provides tech solutions & digital transformation via consulting51-200 Employees
Learn more about The Ksquare Group image