Avatar for Seekr
Seekr
Actively Hiring
Seekr puts trust into every stage of the AI lifecycle, so enterprises can build and run AI
  • Top 10% of responders
    Seekr is in the top 10% of companies in terms of response time to applications
  • Responds within a few days
    Based on past data, Seekr usually responds to incoming applications within a few days

Senior AI Infrastructure Engineer

Posted: 2 weeks ago• Recruiter recently active
Job Location
Visa Sponsorship

Not Available

RelocationNot Allowed
Hiring contact
Susannah Rafferty
Employee
image

About the job

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Senior AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large-scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high-performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving workloads ranging from edge AI deployments to trillion-parameter foundation models.

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform engineering, and production software development. You will collaborate with research scientists, software engineers, product teams, and infrastructure engineers to define the architecture and technical direction of Seekr’s AI platform.

Duties and Responsibilities

  • Design, develop, deploy, and maintain production AI infrastructure supporting model training, fine-tuning, inference, evaluation, and agentic AI workloads.
  • Design and operate scalable Kubernetes-based infrastructure supporting GPU-accelerated workloads across cloud, on-premises, hybrid, and edge environments.
  • Architect and optimize high-performance inference platforms capable of serving models ranging from resource-constrained edge deployments to trillion-parameter foundation models, with a focus on latency, throughput, scalability, reliability, and cost efficiency.
  • Build and maintain distributed systems that enable reliable scheduling, orchestration, deployment, monitoring, and lifecycle management of AI workloads.
  • Develop enterprise platforms supporting autonomous and multi-agent AI systems, including secure tool execution, orchestration, memory, evaluation, governance, and observability.
  • Design, implement, and automate AI infrastructure using Infrastructure-as-Code, GitOps, CI/CD pipelines, and modern software engineering practices.
  • Evaluate and integrate emerging AI infrastructure technologies, model serving frameworks, hardware accelerators, and cloud-native platforms to improve platform performance, scalability, and reliability.
  • Collaborate with engineering, research, product, and cross-functional teams to deliver secure, scalable, and production-ready AI platforms.
  • Lead technical design discussions, perform architecture reviews, mentor engineers, and establish engineering standards and best practices across the AI Infrastructure organization.
  • Participate in production support activities, including troubleshooting complex distributed systems, performance tuning, incident response, and continuous operational improvement.

Skills and Qualifications

  • 5–8 years of professional software engineering experience building distributed systems, cloud infrastructure, or large-scale platform services
  • Strong production ML infra experience, executes complex work independently, owns significant components
  • 4 year or higher degree or additional relevant experience, in addition to years of work experience
  • Demonstrated success designing and operating production Kubernetes environments supporting cloud-native applications and distributed services.
  • Strong software engineering skills using Python and one or more modern programming languages such as Go, Rust, or C++.
  • Proven ability to design, build, and operate production AI or machine learning infrastructure.
  • Expertise developing and optimizing large-scale AI inference platforms, including GPU utilization, distributed inference, batching, caching, quantization, and accelerator performance.
  • Familiarity with modern AI serving technologies such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar platforms.
  • Knowledge of distributed computing, networking, storage systems, cloud-native architectures, and infrastructure automation using technologies such as Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, and Infrastructure-as-Code tools.
  • Experience developing enterprise AI platforms, autonomous agents, or multi-agent systems, including orchestration, tool execution, governance, observability, and evaluation.
  • Familiarity with event-driven architectures, distributed messaging systems, and public cloud platforms including AWS, Azure, Oracle Cloud Infrastructure, or Google Cloud Platform.
  • Demonstrated technical leadership, including driving architectural decisions, mentoring engineers, and leading complex technical initiatives across cross-functional teams.
  • Demonstrated ability to analyze, profile, and optimize AI systems for performance, scalability, reliability, and cost across distributed compute environments.

About the company

Seekr company logo

Seekr

Actively Hiring
Seekr puts trust into every stage of the AI lifecycle, so enterprises can build and run AI
  • Top 10% of responders
    Seekr is in the top 10% of companies in terms of response time to applications
  • Responds within a few days
    Based on past data, Seekr usually responds to incoming applications within a few days
Learn more about Seekr image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Archesys company logo
Archesys
Improving the government services that impact everyday lives
Mobasi company logo
Mobasi
The Autonomous Digital Investigative Agent
Neuralink company logo
Neuralink
Ultra-high bandwidth brain-machine interfaces to connect humans and computers
T-Rex Solutions company logo
T-Rex Solutions
We solve our clients’ critical challenges by leveraging our innovative technical expertise
Boom company logo
Boom
Modern rental financial services for property managers and renters
LogicMonitor company logo
LogicMonitor
We expand what’s possible for businesses by advancing the technology behind them
IgniteTech company logo
IgniteTech
AI-first enterprise software that helps organizations grow revenue and transform