Avatar for Netskope
Netskope
Actively Hiring
Netskope is redefining cloud, network, and data security
  • B2B
  • Scale Stage
    Rapidly increasing operations
  • Top Investors
    This company has received a significant amount of investment from top investors
  • +4

Platform Engineer (Agent Runtime)

Posted: 1 month ago
Job Location
Visa Sponsorship

Not Available

RelocationNot Allowed
Hiring contact
Brendan Lynch
Talent Management • 12 years
San Francisco
image

About the job

Every agent on this platform runs inside its own isolated environment, and someone has to make sure that environment actually behaves the way it's supposed to under real load, not just in a clean test run. As a Senior Agent Runtime Engineer, you'll own the health of the systems agents run on: keeping sessions isolated from each other, catching the failure modes that only show up under real concurrency, and making sure a slow or expensive agent gets caught before it becomes everyone's problem. You'll work closely with the engineers building the agents themselves, since you're often the first call when something behaves strangely in a real environment but not in a test one. If you like being the person who understands a system well enough to know exactly why it's misbehaving, this is that job.

Skills and competencies:

  • Own the health of the agent execution environment day to day — session isolation, resource limits, and catching the failure modes that only appear once real concurrency and real traffic patterns show up, not just in a clean test run.
  • Investigate and resolve cases where an agent behaves differently in production than it did in testing, working directly with the Agent Engineer or Quality Engineer who built it to figure out whether the problem is the agent's design or the environment it's running in.
  • Tune cold-start and concurrency settings for the platform's critical-path functions, and review them on a regular cadence as usage patterns shift rather than setting them once and forgetting them.
  • Understand the platform's session isolation model well enough to reason about its limits — where isolation is guaranteed by the underlying compute layer, and where the platform has to add its own controls because that guarantee doesn't fully hold.
  • Build and maintain the observability that lets someone answer "is this agent actually working correctly," not just "is it technically up" — tracing, behavioral drift detection, and quality signals sitting alongside the usual metrics and logs.
  • Track per-agent and per-session cost and efficiency, and flag agents that are burning more tokens, calling more tools, or running longer than the task should reasonably require.
  • Deploy and manage runtime-layer infrastructure resources through existing CI/CD pipelines as needed — provisioned concurrency settings, runtime-specific IAM roles, observability configurations, etc.
  • Run and improve the tests that validate isolation actually holds — for example, confirming that one agent's session genuinely can't reach or affect another's, not just assuming it because the platform is supposed to guarantee it.

Must-Have:

  • At least 5 years in a platform, SRE, or infrastructure engineering role, with real production experience running serverless or containerized workloads on AWS (Lambda, Fargate, Docker, Kubernetes or equivalent) at meaningful scale.
  • Hands-on experience with AWS observability tooling (CloudWatch, X-Ray, or a comparable distributed tracing stack) — able to go from "something's wrong" to a root cause using traces and logs, not just dashboards.
  • Real experience deploying and managing infrastructure through CI/CD independently — comfortable owning IaC changes (Terraform, CDK, or similar) end to end rather than handing them off to someone else.
  • Working understanding of compute isolation concepts (containers, microVMs, or similar sandboxing models) and where their guarantees actually stop, since a lot of this role is reasoning about the edges of what a platform promises versus what it might not fully cover on its own.
  • Comfort investigating cost and performance problems at a granular level — able to trace an unexpectedly expensive or slow workload back to a specific cause, not just flag that costs went up.
  • Strong incident response instincts: staying calm and methodical while root-causing a live production issue, and following through with an actual fix rather than a workaround. Strong Advantage:
  • Direct experience with AWS Bedrock, SageMaker, or any managed AI/agent runtime platform.
  • Experience with cold-start or concurrency tuning specifically (Lambda provisioned concurrency, container warm pools, or similar).
  • Exposure to a security-conscious or regulated environment where infrastructure changes go through a formal review or approval process.
  • Familiarity with LLM-specific cost drivers (token pricing, tool-call volume, model tiering) even if it wasn't the primary focus of a past role.

#LI-CV1

About the company

Netskope company logo

Netskope

Actively Hiring
Netskope is redefining cloud, network, and data security1001-5000 Employees
Company Size
1001-5000
Company Type
Startup
Company Type
SaaS
Company Type
Enterprise Security
Company Industries
B2B · SaaS · Mobile · Artificial Intelligence / Machine Learning
  • B2B
  • Scale Stage
    Rapidly increasing operations
  • Top Investors
    This company has received a significant amount of investment from top investors
  • Valuation $1B+
    This company has a valuation of $1B or more
  • 4.2
    Highly rated
    Netskope is highly rated on Glassdoor, with 4.2 out of 5 stars
  • 4.1
    Work / Life Balance
    Employees rate Netskope 4.1/5 on Glassdoor for work / life balance
  • 4.1
    Strong Leadership
    Employees rate Netskope 4.1/5 on Glassdoor for faith in leadership
Learn more about Netskope image

Funding

AMOUNT RAISED
$1.2B
FUNDED OVER
10 rounds
Rounds
Co
$401000000
Series Convertible Note - Jan 2023+9

Perks

Insurance, Health & Wellness
● Medical (UHC-HDHP, PPO, & EPO; CA Kaiser- HMO & HDHP) ● Dental ● Vision ● Equitable Life & AD&D Insurance ● Short & Long Term Disability ● Company HSA Contributions ● Employee Assistance Program (EAP)
401(k) Retirement Savings Plan
401(k)/ROTH offering through Newport Group ($20,500 Annual Max Contribution / $6,500 Annual Max Catch-Up). Eligible to start deferring after the 1st paycheck.
Voluntary Life Insurance
Employees have the option to enroll in supplemental life insurance up to a max of $250,000. Netskope is pleased to provide spouse and dependent life insurance offerings upon an employee’s insurance election.
Paid Parental Leave
12 weeks Birth Parent Paid Parental Leave 8 weeks Non-Birth Parent Parental Leave
Commuter Benefits
Employees can contribute up to $280/month to a pre-tax account for mass transit and/or parking.
Additional Netskope Perks
● 13+ Company Observed Holidays ● Quarterly Global Wellness Days ● Unlimited Paid Time Off ● Discount program for popular brands, 30,000 national/local offers, and devices ● Meditation Hours ● Family Planning Assistance ● Travel Assistance

Founders

Krishna Narayanaswamy
Founder
image
Sanjay Beri
Founder
image
View the team image