Skip to content

Data, Infrastructure & Security

Site Reliability Engineer (SRE)

Build infrastructure, automation, and observability to deploy and operate NebulaOS, recover from failures, and keep capacity and costs appropriate to usage.

Talent pool

About Starya

Starya develops AI for complex operations. We connect artificial intelligence, software, and data to transform real business processes with security, reliability, and control.

Our platform, NebulaOS, is the foundation for building, integrating, and operating these solutions. Our work combines engineering close to customers, product development, and a foundation of data, infrastructure, and security.

The challenge

Build mechanisms and practices for operating NebulaOS with observable behavior and known recovery paths. Connect infrastructure, services, and dependencies to the journeys they serve, using evidence to guide changes, sizing, and reductions in manual work. Work with component developers so that diagnosis, recovery, and cost are part of ongoing development.

What you will do

  • Define service level indicators (SLIs) and reliability objectives (SLOs) with journey owners. Document success, failure, the measurement window, and dependencies; use the agreed error budget to discuss the pace of change and improvement priorities with teams.
  • Instrument metrics, traces, and technical logs to connect symptoms to the relevant service, queue, or integration. Build actionable alerts, dashboards, and diagnostic procedures without recording personal data, credentials, prompts, or conversation content, distinguishing a submitted attempt, an unknown outcome, and confirmed completion.
  • Automate provisioning and configuration with infrastructure as code, covering networking, identity, resources, and dependencies. Maintain versioned modules, change review, plan validation, and drift detection between declared configuration and the environment, with reproducible procedures for creating or recovering components.
  • Implement progressive deployment with promotion, observation, and suspension criteria. Prepare rollbacks with each service’s maintainers, checking configuration and schema compatibility; for changes with persistent effects, agree on reconciliation or compensation and test the procedure before broader rollout.
  • Exercise backup restoration and service recovery with representative scenarios, recording the recovered state, dependencies, and limitations found. Participate in incident response and investigation, turning recurring issues into actions with owners and evidence of completion alongside component owners.
  • Plan capacity based on demand, queues, saturation, and latency. For inference services, analyze concurrency, resource usage, and cost per useful execution with Engineering and Research; compare sizing, limits, and optimizations using representative loads and expected service quality.

What we look for

  • Knowledge of Linux, networking, and containers to investigate communication, resources, and processes, relating application behavior to execution environment conditions.
  • Experience with Kubernetes or equivalent orchestration and declarative infrastructure, understanding deployment, configuration, permissions, dependencies, and component recovery.
  • Ability to automate operations in Python, Go, or TypeScript, with versioned code, tests, and failure handling for repeatable tasks that can be investigated.
  • Practice in observability and incident analysis, using operational signals to formulate hypotheses, locate causes, and choose recovery and prevention actions.
  • Experience collaborating on production changes, connecting availability, compatibility, and capacity to release criteria and rollback or reconciliation procedures.

Additional experience

  • Experience with AWS, OCI, or equivalent clouds and Terraform or similar tools, including reusable modules, plan review, and dependency management.
  • Experience with OpenTelemetry, Datadog, or equivalent instrumentation and diagnostic solutions, considering cardinality, collection costs, and the usefulness of signals to teams.
  • Experience operating inference services, accelerators, or data-intensive workloads, investigating queues, latency, memory, and consumption to guide sizing and optimization.

Your impact

Impact will be tracked through journey reliability objectives, incident duration and recurrence, recovery demonstrated in tests, and recurring manual effort. Latency, saturation, and cost per useful unit of work help assess capacity. Targets and windows will be agreed per service, considering dependencies and consequences for NebulaOS users.

Who you build with

Production is shared work with the teams building services and integrations. SRE develops mechanisms and participates in diagnosis; component owners implement fixes and maintain procedures. You collaborate with Data on load recovery, Research on model execution, and Security on operational controls, aligning priorities with leads.

Data, Infrastructure & Security

Other roles on this team