Skip to content

Product Engineering

AI Research Engineer

Develop AI methods and evaluations that guide release and evolution at scale, with traceability, quality, and control for operations in regulated environments.

Talent pool

About Starya

Starya develops AI for complex operations. We connect artificial intelligence, software, and data to transform real business processes with security, reliability, and control.

Our platform, NebulaOS, is the foundation for building, integrating, and operating these solutions. Our work combines engineering close to customers, product development, and a foundation of data, infrastructure, and security.

The challenge

Investigate, implement, and evaluate AI models and systems for NebulaOS, focusing on operations that require verifiable decisions and control over data and actions. You build evidence to select methods, define limits of use, and monitor behavior as volume, concurrency, and operational diversity grow. The work connects applied research, rigorous evaluation, and engineering to sustain product evolution in regulated environments.

What you will do

  • Turn operational problems and risks into hypotheses, comparison baselines, and acceptance criteria with Product, QA, and domain specialists. Measure correct outcomes, omissions, unauthorized actions, and abstention; evaluate escalation criteria and the effectiveness and workload of human review, considering the consequence of error in each task.
  • Build datasets and evaluation protocols with authorized provenance and use, versions, and annotation criteria. Include frequent, rare, and critical cases; separate development and test sets, examine contamination and context-specific slices, and preserve isolation and protection of sensitive data.
  • Evaluate models, retrieval-augmented generation (RAG), context, and tool use as parts of a system. Test unsupported responses, conflicting sources, dependency failures, and incorrect actions in controlled environments, verifying task outcomes and their effects with Engineering and QA.
  • Combine automated and human evaluations. Calibrate model-based evaluators against expert judgments, investigate disagreements, and estimate variability and uncertainty. Document evaluation coverage and limits, including when available evidence is insufficient to recommend broader use.
  • Design scale experiments with SWE and SRE to measure concurrency, throughput, tail latency, errors, and timeouts under representative loads. Calculate cost per completed task, including failed attempts and retries. Compare inference, models, and context, checking whether performance gains preserve quality and controls in critical slices.
  • Version code, datasets, retrieval corpora, prompts, and configurations; record external model versions or identifiers and limits on reproducibility. Link results to the evaluated artifacts and deliver experiments, evaluation suites, and evidence that support review and release of changes, without exposing sensitive content in logs or reports.
  • Monitor evaluations in production using authorized data and samples, comparing observed behavior with validated behavior. Investigate distribution and quality shifts, incorporate failures into tests, and propose criteria for gradual expansion, suspension, or rollback with service owners.

What we look for

  • Foundations in machine learning, statistics, and experimental design, with the ability to formulate hypotheses, choose baselines, analyze samples, and interpret uncertainty, bias, and measurement limitations.
  • Practical experience with Python and software engineering to implement experiments and evaluation mechanisms, organize tests and dependencies, and produce code that others can run, review, and maintain.
  • Experience evaluating AI models or applications, with task-level metrics, versioned test sets, error analysis, and comparisons across versions. Understand contamination, coverage, human evaluation, and the limits of automated evaluators.
  • Ability to investigate AI system performance and relate quality to latency, concurrency, resource consumption, and cost. Translate experimental results into execution and validation requirements with engineering.
  • Practice in traceability and change control, with attention to authorized access, restricted data, and evaluation records. Ability to explain which version was evaluated, with what data, criteria, and limits of use.
  • Ability to read technical papers, turn methods into executable comparisons, and discuss results with domain specialists. Communicate uncertainty and operational consequences in a way that supports shared product and release decisions.

Additional experience

  • Experience with AI in healthcare, insurance, financial services, or other regulated environments, working with specialists on criteria, human review, validation evidence, and monitoring changes.
  • Experience with inference at scale, model adaptation, or optimization, including experiments comparing efficiency and quality under representative loads, in collaboration with product and infrastructure engineering.
  • Experience with adversarial evaluations, calibration of automated evaluators, or human evaluation programs, investigating rare errors, disagreements, and multi-step agent behavior.

Your impact

Your impact will be assessed through evidence-supported decisions about product evolution: reproducible evaluations, important regressions detected before broader use, and clear limits of use. Quality by task and slice, critical errors, and cost and latency as load grows guide this assessment. Each change must be linked to the version, criteria, and results that supported its release.

Who you build with

AI Research develops and maintains methods, evaluators, and quality criteria; QA integrates these evaluations into journey regression testing. SWE incorporates capabilities into the product, and SRE collaborates on infrastructure and scale experiments. Data supports evaluation datasets; Security and domain specialists help define risks and controls. Product and service owners share decisions on use and release.

Product Engineering

Other roles on this team

  • Product Engineering

    Product Designer

    Design the operational experience of NebulaOS so people can configure agents and workflows, monitor executions, and intervene with clarity when a situation requires a human decision.

  • Product Engineering

    Product Manager

    Lead discovery, priorities, and adoption for NebulaOS, connecting operational problems to product choices and results that users and teams can recognize.