About Starya
Starya develops AI for complex operations. We connect artificial intelligence, software, and data to transform real business processes with security, reliability, and control.
Our platform, NebulaOS, is the foundation for building, integrating, and operating these solutions. Our work combines engineering close to customers, product development, and a foundation of data, infrastructure, and security.
The challenge
Design, implement, and maintain data flows between enterprise systems and NebulaOS. Organize ingestion, transformation, and consumption models so products, analytics, and operations use information with verifiable meaning and quality. Preserve the relationship with the source of record and make history, freshness, permissions, and procedures for correcting or reprocessing a load explicit.
What you will do
- Build connectors for databases, APIs, files, or events according to source needs. Define initial loads, backfills, and change data capture (CDC), documenting read limits, continuation points, and reconciliation between loaded history and changes received during processing.
- Implement pipelines with checkpoints, deduplication, late-event handling, and bounded retries. Prepare idempotent reprocessing, rejected-record diagnostics, and schema evolution tests; record how changes to fields, types, and contracts reach consumers and how to recover an interrupted run.
- Model entities, events, and relationships with explicit granularity and agreed semantics. Preserve history and temporal validity, distinguishing event time from processing time; define how to handle corrections, deletions, and changes to records so the state relevant to an analysis or operation can be reconstructed.
- Develop transformations in SQL and Python and organize raw, processed, and consumption layers according to the design. Version code and models, document lineage, and separate reusable rules from customer adaptations, with tests and maintenance owners for each part.
- Implement completeness, consistency, uniqueness, and freshness checks for each dataset. Connect alerts to consumers and process effects; provide quality evidence and procedures to investigate discrepancies, correct transformations, and validate loads before resuming consumption.
- Keep customer systems as the sources of record for the data they originate, and perform writes through authorized integrations. Apply agreed organization isolation, access, and retention to loads, models, and copies, checking destinations and preserving traceability during corrections and reprocessing.
What we look for
- SQL knowledge for modeling, transformations, and query investigation, including joins, aggregations, granularity, and the effects of duplicates or missing values on results.
- Experience with Python for data integration and processing, organizing testable code, error handling, and automation of tasks that can be safely rerun.
- Understanding of ingestion and data contracts: incremental loads, ordering, schema changes, and failure recovery; ability to choose strategies compatible with the source.
- Practice in analytical and historical modeling, with the ability to explain keys, relationships, temporal validity, and the meaning of measures with specialists and consumers.
- Experience maintaining versioned pipelines and quality checks, using evidence to investigate discrepancies and document permissions, dependencies, and correction procedures.
Additional experience
- Experience with dbt or an equivalent approach to organize SQL transformations, tests, documentation, and dependencies, considering model evolution and consumption by different teams.
- Experience with BigQuery, ClickHouse, or equivalent analytical engines, including partitioning, physical organization, performance, and cost choices appropriate to queries and workloads.
- Experience with CDC, events, or heterogeneous enterprise integrations, especially reconciling history, late data, and differences in meaning between systems.
Your impact
Your work will be tracked through data freshness and quality-check coverage, time to make a source available, reprocessing effort, and cost per useful volume processed. Regressions after source changes and discrepancies noticed by consumers help determine improvements. Measures must account for the granularity, window, and needs of each operation.
Who you build with
You work with PM and customer specialists on semantics and acceptance criteria. With SWE, agree on API and event contracts; with FDEs, understand source-specific details. SRE contributes to capacity and recovery, and Security to access and protection requirements. Source owners authorize changes and write integrations.