Data Engineer
🌍 Hello World! We are The Codest - International Tech Software Company with tech hubs in Poland delivering global IT solutions and projects. Our core values lie in “Customers and People First” approach that prioritises the needs of our customers and a collaborative environment for our employees, enabling us to deliver exceptional products and services. Our expertise centers on web development, cloud engineering, DevOps and quality. After many years of developing our own product - Yieldbird, which was honored as a laureate of the prestigious Top25 Deloitte awards, we arrived at our mission: to help tech companies build impactful product and scale their IT teams through boosting IT delivery performance. Through our extensive experience with product development challenges, we have become experts in building digital products and scaling IT teams. But our journey does not end here - we want to continue our growth. If you’re goal-driven and looking for new opportunities, join our team! What awaits you is an enriching and collaborative environment that fosters your growth at every step. Project description: We are building the data platform that powers a Demand-Side Platform (DSP) operating at real-time bidding (RTB) scale. Our stack ingests billions of auction events per day from a Kafka-based internal event bus via Google Pub/Sub, processes them through Apache Beam streaming pipelines, and lands data into BigQuery - from which we build a full analytical warehouse serving dashboards, attribution models, audience pipelines, and the bidder itself. We are looking for a Junior OR Mid-level data engineer to take ownership of pipeline components, drive architectural decisions, and collaborate across mutliple teams. The role: What You Will Work On Streaming pipelines & data ingestion Build and maintain Apache Beam (Dataflow) streaming pipelines that consume events from Pub/Sub and land them into BigQuery, implementing efficient parsing techniques to handle high volume cost-effectively Apply the correct streaming patterns to ensure resilience, data integrity, and strict deduplication Implement incremental and merge load strategies in dbt: detailed incremental filters utilizing partition pruning and time ranges to scan only the necessary data blocks, maximizing query performance and ensuring cost optimization; perform MERGE actions for state synchronization of dimension tables Integrate data from multiple source systems using highly performant ingestion processes and optimal database schemas Data warehouse & transformation Design and implement dbt models across staging, warehouse, and marts layers, following the Medallion architecture Build aggregation and mart tables (hourly campaign aggregates, daily creative stats, funnel metrics) powering dashboards and the Panel UI PostgreSQL export Own the attribution pipeline : bucket accumulator tables, time-decay scoring at conversion time, product hierarchy cascade, config versioning Audience & ID graph Build audience activation pipelines in Airflow + dbt that resolve simple and compound audience segments, join them to ID graph clusters, and export to Couchbase Keep audience and ID graph documents in Couchbase in sync with upstream changes (batch baseline + incremental streaming updates) Schema & cross-language contracts Design Couchbase document schemas (audience, ID graph clusters, reverse mappings, activity events) shared between Python pipelines and the Go bidder Maintain YAML JSON Schema as the single source of truth; codegen produces Pydantic v2 models for Python services and Go structs for the bidder — schema changes require regenerating both artefacts and updating all consumers Infrastructure & CI/CD Contribute to Terraform infrastructure (Pub/Sub topics with dead-letter, Dataflow worker configurations, GCS buckets, KMS keys, Secret Manager secrets, Artifact Registry) Maintain CI/CD pipelines : automated linting (sqlfluff, pre-commit), DAG syntax validation, schema contract checks, containerised Dataflow worker builds and releases to Artifact Registry via GitHub Actions Write Architecture Decision Records (ADRs) and review PRs Streaming & Data Engineering Concepts You Must Know Streaming Fundamentals (Apache Beam / Dataflow) Event time vs processing time — events are produced at one time and arrive later; all business logic must use event time; processing time is only for system metrics Watermarks — Beam’s estimate of how far behind event time the pipeline is; when the watermark advances past a window boundary, that window is considered complete and results are emitted; a watermark that stalls means the pipeline is backlogged Windowing — grouping an unbounded stream into finite buckets for aggregation: Tumbling (fixed, non-overlapping) — e.g. hourly campaign spend buckets Sliding (overlapping) — e.g. rolling 7-day reach Session (gap-based) — e.g. user activity sessions with inactivity timeout Triggers — control when partial or final results fire out of a window before it closes; early firings give low-latency approximations; late firings correct for late-arriving data Late data & allowed lateness — data arriving after the watermark has passed; we allow up to 3 hours of lateness and re-emit corrected window results when they arrive State and timers — Beam stateful transforms maintain per-key state across elements; used for enrichment joins, deduplication caches, and session stitching Kafka / Pub/Sub Messaging Model Understanding the differences between Kafka and Pub/Sub is required: Kafka concepts : topics, partitions, consumer groups, offsets, offset commit, compacted topics (for changelog/CDC), retention by offset or time Pub/Sub concepts : topics, subscriptions (pull vs push), message acknowledgement, ack deadline, subscription backlog, oldest unacked message age Key difference : Kafka consumers own their offset (replay is free); Pub/Sub delivers to any subscriber and relies on ack to determine progress — a message not a
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: thecodest.recruitee.com