Senior Site Reliability Engineer for Fuse Team
Your responsibilities a. Platform reliability and observability Own and improve the reliability posture of Fuse services, workers, APIs, queues, storage systems, and destination synchronization pipelines. Establish meaningful SLIs, SLOs, and error budgets for customer-facing APIs, asynchronous jobs, catalog data freshness, destination synchronization, and indexing. Build end-to-end observability across Data Hub item collections, from API request and job submission through processing, persistence, indexing, and downstream delivery. Ensure engineers can trace a workspace, item collection, catalog, or job across services without manually correlating disconnected logs and database records. Create and maintain actionable dashboards, alerts, and service health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, PagerDuty, and GCP tooling. Detect missing, stalled, duplicated, or inconsistent processing before customers or downstream teams report it. Improve capacity planning and autoscaling using workload telemetry, queue depth, processing throughput, latency, memory usage, storage growth, and customer-level traffic patterns. Reduce noisy alerts and replace symptom-based monitoring with signals tied to customer impact. b. Reliability of catalog storage and indexing Improve the availability, scalability, and operability of catalog data across PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, Kafka, and related storage systems. Support catalog placement, routing, index lifecycle, shard management, safe migration, and recovery across multiple Elasticsearch clusters. Develop safeguards for full replacements, delta updates, deletions, schema changes, destination changes, and catalog reindexing. Define and automate data-consistency checks between source records, transformed items, job state, Bigtable, Elasticsearch, and downstream destinations. Help establish practical platform limits and quotas for catalog size, API traffic, job concurrency, queue depth, payload size, and expensive operations. Partner with engineers on performance testing for large catalogs and high-throughput customer workloads. c. Infrastructure, deployments, and release safety Own and evolve Kubernetes configuration and operational infrastructure for Fuse components. Improve deployment automation, progressive rollout, rollback, and validation across development and production environments. Make coordinated releases safer when changes span app/app , Fuse workers, Kubernetes configuration, and PostgreSQL migrations. Automate operational procedures that currently depend on manual commands, one-off scripts, or specialist knowledge. Maintain CI/CD pipelines with tests, linters, dependency management, security checks, image publication, and release verification. Create reusable tooling for local development, ephemeral environments, end-to-end testing, load testing, and production diagnosis. Ensure runbooks remain executable and are validated through exercises rather than existing only as documentation. d. Incident management and L3 support Participate in and help improve the Fuse L3/on-call rotation . Lead incident investigation, mitigation, stakeholder communication, and follow-up for Fuse-owned systems. Use logs, metrics, traces, database state, queue state, and Kubernetes signals to diagnose failures across distributed workflows. Build safe operational tools for common support activities such as job tracing, queue inspection, rate-limit diagnosis, catalog health checks, and index recovery. Facilitate blameless incident reviews and ensure resulting actions address root causes rather than only immediate symptoms. Improve the handoff between customer support, L2, Fuse L3, Infrastructure, and dependent engineering teams. Reduce recurring support demand by turning incident knowledge into safeguards, automation, tests, dashboards, and clear documentation. e. Security, isolation, and compliance Help Fuse meet Bloomreach security and compliance requirements, including ISO and SOC 2 controls. Enforce least-privilege access, workload identity, service-level authentication and authorization, secret rotation, encryption, and auditability. Protect customer isolation across workspaces, item collections, projects, accounts, databases, indexes, buckets, and asynchronous jobs. Ensure operational tooling and incident procedures respect production-access restrictions and PII-handling requirements. Partner with engineering teams to make security controls observable and testable rather than relying on undocumented assumptions. f. Reliability by design Participate early in the design of new Fuse capabilities so reliability, recovery, observability, limits, and operational ownership are defined before implementation. Review designs for failure modes, retry behavior, idempotency, backpressure, ordering, consistency, timeout handling, cancellation, and safe rollout. Clarify ownership boundaries and service contracts with teams including Campaigns, Data Pipeline, Integrations, Discovery, Recommendations, Infrastructure, Frontend, and QA. Help teams choose architectures that balance immediate delivery with long-term operability and cost. Coach engineers in production readiness, operational testing, debugging, and sustainable on-call practices. Our tech stack Primary languages: Go, Python, SQL APIs and application: REST APIs, Python application monolith, Go workers and services, Java Infrastructure: GCP, Kubernetes/GKE, internal Kubernetes deployment tooling Databases and storage: PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, MongoDB, Redis Messaging and coordination: Kafka, ETCD, asynchronous job queues Observability: Grafana, Prometheus-compatible metrics, OpenTelemetry, GCP Logging and Monitoring, PagerDuty CI/CD and collaboration: GitLab, Jira, Confluence Testing: Go and Python unit/integration tests, API and end-to-end automation, performance testing AI-assisted engineering: Claude Code, Cursor, Copilot, Gemini CLI, or comparable to
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: job-boards.greenhouse.io