Napravite profil kako bi poslodavci mogli da vas pronađu, da bi dobijali odgovarajuće poslove i brže se prijavljivali.
  • Pretraga posla
  • Omiljeno
  • Napravi CV
    Novo
  • Upisi

Lead Site Reliability Engineer - Imunify Reliability Platform (remote-only)

Cloudlinux

The problem you'd own

Imunify360 is a multi-layer Linux server security suite — WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, reputation — running as an agent on hundreds of thousands of customer servers, backed by a cloud estate of scanning, correlation and signature-delivery services on our own bare metal.

Roughly 70 components currently ship without a defined service level indicator. Some are internal services we can scrape. Many are agent-side subsystems running on machines we do not own, reporting through a heartbeat we designed for something else. There is monitoring, and there are dashboards, and there is no coherent answer to the question "is this component doing its job right now, and how would we know if it stopped?".

We know the cost of that gap precisely, because we recently paid it: a security control was silently disabled across a large fraction of the fleet for 61 days. Every dashboard was green. The telemetry reported a ruleset version but not whether the control that consumed it was switched on, so a configuration change was indistinguishable from a broken updater. Three independent safety mechanisms existed and all three were gated behind the same condition that caused the failure.

Your job is to make that class of failure detectable in hours instead of months, across the whole product line, and to build the system that keeps it detectable as the product changes.

This is a greenfield charter inside a brownfield estate. You are not inheriting an SRE team, an SLO framework or a paging culture. You are defining them, with the engineering leads, and then making them stick.

What you'll do

1. Define what "working" means for ~70 components

  • Run SLI definition with squad leads and senior engineers. You facilitate and hold the standard; the owning squad signs the SLI.
  • Build the taxonomy this product actually needs, which is broader than availability and latency:
    • Service SLIs — availability, latency, error rate for cloud-side services.
    • Fleet SLIs — heartbeat reachability, version and configuration convergence across the installed base.
    • Control-efficacy SLIs — the differentiator. What fraction of protected units have the control effectively enabled and current , not merely installed. Ruleset generation drift, signature age, scan coverage, enforcement-mode distribution.
    • Delivery SLIs — artifact publish success, rule-to-fleet lead time, hotfix time-to-convergence.
    • Pipeline SLIs — ingest lag, verdict latency, queue age, backlog burn.
  • Enforce one non-negotiable design rule: an SLI must be measurable from outside the gate of the thing it measures. If the control being off also switches off the signal that would tell you it is off, the SLI is invalid. This is the lesson of the incident above and it is the reason this role exists.
  • Attach an SLO, an error budget and an owning squad to each. Tiering is expected — not every component earns a 99.9% target or a pager.

2. Build the collection system

  • Design and build the pipeline that gets these indicators off the fleet and into a queryable store: push-based, sampled, privacy-constrained, and with a cardinality budget you set and defend.
  • Extend agent-side and service-side instrumentation where the signal does not exist yet, in Python, Go and Rust, working with the owning squads.
  • Consolidate the current sprawl of dashboards, ad-hoc queries and reporting paths into a defensible set of instruments, and retire what does not earn its keep.

3. Build alerting and alert management

  • Symptom-based, SLO-anchored alerting with multi-window burn-rate semantics. Not threshold soup.
  • A three-tier taxonomy — page / ticket / dashboard — with an explicit rule for what is allowed to page a human at 03:00.
  • Every alert ships with an owner, a runbook and a documented failure mode, or it does not ship.
  • Alert hygiene as a standing practice: quarterly review, deletion counted as a win, actionable-rate tracked. A persistent inability to perform a security-relevant refresh should page. It currently logs a warning.

4. Build escalation

  • Component → owning squad ownership map, kept current, machine-readable, and wired into routing so an alert reaches the right seven people rather than a shared channel.
  • Severity matrix, acknowledgement SLAs, follow-the-sun rota design across UTC−5 … UTC+8, and clean handoff protocol.
  • Incident command practice and blameless postmortems within 24 hours. We already do postmortems and do them honestly, including publicly retracting our own wrong findings; you raise the floor on the mechanical parts — timelines, ownership, action-item follow-through.
  • Design the escalation system so that squads carry their own pagers. You build and operate the platform and coach on the practice; you are not the buffer that absorbs everyone else's alerts.

Requirements

What you'll bring

Required (Must-haves):

  • Substantial production-engineering or SRE experience, including at least one environment where you defined the SLO framework rather than inherited it. We will ask you to walk through SLIs you personally wrote and how you negotiated them with resistant teams.
  • Strong Python. Comfortable reading and modifying Go or Rust — our agents are written in them and instrumentation lands there.
  • Deep practical grip on time-series and event telemetry at scale: Prometheus/OpenMetrics, Grafana, an Alertmanager-class routing layer, and a columnar store for high-cardinality fleet data (ClickHouse or equivalent).
  • Distributed systems debugging on bare metal and long-lived hosts. Most of this estate is not Kubernetes, and the reflexes that assume an orchestrator will not transfer cleanly.
  • Configuration management and CI at production scale — Ansible, GitLab CI, Jenkins or close equivalents.
  • The judgement to design measurement for machines you do not own and cannot scrape: push telemetry, sampling, clock skew, partial reporting, and the privacy constraints that come with running on a customer's server.
  • Written communication that holds up async. This role is 40% telemetry engineering and 40% getting sixty engineers to agree on what "healthy" means; the remaining 20% is refusing to let the answer be a dashboard nobody reads.

Valuable (Nice-to-haves):

  • Security product background — WAF, EDR, AV, vulnerability management — and the instinct that a security control's SLI is about enforcement, not uptime.
  • Monitoring under audit: SOC 2 CC7.x, ISO 27001 A.8.16, NIST SP 800-137 continuous monitoring. Some of this work is audit evidence and it helps if you have written for that audience.
  • OpenTelemetry, eBPF, Sentry.
  • Cost- and cardinality-aware telemetry design.
  • Fluency with agentic development tooling — we run a Cursor/Claude-first SDLC with internal and third-party MCP servers, and engineers here are assessed on how well they work with it.
  • Kubernetes, for the one workload that is on it.

Not this role

  • Not a DevOps ticket queue, not build-system ownership, not cloud cost management, not the on-call rota for other squads' services.

First year, in outcomes

30 days: Component inventory with named owners. SLI taxonomy and tiering agreed. 3 pilot components fully instrumented end to end as the reference implementation.

90 days: Collection pipeline in production. Tier-1 components (the ones whose failure is a customer security exposure) carry SLO, alert, runbook, owner. Escalation routing live for tier-1.

180 days: All ~70 components have a defined SLI and an owner. Alert taxonomy enforced; page volume and actionable-rate measured and published. Squad on-call operating.

365 days: Mean time to detect a silent control-degradation is under 24 hours, measured, against a 61-day baseline. Error-budget policy influences release decisions. The function is documented well enough that hire #2 and #3 are additive, not archaeological.

How we work

Remote-first and async across nine time zones. Weekly PO sync and architecture sync; monthly demo and OKR review; quarterly architecture summit. Decisions land as ADRs. Every output carries an owner and a due date. Postmortems are blameless and published, and we correct ourselves on the record when we get something wrong.

Benefits

What's in it for you?

  • A strong focus on professional development with opportunities for learning and growth:
    • Interesting and challenging projects,
    • Mentor and other knowledge-exchange programs;
  • Fully remote work with flexible working hours, that allows you to schedule your day and work from any location worldwide;
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves to ensure you maintain a healthy work-life balance;
  • Compensation for private medical insurance;
  • Co-working and gym/sports reimbursement;
  • The opportunity to receive a reward for the most innovative idea that the company can patent, fostering a culture of creativity and innovation.

By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (), which provides detailed information on how we maintain and handle your data.

Oglas je objavljen pre 5 sati
Slični poslovi
  •  ...CloudLinux is a global remote-first company. We are driven by our...  ...stability for hosting providers. Imunify is an innovative security...  ...We are looking for a DevOps Engineer to join the Patchman team....  ...integrate it with our broader platform. This includes tackling technical... 
    Rad na daljinu
    Fleksibilno radno vreme

    Cloudlinux

    Beograd
    pre 24 dana
  •  ...through cohort-based programs built on our own platform and curriculum, developed in partnership with Nebius AI. Our team is fully remote and globally distributed, and we serve a...  ...We are looking for a Learning Experience Lead to design end-to-end learning experiences,... 
    Rad na daljinu

    TripleTen

    Beograd
    pre 1 dan
  •  ...Positions Overview IGT is seeking a Staff Platform Engineer to help shape the future of software...  ...teams to deliver software faster, more reliably, and at scale. The ideal candidate...  ...distributed systems, cloud platforms, DevOps, Site Reliability Engineering, Platform... 
    Predloženo
    Hibridni rad

    IGT, a Nevada Corporation

    Beograd
    pre 3 dana
  •  ...currently looking for a Senior Java Software Engineer (Team Lead) with strong technical expertise,...  .... Their teams operate in a fully remote environment, fostering flexibility, autonomy...  ...developing, and evolving our Java-based SaaS platform. In addition to being a hands-on... 
    Rad na daljinu
    Puno radno vreme
    Rad u kancelariji

    Holycode

    Beograd
    pre mesec dana
  •  ...includes social entertainment platforms designed to connect people...  ...team of digital nomads works remotely from all over the world. We...  ...looking for an experienced Lead Graphic Designer to drive creative...  .... ~ Experience with prompt engineering, reference-based generation,... 
    Rad na daljinu
    Rad od kuće
    Frilens
    Puno radno vreme
    Rad u kancelariji
    Rad sa bilo kog mesta
    Rad od kuće

    Social Discovery Group

    Beograd
    pre 21 dan
  •  ...our team shares a love of the outdoors and a desire to protect it for future generations.  Role Summary As a Senior Platform Engineer - Site Reliability Engineering, you will help build and deploy the self-service tools and resilient systems that allow our engineering... 
    Ugovor o radu

    Rivian

    Beograd
    pre 26 dana
  • 10.000 - 15.000 $

     ...This is a remote position. We are looking for a Content Manager & Content Quality Lead to help us improve the quality, efficiency, and scalability of our content production...  ...Upload and format articles in CMS platforms such as WordPress Assist with publishing... 
    Rad na daljinu
    Frilens

    AWISEE

    Beograd
    pre 1 dan
  •  ...We are looking for an experienced and performance-driven Team Lead to lead our IT, DE and TR sales team — our most strategically important...  ...Competitive salary with performance-based bonuses ~100% fully remote (if Belgrade hire then hybrid) ~ Full onboarding and ongoing... 
    Rad na daljinu
    Puno radno vreme
    Hibridni rad
    Popodnevna smena

    Best Service Team

    Beograd
    pre 24 dana
  •  ...Join FxPro : a leading international fintech company. Be a part of our...  ...of our success story. As a Lead SEO , you will own the growth...  ...on-page SEO audits to optimize site performance and health Monitor...  ...package ~ Flexibility of remote work ~ All the equipment you... 
    Rad na daljinu
    Puno radno vreme
    Rad u kancelariji
    Smenski rad

    FxPro Financial Services Limited

    Beograd
    pre 10 dana
  •  ...Rivian is seeking a Senior Cybersecurity Engineer, Identity DevOps and Platform Automation to help design, build,...  ...support. You will help deliver reliable, auditable IAM services across Microsoft...  ...onsite/hybrid at a Rivian location; remote is not available). Participation in... 
    Rad na daljinu
    Ugovor o radu
    Hibridni rad
    Dežurstvo

    Rivian

    Beograd
    pre 26 dana
  •  ...dynamic and collaborative workplace. We are on the lookout for a  Lead Brand Designer to join our  Design Department. Role overview...  ...influences engagement and conversion. Benefits Remote work opportunity. Flexible working schedule. Interesting product... 
    Rad na daljinu
    Puno radno vreme
    Fleksibilno radno vreme

    Neo Group

    Beograd
    pre mesec dana
  •  ...growth partner for the world’s leading brands. With exceptional talent...  ...our pioneering agentic marketing platform, WPP Open – we help clients...  ...wppmedia.com . Job Title: Search Lead Department: Media...  ...appreciate all applications received, only those candidates selected for... 
    Ugovor o radu
    Rad u kancelariji
    Rad od kuće
    Hibridni rad

    wppmedia

    Beograd
    pre mesec dana
  •  ...Description About the Role We are hiring a Cloud DevOps Engineer to build and operate our AWS infrastructure, improve deployment...  ...in a software-defined world with our industry-first cloud-based platform that unites every stakeholder and phase of electronics... 
    Puno radno vreme
    Dežurstvo

    Altium

    Beograd
    pre 21 dan
  •  ...Do you enjoy leading high-performing engineering and consulting teams to build impactful CRM solutions? Are you passionate about modern platform engineering using Microsoft Dynamics 365 | Sales & Customer...  ...platform maturity, and deliver reliable, high‑impact CRM solutions that... 
    Puno radno vreme
    Hibridni rad

    SoftwareOne

    Beograd
    pre 1 dan
  •  ...considered yourself a natural leader? Are you proud of your Bulgarian...  ...one we are looking for! While leading a team of agents, supporting exceptional...  ...What you will get from us:   Remote model of work Opportunity to...  ...new role, apply and join us! Only shortlisted candidates will be... 
    Rad na daljinu

    M Plus

    Beograd
    pre 8 dana
  •  ...Location: Fully Remote (Global) Compensation: $24k to $36k USD...  ...impact-driven Senior Systems Engineer to own, scale, and protect our...  ...performance. 2. Deliverability Engine & Authentication Protocol...  ...aggregate DMARC reports using platforms like DMARCian, Valimail, or Postmark... 
    Rad na daljinu
    Puno radno vreme
    Rad u turnusima

    Mission Inbox

    Beograd
    pre mesec dana
  •  ...As one of the leading insurance companies, we know that together we...  ...across SEE markets. DevOps Engineer (m/f/d) What You Will Be Doing...  ...goals. Benefits Hybrid/remote work options Flexible working...  ...for their interest; however, only shortlisted candidates will be... 
    Rad na daljinu
    Hibridni rad
    Popodnevna smena
    Fleksibilno radno vreme

    SEE Digital d.o.o.

    Beograd
    pre 9 dana
  •  ...based programs built on our own platform and curriculum, developed in...  ...Nebius AI. Our team is fully remote and globally distributed, and...  ...a program for experienced engineers transitioning into system architecture...  ...(LLM integration, RAG, ML reliability, AI governance) Module 8 —... 
    Rad na daljinu
    Rad u kancelariji
    Fleksibilno radno vreme

    TripleTen

    Beograd
    pre 9 dana
  •  ...Publishing As Switzerland's leading digital hub, we provide our media and platforms with enabling technology solutions...  ...and user experience to engineering, quality assurance and platform operations...  ...are looking for a Product Studio Lead who will build and scale this organization... 
    Rad u kancelariji

    TX Services

    Beograd
    pre 22 dana
  •  ...based programs built on our own platform and curriculum, developed in...  ...Nebius AI. Our team is fully remote and globally distributed, and...  ...a program for experienced engineers transitioning into system architecture...  ...(LLM integration, RAG, ML reliability, AI governance) Module 8 —... 
    Rad na daljinu
    Rad u kancelariji
    Fleksibilno radno vreme

    TripleTen

    Beograd
    pre 3 dana
  •  ...highly skilled and experienced Senior DevOps Engineer to join our dynamic technology team. We...  ...(CI/CD) processes, and overall system reliability. Job Overview The Senior DevOps...  ...) ~ Expert-level knowledge of cloud platforms ( Azure , AWS or GCP) ~ Proficiency in... 
    Dežurstvo

    KMS Lighthouse

    Beograd
    pre 2 dana
  •  ...Some of our key Benefits: ~ Competitive salary ~ Global Remote working for up to 2 week per year for those who are able to work...  ...your birthday, move house and volunteer ~ Access to a wellbeing platform, Rezilient ~ Partnership with SOS Children's Village ~ Private... 
    Rad na daljinu
    Stalno zaposlenje
    Rad u kancelariji
    Rad od kuće
    Hibridni rad
    40 sati nedeljno
    Ponedeljak-petak
    Smene po rasporedu
    Smenski rad

    The Opportunity Hub UK

    Beograd
    pre 1 dan
  •  ...management of our marketplace activities (e.g. Amazon, eBay, and other platforms) with a focus on performance marketing in Europe Planning,...  ...in an open-minded multinational team Flexibility with remote work, allowing you to create your ideal work environment All... 
    Rad na daljinu
    Puno radno vreme
    Rad u kancelariji
    Hibridni rad

    Holycode

    Beograd
    pre 14 dana
  •  ...design and build the back-office platform that powers our games. The...  ...Optimize system performance, reliability, and maintainability ~...  ...with game developers, frontend engineers, QA, and product managers ~Participate...  ...based on experience ~Remote or hybrid work options ~Long... 
    Rad na daljinu
    Rad u kancelariji
    Hibridni rad

    Spinsoft Gaming

    pre 1 dan
  •  ...performance Have hands-on experience with the Klaviyo email marketing platform Have a proactive, ownership-driven mindset with a strong...  .... For our Vilnius office, we follow a 4 days onsite, 1 day remote hybrid model. For roles based outside Vilnius, the work... 
    Rad na daljinu
    Rad u kancelariji
    Rad sa bilo kog mesta
    Hibridni rad
    Popodnevna smena

    Kilo

    Beograd
    pre 8 sati
  • We are looking for a Senior/Lead Full-Stack Python Developer to drive...  ...for building scalable backend engines, sophisticated data...  ...performance financial or trading platforms....  ...real-time data needs, and cross-platform (Linux/Windows) requirements.... 

    Luxoft

    Beograd
    pre 3 dana
  •  ...are looking for a  Technical Lead: Level : Senior Type of...  ...(3 days in the office, 2 days remote)   The Opportunity: Our...  ...operating high-performance digital platforms at scale. Their...  ...together professionals across engineering, design, QA, DevOps and operations... 
    Rad na daljinu
    Puno radno vreme
    Rad u kancelariji
    Hibridni rad

    People Focus d.o.o.

    Beograd
    pre 2 meseci
  •  ...Childcare Marketing Group | Remote-First | Full-Time Position...  ...Marketer owns the paid media engine behind a portfolio of child care...  ...turning ad spend into qualified leads for the families and providers...  ...know what to do next.  • Platform Fluent: You move confidently... 
    Rad na daljinu
    Puno radno vreme
    Rad u kancelariji
    Rad od kuće
    Dodatna zarada
    Smenski rad

    Childcare Marketing Group

    Beograd
    pre mesec dana
  •  ...As one of the leading insurance companies, we know that together we...  ...across SEE markets. Cloud Engineer  (m/f/d)  What You Will Be...  ...goals. Benefits Hybrid/remote work options Flexible working...  ...for their interest; however, only shortlisted candidates will be... 
    Rad na daljinu
    Hibridni rad
    Popodnevna smena
    Fleksibilno radno vreme

    SEE Digital d.o.o.

    Beograd
    pre 9 dana
  •  ...innovation was creating the market’s only FDA-compliant at-home hormone...  ...understanding of paid social platforms (Meta, TikTok) and what makes...  ...Details The role is a remote position, with a 40-hour workweek...  ...Step 3 ‘Interview with department lead’ - Step 4 - ‘Cognitive Assessment... 
    Rad na daljinu
    Frilens
    Rad od kuće
    40 sati nedeljno
    Fleksibilno radno vreme

    Mira

    Beograd
    pre 21 dan