Site Reliability Engineer - Cycloid
Posted 6 days 12 hours ago by LCA Consulting Services
Site Reliability Engineer
We are seeking an experienced Site Reliability Engineer to join our platform team.
Job Description
As an SRE, you will be a key contributor to building and maintaining our internal on-prem enterprise grade automation and orchestration platforms, leveraging technologies such as n8n, Cycloid, Ansible and Kubernetes (OpenShift).
You will help design, deploy, and operate secure, scalable orchestration services, enabling internal teams to build their own automation workflows. You will also drive automated platform testing, observability, disaster recovery, security, and standardized onboarding for squads and tools.
Professional skills:
Practical experience with workflow/orchestration platforms (n8n or comparable tools): workflow design, triggers/webhooks, resilient execution (timeouts, retries, error handling, dead-letter-like patterns), secrets and credentials management, environment promotion, versioning, reusable components/templates, and governance/guardrails for safe multi-team adoption
Deep knowledge of containerization and container orchestration: Docker, Kubernetes, and Kubernetes operators; OpenShift experience is a plus
Experience operating and supporting IT services at scale, including on-call participation
Familiarity with GitOps principles and tooling; GitLab experience is a plus
Proficiency in Scripting and automation (eg, Python, Bash, Go)
Understanding of Platform Engineering principles and practices
Working knowledge of observability practices (monitoring, logging, alerting, tracing)
Strong problem-solving skills and the ability to perform under pressure Activities:
Implement highly scalable, resilient, and performant platform solutions
Identify and develop cross-cutting reusable building blocks to be provided to domain applications
Automate infrastructure provisioning and configuration to reduce manual interventions and improve efficiency
Develop automated tests for internal platforms