Site Reliability Engineer (SRE)
О должности
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of critical systems by combining software engineering with infrastructure and operations expertise. The role focuses on automation, cloud infrastructure, monitoring, incident management, and close collaboration with engineering teams to deliver highly available and resilient services.
+ ' ' +Strong knowledge of computer science fundamentals (data structures, algorithms, OS, networking);
Software development experience with at least one modern language (e.g., Go, Python, Java, C#);
Experience designing and supporting distributed systems and microservices;
Proficiency in troubleshooting complex issues in production environments;
Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry);
Working knowledge of Kubernetes and containerized infrastructure;
Familiarity with cloud platforms (AWS, GCP, Azure);
Solid understanding of Linux systems and scripting;
Experience with CI/CD pipelines and DevOps practices;
Strong problem-solving skills and a curious, analytical mindset;
Effective communicator - able to clearly explain technical concepts to both engineers and non-engineers;
Team player with a collaborative approach to working across engineering, product, and operations;
Takes ownership and initiative;
Comfortable working in a fast-paced, evolving environment;
Attention to detail while keeping an eye on the bigger picture;
Eager to learn continuously and stay up to date with emerging technologies and practices.
+ ' ' +- Opportunities for professional growth and development;
- Competitive salary and bonuses;
- Comprehensive insurance coverage;
- Supportive work environment;
- Visa Premium salary card;
- Corporate discounts and events;
- Additional vacation days;
- Discounted education and employee loans;
- Multicultural environment with foreign colleagues sharing their best experience.
Collaborate with development and infrastructure teams to design and maintain scalable, reliable systems;
Write automation and tooling for infrastructure management, deployments, and operational tasks;
Build and maintain observability stacks (metrics, logs, traces) to ensure visibility and fast issue resolution;
Lead and participate in troubleshooting and debugging efforts across the full stack - application, platform, and infrastructure;
Conduct post-incident analysis, drive root cause investigations, and implement long-term fixes;
Define and monitor service level objectives (SLOs), indicators (SLIs), and error budgets;
Contribute to CI/CD pipelines and infrastructure as code efforts. - Continuously seek to improve system performance, resilience, and developer experience
Kapital Bank iş mühiti, əlavə fürsətlər və digər vakansiyaları görüntüləmək üçün Kapital Bank Life səhifəsinə keçid edin.
Kapital Bank Azərbaycan Əmanət Bankının varisi kimi 140 ildən çox, uğurla fəaliyyət göstərir. Hazırda Kapital Bank Azərbaycanda ən böyük xidmət şəbəkəsinə malik maliyyə qurumudur. Universal bank olan…